Multi-terminal Translation Method, Device, Wearable Device, Terminal Device and Storage Medium

By establishing target groups between multi-end devices and using image information and real-time audio transmission to achieve multi-end translation, the problems of convenience and efficiency of existing translation devices and software in multi-language communication scenarios are solved, and efficient and accurate multi-end translation services are achieved.

CN118673936BActive Publication Date: 2025-06-27BEIJING SUPERHEXA CENTURY TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410975511.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2025-06-27
Estimated Expiration
2044-07-19

AI Technical Summary

Technical Problem

Existing translation equipment and software have problems of convenience and inefficiency in multilingual communication scenarios, especially the lack of effective solutions in real-time translation between multi-end devices.

Method used

By establishing a target group between the first device and the second device, establishing a group using image information, and transmitting and translating audio information in real time, the convenience and efficiency of multi-terminal translation are achieved.

Benefits of technology

It improves the accuracy and efficiency of translation, provides a more considerate and efficient language communication experience, meets users' diverse needs in different scenarios, and achieves seamless cross-language communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118673936B_ABST
    Figure CN118673936B_ABST
Patent Text Reader

Abstract

The present disclosure provides a multi-terminal translation method, apparatus, wearable device, terminal device and storage medium, belonging to the technical field of device interaction. The method includes: sending first image information to a second device, where the first image information is image information of the environment where the first device is located, and the first image information is used to instruct the second device to establish a target group according to the received multiple pieces of first image information. In response to receiving a first connection instruction sent by the second device, sending a first confirmation instruction to the second device, where the first confirmation instruction is used to confirm that the first device joins the target group. In response to the first device being in a first response mode, receiving first translation information sent by the second device, where the first translation information is obtained by the second device translating audio information sent by a user other than the target user. The present disclosure enables the target user to seamlessly understand and participate in a multilingual environment without relying on other translation tools or interrupting communication, improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the technical field of device interaction, and more specifically, relates to a multi-terminal translation method, device, wearable device, terminal device, and storage medium. Background Art

[0002] In the current context of the continuous in-depth development of globalization, the communication needs between users of different languages are increasing day by day. Whether it is in international business meetings, academic exchanges, travel, or cross-language communication scenarios in daily life, accurate and efficient translation has become crucial.

[0003] Traditional translation methods usually rely on dedicated translation devices or software, but these methods often have certain limitations. For example, a single translation device may have relatively simple functions and can only perform basic text translation, unable to meet the diverse scenario requirements; while some translation software, although rich in functions, requires frequent operation of terminal devices such as mobile phones and computers during use, the operation is rather cumbersome, and it is extremely inconvenient to use in certain specific scenarios.

[0004] With the rapid development of wearable device and intelligent terminal technologies, people have higher expectations for the implementation method and usage experience of the translation function. How to provide users with more convenient, efficient, and accurate translation services has become an urgent problem to be solved in the technical field of device interaction. Summary of the Invention

[0005] The purpose of the present disclosure is to provide a multi-terminal translation method, device, wearable device, terminal device, and storage medium, which can provide users with more convenient, efficient, and accurate translation services, and improve the user experience in multi-language translation scenarios.

[0006] In the first aspect of the embodiments of the present disclosure, a multi-terminal translation method is provided, which is applied to a first device and includes:

[0007] Sending first image information to a second device, where the first image information is image information of the environment where the first device is located, and the first image information is used to instruct the second device to establish a target group according to the received multiple first image information, and the second device is a device having a connection relationship with the first device;

[0008] In response to receiving a first connection instruction sent by the second device, sending a first confirmation instruction to the second device, where the first confirmation instruction is used to confirm that the first device joins the target group; the first connection instruction is an instruction generated by the second device based on the target group;

[0009] In response to the first device being in the first response mode, receive the first translation information sent by the second device, where the first translation information is obtained by the second device translating audio information sent by a user other than the target user, and the target user is the user wearing the first device.

[0010] In a second aspect of the embodiments of the present disclosure, there is provided a multi-terminal translation method applied to a second device, including:

[0011] Receive multiple first image information sent by multiple first devices, calculate the similarity between the multiple first image information; establish a target group for the multiple first devices according to the similarity, and send a first connection instruction to all first devices in the target group;

[0012] The first connection instruction is used to instruct the first device to send a first confirmation instruction to the second device when the first device receives the first connection instruction sent by the second device;

[0013] In response to receiving the first confirmation instruction, determine that the first device joins the target group, translate audio information sent by a user other than the target user to obtain first translation information, and send the first translation information to the first device.

[0014] In a third aspect of the embodiments of the present disclosure, there is provided a multi-terminal translation device applied to a first device, including:

[0015] A first sending module, configured to send first image information to a second device, where the first image information is image information of the environment where the first device is located, and the first image information is used to instruct the second device to establish a target group according to the received multiple first image information, and the second device is a device having a connection relationship with the first device;

[0016] A first confirmation module, configured to send a first confirmation instruction to the second device in response to receiving the first connection instruction sent by the second device, where the first confirmation instruction is used to confirm that the first device joins the target group; the first connection instruction is an instruction generated by the second device based on the target group;

[0017] A first receiving module, configured to receive the first translation information sent by the second device in response to the first device being in the first response mode, where the first translation information is obtained by the second device translating audio information sent by a user other than the target user, and the target user is the user wearing the first device.

[0018] In a fourth aspect of the embodiments of the present disclosure, there is provided a multi-terminal translation device applied to a second device, including:

[0019] A calculation module, configured to receive multiple first image information sent by multiple first devices, calculate the similarity between the multiple first image information; establish a target group for the multiple first devices according to the similarity, and send a first connection instruction to all first devices in the target group; the first connection instruction is used to instruct the first device to send a first confirmation instruction to the second device when receiving the first connection instruction sent by the second device.

[0020] A second confirmation module, configured to, in response to receiving the first confirmation instruction, determine that the first device joins the target group, translate audio information sent by users other than the target user to obtain first translation information, and send the first translation information to the first device.

[0021] In a fifth aspect of the embodiments of the present disclosure, a wearable device is provided, including a first memory, a first processor, and a computer program stored in the first memory and running on the first processor. When the first processor executes the computer program, the steps of the multi-terminal translation method in the first aspect are implemented.

[0022] In a sixth aspect of the embodiments of the present disclosure, a terminal device is provided, including a second memory, a second processor, and a computer program stored in the second memory and running on the second processor. When the second processor executes the computer program, the steps of the multi-terminal translation method in the second aspect are implemented.

[0023] In a seventh aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above multi-terminal translation method are implemented.

[0024] The beneficial effects of the multi-terminal translation method, device, wearable device, terminal device, and storage medium provided by the embodiments of the present disclosure are as follows:

[0025] On the one hand, the second device can dynamically establish a target group for users according to the image information provided by the first device, ensuring that users in the same target group can communicate in real time. This multi-terminal translation method not only improves the accuracy of translation, but also provides a more considerate and efficient language communication experience for users by continuously optimizing the translation model. At the same time, by dynamically adjusting service content and strategies, the diverse needs of users in different scenarios are met.

[0026] On the other hand, by the first device receiving the translation information sent by the second device in real time, the target user (i.e., the user wearing the first device) can seamlessly understand and participate in a multilingual environment without relying on other translation tools or interrupting communication. This instant translation service greatly improves the efficiency and convenience of communication, making cross-language communication easy and natural. Description of the Drawings

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0028] Figure 1 Schematic diagram of the application scenario of the multi-terminal translation method provided by an embodiment of the present disclosure;

[0029] Figure 2 Schematic diagram of the process of the multi-terminal translation method provided by an embodiment of the present disclosure;

[0030] Figure 3 Schematic diagram of the process of the multi-terminal translation method provided by another embodiment of the present disclosure;

[0031] Figure 4 Schematic diagram of the signaling interaction between devices provided by an embodiment of the present disclosure;

[0032] Figure 5 Block diagram of the structure of the multi-terminal translation device provided by an embodiment of the present disclosure;

[0033] Figure 6 Block diagram of the structure of the multi-terminal translation device provided by another embodiment of the present disclosure;

[0034] Figure 7 Schematic block diagram of a wearable device provided by an embodiment of the present disclosure;

[0035] Figure 8 Schematic block diagram of a terminal device provided by an embodiment of the present disclosure. Detailed Embodiments

[0036] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.

[0037] To make the purpose, technical solutions, and advantages of the present disclosure clearer, the following will be described through specific embodiments with reference to the drawings.

[0038] Please refer to Figure 1 and Figure 2 , Figure 1Schematic diagram of the application scenario of the multi-terminal translation method provided by an embodiment of the present disclosure. Figure 2 Flow chart of the multi-terminal translation method provided by an embodiment of the present disclosure. The method includes:

[0039] S101: Send the first image information to the second device. The first image information is the image information of the environment where the first device is located. The first image information is used to instruct the second device to establish a target group according to the received multiple first image information. The second device is a device having a connection relationship with the first device.

[0040] In this embodiment, the first device refers to the device worn by the target user. The first device can be a Bluetooth glasses, a smart audio glasses, etc. The first image information refers to the image information of the environment around the target user collected by the first device. The first image information may include multiple images. For example, the first image information may be the environmental image information in front of the target user's line of sight captured by the camera built in the smart glasses. These image information can be static pictures, video streams or any form of visual data, used to describe or display the current environmental state of the first device.

[0041] The second device can be a device that receives the image information from the first device and has a certain form of connection relationship with the first device. This connection can be Wi-Fi, Bluetooth, Internet connection, etc. The second device may also be a smart phone, a tablet computer, a desktop computer, a cloud server or any other terminal device capable of receiving and processing image information. The target group is a group established by the second device according to the received multiple first image information (from multiple first devices). This group can be created based on certain common features or conditions (such as geographical location, time, image content similarity). Since the second device can process the received multiple first image information at the same time, in a basic multi-terminal translation system, there can be one second device and multiple first devices.

[0042] Reference Figure 1, the first device can be the Bluetooth glasses 200, and the second device can be the server 300. Before the Bluetooth glasses 200 and the server 300 exchange information, a communication connection needs to be established first, which is usually achieved through a wireless network (such as Wi-Fi, Bluetooth, or cellular network) or a wired connection. Once the connection is successful, the two devices can start data exchange. For example, after the user 100 wears the Bluetooth glasses 200, the Bluetooth function of the Bluetooth glasses 200 can be turned on through triggering methods such as pressing a button, pressing and sliding, or double-clicking the side of the glasses temple, so that the Bluetooth glasses 200 can be connected to the server 300. After the Bluetooth glasses 200 and the server 300 establish a communication connection, the image acquisition device built into the Bluetooth glasses 200 can collect the first image information in front of the line of sight of the target user 100 and send the first image information to the server 300. The server 300 can use image processing technology and user behavior analysis to identify user groups with similar environmental characteristics or common activity areas. The first devices corresponding to these user groups are dynamically determined by the second device as the same "target group". For example, users participating in the same meeting or visiting the same scenic spot.

[0043] After establishing the target group and obtaining the environmental information of the user wearing the first device, the second device can provide a more accurate translation service. For example, when it is detected that the user group is in a foreign language environment, the second device can translate the audio information of the target user who is speaking in the target group into a language that other users in the group can understand. At the same time, the second device can also continuously optimize the accuracy and response speed of the translation model according to the interactions and feedback within the target group, improving the overall user experience.

[0044] S102: In response to receiving the first connection instruction sent by the second device, send a first confirmation instruction to the second device. The first confirmation instruction is used to confirm that the first device joins the target group; the first connection instruction is an instruction generated by the second device based on the target group.

[0045] In this embodiment, the first connection instruction can be an instruction sent by the second device to the first device, and its purpose is usually to invite or instruct the first device to join the target group. This instruction can contain the necessary information required for the first device to join the target group, such as group identifier, security key, or connection parameters, etc.

[0046] The first confirmation instruction can be a response instruction sent by the first device to the second device after receiving the first connection instruction. The purpose of this instruction is to confirm that the first device has successfully received the connection instruction and agrees to join the specified target group. By sending the first confirmation instruction, the first device indicates to the second device that it is ready to participate in the communication or activities within the group.

[0047] When the second device determines that the first devices corresponding to the first image information with relatively high similarity are in the same environment, meeting, or other scenarios, it will send a first connection instruction to these first devices. After receiving the first connection instruction, the first device (such as Bluetooth glasses) will first verify the validity of the instruction, confirm that the instruction indeed comes from the second device connected to it and is targeted at the current device. Then the first device will parse the instruction content to understand which specific target group it is invited to join, which usually involves identifying the group identifier (such as group ID). Once the validity of the first connection instruction is verified and the group information is confirmed, the first device will generate a first confirmation instruction, which is used to confirm to the second device that the first device is ready and willing to join the specified target group. After receiving the first confirmation instruction, the second device will update the member list of the target group and mark the device that sent the confirmation instruction as a member who has joined the target group. During the process of determining the target group members, the second device can perform additional authentication steps to ensure that only authorized devices can join the group. For example, through the first device ID, user account information, etc. After the first device joins the target group, the second device can start providing translation services, information sharing, or other related functions to all members in the group. From the perspective of the second device, the management of the target group is dynamically changing, and members can be added or deleted as needed, or the service content and policies can be adjusted according to the real-time behavior of the group members. From the user's perspective, the entire process of joining the target group should be seamless and automatic. They may only need to wear the first device and turn on the relevant services to immediately enjoy the convenience brought by group translation.

[0048] S103: In response to the first device being in the first response mode, receive the first translation information sent by the second device, where the first translation information is obtained by the second device translating the audio information sent by users other than the target user, and the target user is the user wearing the first device.

[0049] In this embodiment, the first response mode is a working state of the first device, indicating that the device is currently ready to receive and process information from other devices (such as the second device). This mode can be the working state after the first device sends the first confirmation instruction to the second device. When the first device is in the first response mode, it can receive and process the first translation information sent by the second device, thereby providing real-time translation services for users.

[0050] The first translation information is the result obtained by translating the audio information sent by the second device to users other than the target user. It contains the translated text of the original audio information and may also include other relevant information (such as the identity of the speaker, timestamp, etc.). The first translation information is crucial for the target user to understand the speech of other users, making language no longer a barrier to communication and facilitating communication between users of different languages.

[0051] From the above, on the one hand, the second device can dynamically establish a target group for users based on the image information provided by the first device, ensuring real-time communication among users in the same target group. This multi-device translation method not only improves the accuracy of translation but also provides a more considerate and efficient language communication experience for users by continuously optimizing the translation model. At the same time, by dynamically adjusting the service content and strategies, it meets the diverse needs of users in different scenarios.

[0052] On the other hand, by the first device receiving the translation information sent by the second device in real time, the target user (i.e., the user wearing the first device) can seamlessly understand and participate in the multilingual environment without relying on other translation tools or interrupting the communication. This instant translation service greatly improves the efficiency and convenience of communication, making cross-language communication easy and natural.

[0053] In an embodiment of the present disclosure, before sending the first image information to the second device, it further includes:

[0054] Collecting the image information of the first area at the first acquisition frequency to obtain the first image information, where the first area is the area where the target user's gaze duration is greater than the first duration;

[0055] Or, collecting the image information of the second area at the second acquisition frequency to obtain the first image information, where the second area is the environmental area where the target user is located.

[0056] In this embodiment, the first image information is the image information collected by the image acquisition device on the first device. There are two methods for image acquisition. The first one is: the first device records and calculates the time when the target user gazes at each area. When the gaze duration of a certain area exceeds the preset first duration, it is determined that this area is an important area, and the image information in this important area is collected at the first acquisition frequency to obtain higher-quality or more detailed image information. For example, for each user in a meeting state (generally looking in the same direction), the first device worn by them determines the line-of-sight area in the direction of the podium as the first area (important area) according to the time when the user gazes at each area, then collects the image information in the first area at the first acquisition frequency and sends this image information to the second device for corresponding data processing.

[0057] The second method for image acquisition is as follows: The first device acquires image information within a 360-degree range of the environment where the first device is located at a second acquisition frequency (which can be a fixed frequency or a dynamically adjusted frequency). For example, during a trade fair, when users (generally looking in different directions) are negotiating, the first devices they wear can acquire environmental image data within a 360-degree range around them and send this image information to the second device for corresponding data processing.

[0058] When the second device receives the first image information acquired by the two image acquisition methods, it uses two different processing methods to establish a target group respectively. The group establishment method corresponding to the first image acquisition method is: group establishment by image recognition. When a user wears the first device and looks in the same direction, the first device takes pictures through the built-in camera and uploads the first image information to the second device. The second device uses a deep learning algorithm to compare the similarity between all the submitted first image information. When the similarity exceeds a preset threshold, the second device automatically determines that the users have a common visual focus and suggests establishing a target group.

[0059] The group establishment method corresponding to the second image acquisition method is: group establishment by environmental feature point analysis. When a user wears the first device and scans the surrounding environment through the built-in camera, the first device takes pictures and uploads the first image information to the second device. The second device analyzes and creates a point cloud model of the environment in the first image information and analyzes the similarity of environmental feature points. When the similarity exceeds a preset threshold, the second device automatically determines that the users are in the same environment and suggests establishing a target group.

[0060] As can be seen from the above, in this embodiment, through an intelligent image acquisition and processing mechanism, the accuracy and efficiency of group establishment are effectively improved. Whether it is to collect high-quality images in the important area determined by the fixation duration or to comprehensively scan the environment to obtain environmental images, the user behavior and environmental characteristics can be accurately captured. By using image recognition and environmental feature point analysis, the visual focus or environmental similarity between users is automatically judged, and then a target group is intelligently recommended to promote the effective exchange of information and the accurate docking of resources.

[0061] In an embodiment of the present disclosure, the multi-terminal translation method further includes:

[0062] In response to recognizing the first audio information of the target user, the first audio information is sent to the second device; the first audio information is used to instruct the second device to recognize the language type in the first audio information and determine the language type as the target language type of the target user;

[0063] The first audio information is also used to instruct the second device to extract the voice features in the first audio information.

[0064] In this embodiment, the first audio information is the audio data generated by the target user, which can be collected by the first device through the built-in microphone. This audio information contains the content that users other than the target user want to hear and is the key input data in the multi-terminal translation process.

[0065] The language type refers to the type of natural language used in the audio information, such as Chinese, English, French, etc. Identifying the language type is an important step in the translation process, which can determine which translation model or strategy should be used for translation. The target language type is the native language type of the target user, that is, the language type that the target user can understand.

[0066] The voice characteristics can refer to other acoustic characteristics in the audio information besides the language content, such as pitch, speech rate, volume, etc. Extracting the voice characteristics helps to more accurately understand the user's intention and emotion, and may also be used to improve the translation quality or enhance the user experience during the translation process.

[0067] The built-in audio collection device (such as a microphone) in the first device can actively capture the first audio information of the target user and then send the first audio information to the second device (cloud server). After receiving the first audio information, the second device can use natural language processing technologies (such as machine learning models) to analyze the first audio information and identify the language type. Based on the language recognition result, the second device determines the language type in the audio information and uses this language type as the target language type (i.e., the native language) of the target user. When or after identifying the language type, the second device can also extract the voice characteristics from the first audio information. These voice characteristics may be used for various purposes, such as improving the accuracy of translation (by understanding intonation, speech rate, etc.), enhancing the user experience (such as through voice synthesis by imitating the user's voice), or performing other advanced processing (such as sentiment analysis). It can be concluded from the above that through real-time audio transmission and recognition, this embodiment can quickly determine the language needs of the target user and synchronously extract the voice characteristics, providing an accurate language type and personalized voice processing basis for subsequent multi-terminal translation, and significantly improving the accuracy of translation services and the user experience.

[0068] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of a multi-terminal translation method provided in another embodiment of the present disclosure. The method includes:

[0069] S201: Receive multiple first image information sent by multiple first devices, calculate the similarity between the multiple first image information; establish a target group for the multiple first devices according to the similarity, and send a first connection instruction to all first devices in the target group; the first connection instruction is used to instruct the first device to send a first confirmation instruction to the second device when receiving the first connection instruction sent by the second device.

[0070] Establish a target group for multiple first devices according to the similarity, including:

[0071] Mark the first devices corresponding to the first image information with a similarity greater than a preset threshold as target devices;

[0072] Establish a target group according to multiple target devices.

[0073] In this embodiment, the second device can receive multiple first image information sent by multiple first devices, then perform image processing on the received multiple first image information, and extract key features (such as color histograms, textures, edge information, specific object detection, etc.). Using an image feature comparison algorithm, calculate the similarity between these image information. The similarity can be quantified by comparing the similarity degree of image features, and common methods include Euclidean distance, cosine similarity, etc.

[0074] If the similarity between the image information is greater than a preset threshold (for example, 0.8, that is, 80% similarity), the first devices corresponding to the first image information with a similarity greater than the preset threshold can be marked as target devices, and then a target group is established according to multiple target devices. After establishing the target group, the second device needs to send a first connection instruction to the first devices in the target group (i.e., the target devices), informing the devices in the group that they have been recognized as paying attention to the same or similar content, so that the users wearing the first devices can enjoy multi-terminal translation services in the target group and enhance the communication experience. After receiving the first connection instruction, the first device can send a first confirmation instruction to the second device as a confirmation and response to the first connection request.

[0075] After the first confirmation instruction is received and confirmed, the connection between the second device and the target group where the first device is located is officially established. After that, data transmission, sharing, synchronization or other forms of interaction can start between the devices, depending on the requirements of the application scenario.

[0076] S202: In response to receiving the first confirmation instruction, determine that the first device joins the target group, translate the audio information sent by users other than the target user to obtain the first translation information, and send the first translation information to the first device.

[0077] In this embodiment, after the second device receives the first confirmation instruction sent by the first device, it can determine all the first devices that send the first confirmation instruction as the devices in the target group. Then, the second device can listen for and receive the audio information of the user sent by the first device, and after translating the audio information, it can send the translated first translation information to the first device. Since there are multiple first devices and the language types used by the target users corresponding to different first devices are different, there can be multiple types of first translation information, such as English translation information, Russian translation information, etc.

[0078] It can be concluded from the above that in this embodiment, by automatically establishing a group through image similarity, intelligent connection and translation services between devices can be achieved. This method can improve the user experience, quickly respond to the audio translation needs of non-target users in the group, enhance the convenience and real-time nature of information exchange, and at the same time ensure the accuracy of the translated content and privacy protection.

[0079] In an embodiment of the present disclosure, calculating the similarity between multiple first image information includes:

[0080] Extract the key feature points in each first image information, convert the key feature points in each first image information into point cloud data, and calculate the similarity between the point cloud data corresponding to each first image information.

[0081] In this embodiment, in image processing, key feature points usually refer to the points in the image that are of great significance for image recognition, matching, or description. These points can be corner points, edge points, high-curvature points, etc. in the image, and they can provide stable reference points between different images for subsequent image comparison or matching.

[0082] Point cloud data is a set of points in three-dimensional space, and these points usually represent the surface of an object. In this embodiment, converting the key feature points in the image into point cloud data means mapping the feature points in the two-dimensional image to three-dimensional space for more accurate spatial position comparison and analysis.

[0083] The similarity between point cloud data refers to evaluating the similarity degree between them by comparing the point cloud data corresponding to different images. Since point cloud data contains position information in three-dimensional space, the similarity between them can be evaluated by calculating indicators such as the spatial distance and shape matching degree between point clouds. In this embodiment, calculating the similarity between point cloud data is one of the key steps in realizing image similarity comparison and group division.

[0084] Exemplarily, first, the second device can extract key feature points from the images sent by each first device, and these feature points represent the significant elements in the images. Then, the second device converts these key feature points into point cloud data in three-dimensional space for more accurate spatial position comparison. Finally, by calculating the similarity between the point cloud data corresponding to different images (such as using spatial distance, shape matching algorithms, etc.), the second device can evaluate the similarity between the images, thereby providing a basis for establishing the target group. This process effectively utilizes the spatial information in the images and improves the accuracy and efficiency of group division.

[0085] As can be seen from the above, in this embodiment, by extracting key feature points and converting them into point cloud data to calculate the similarity, the similarity between images can be more accurately identified, and the accuracy of group division is improved. This method effectively utilizes the spatial information of the images, enhances the intelligence and personalization of the translation service between devices, and improves the user experience.

[0086] In an embodiment of the present disclosure, the multi-terminal translation method further includes:

[0087] In response to receiving the first audio information sent by the first device, identify the language type in the first audio information and determine the language type as the target language type of the target user.

[0088] In this embodiment, when the second device detects and receives the audio information from the first device, it will trigger a response mechanism. This audio information can be a voice message recorded by the user through the microphone. After receiving the audio information, the second device will use the built-in language recognition technology or algorithm to analyze the audio information. This analysis process aims to determine the language type of the user in the audio information. Modern language recognition technology is usually based on machine learning or deep learning models, which are trained with a large amount of multilingual data and can accurately identify the language in the audio.

[0089] After the second device identifies the language type in the audio information, it will store this language type as the target language type of the target user (i.e., the mother tongue of the target user) in the storage space of the second device for use during multi-terminal translation interaction.

[0090] As can be seen from the above, the second device greatly improves the personalization and intelligence level of the user experience by automatically identifying the language type in the first audio information sent by the first device and determining the target language type of the target user accordingly. It can achieve precise matching and presentation of content without the need for the user to manually set language preferences, reducing the complexity of user operations and promoting the efficient transmission and understanding of information.

[0091] In one embodiment of the present disclosure, translating the audio information sent by users other than the target user to obtain first translation information and sending the first translation information to a first device includes:

[0092] In response to receiving second audio information sent by the first device, converting the second audio information into text information;

[0093] Synthesizing the text information with the voice characteristics of the target user in the first audio information to obtain first translation information, and sending the first translation information to a third device; the third device is the first device that has not sent the second audio information.

[0094] In this embodiment, first, the second device can convert the received second audio information into text information, and then translate the text information into multiple languages. The translated text information will use the voice characteristics (such as voiceprint information) of the source voice, and through text-to-speech (TTS) technology, synthesize a voice output similar to the original speaker, including timbre, intonation, and emotion.

[0095] For example, there are 5 first devices in the target group, numbered device 1, device 2, device 3, device 4, and device 5 respectively. The target language types of the users corresponding to the 5 first devices are Chinese, English, French, Russian, and German respectively. If the target user currently speaking is the user of wearable device 1, device 1 sends the first audio information of the target user collected to the second device. The second device identifies that the language type of the first audio information is Chinese, then converts the first audio information into text information, and translates the text information into four different translated text information in English, French, Russian, and German. The second device combines the translated text information of the English language type with the voiceprint information of the user corresponding to device 1, synthesizes the translated information in the English version and sends it to device 2, and device 2 plays it to the user of wearable device 2. The translation methods for other language types are the same as above. Therefore, in this embodiment, when a user corresponding to one of the first devices in the target group speaks, other users can listen to the translated information corresponding to their mother tongue in real time.

[0096] It can be concluded from the above that in this embodiment, by translating the audio information of non-target users and integrating the voice characteristics of the target user to generate first translation information, the naturalness and fluency of communication are ensured, the immersion and authenticity of communication are enhanced, at the same time, the accuracy and efficiency of information transmission are improved, and seamless communication among multiple users is promoted.

[0097] In one embodiment of the present disclosure, translating the audio information sent by users other than the target user to obtain first translation information and sending the first translation information to a first device includes:

[0098] In response to receiving the second audio information sent by the first device, convert the second audio information into text information;

[0099] Synthesize the text information and the communication characteristics of the corresponding user of the third device to obtain the first translation information, and send the first translation information to the third device;

[0100] The third device is the first device that has not sent the second audio information.

[0101] In this embodiment, first, the second device can convert the received second audio information into text information, and then translate the text information into multiple languages. The translated text information will use the voice characteristics of the source voice (such as voiceprint information), combined with the communication characteristics of the listening user (such as speech rate, pauses in sentences, and other listening habits), and synthesize a voice output that is similar to the original speaker and has the same communication characteristics as the listening user through text-to-speech (TTS) technology. For example, in the scenario of an international conference, there are three participants in the same target group:

[0102] Participant A (the target user) uses his own device to speak, and his native language is English.

[0103] Participant B (non-target user) uses his own device to speak, and his native language is French.

[0104] Participant C (another non-target user) hopes to understand B's speech, but his native language is Chinese, and he uses his own device (the third device) to receive the translation information.

[0105] The second device receives the second audio information (French speech) sent by Participant B through his first device. Using speech recognition technology, convert the French audio information into text information. Translate the text information into multiple languages, including Chinese, to meet the needs of different participants. For the Chinese translation for Participant C, the second device not only translates the text information into Chinese, but also uses the original voice characteristics of Participant B (such as voiceprint information), combined with the communication characteristics of Participant C (such as his accustomed speech rate, sentence pauses, etc.), and synthesizes a Chinese voice that contains both B's voice characteristics and conforms to C's listening habits through text-to-speech (TTS) technology. Then the second device sends this specially synthesized Chinese first translation information to Participant C's third device. Participant C hears a voice on his own device that almost seems like Participant B is speaking directly in Chinese, enhancing the immersion and authenticity of communication.

[0106] It can be concluded from the above that the multi-terminal translation method provided in this embodiment not only solves the language barrier problem, but also greatly improves the efficiency and quality of cross-border and cross-cultural communication through an intelligent processing method.

[0107] Figure 4A signaling interaction diagram between devices provided by an embodiment of the present disclosure. Among them, the devices refer to the signaling interaction diagram among a first device, a second device, and a third device (a special first device).

[0108] A111: The first device collects first image information;

[0109] A112: The first device sends the first image information to the second device;

[0110] A113: The second device establishes a target group according to the first image information;

[0111] A114: The second device sends a first connection instruction to the first device;

[0112] A115: The first device sends a first confirmation instruction to the second device;

[0113] A116: The first device collects second audio information of the target user;

[0114] A117: The first device sends the second audio information to the second device;

[0115] A118: The second device translates the second audio information to obtain first translation information;

[0116] A119: The second device sends the first translation information to the third device.

[0117] In the above signaling interaction diagram, the first device can be a device worn by the target user who is speaking, the second device can be a server, and the third device can be a device worn by the user who is not speaking. The second device can translate the second audio information of the user who is speaking into first translation information and send the first translation information to the third device worn by the user who is not speaking.

[0118] Corresponding to the multi-terminal translation method in the above embodiment, Figure 5 A structural block diagram of a multi-terminal translation device applied to a first device provided by an embodiment of the present disclosure. For the sake of convenience of description, only the parts related to the embodiments of the present disclosure are shown. Refer to Figure 5 , the multi-terminal translation device 30 includes: a first sending module 31, a first confirmation module 32, and a first receiving module 33.

[0119] Among them, the first sending module 31 is used to send the first image information to the second device. The first image information is the image information of the environment where the first device is located. The first image information is used to instruct the second device to establish a target group according to the received multiple first image information. The second device is a device having a connection relationship with the first device;

[0120] The first confirmation module 32 is configured to send a first confirmation instruction to the second device in response to receiving a first connection instruction sent by the second device. The first confirmation instruction is used to confirm that the first device joins the target group. The first connection instruction is an instruction generated by the second device based on the target group.

[0121] The first receiving module 33 is configured to receive first translation information sent by the second device in response to the first device being in the first response mode. The first translation information is obtained by the second device translating audio information sent by a user other than the target user. The target user is the user wearing the first device.

[0122] In an embodiment of the present disclosure, the multi-terminal translation device 30 further includes an image acquisition module.

[0123] The image acquisition module: is configured to acquire image information of a first area at a first acquisition frequency to obtain first image information. The first area is a fixation area where the target user's fixation duration is greater than a first duration.

[0124] Alternatively, acquire image information of a second area at a second acquisition frequency to obtain first image information. The second area is the environmental area where the target user is located.

[0125] In an embodiment of the present disclosure, the multi-terminal translation device 30 further includes a first information recognition module.

[0126] The first information recognition module is configured to send the first audio information to the second device in response to recognizing the first audio information of the target user. The first audio information is used to instruct the second device to recognize the language type in the first audio information and determine the language type as the target language type of the target user.

[0127] The first audio information is further used to instruct the second device to extract the voice feature in the first audio information.

[0128] Corresponding to the multi-terminal translation method in another embodiment above, Figure 6 This is a structural block diagram of a multi-terminal translation device applied to a second device provided in another embodiment of the present disclosure. For ease of description, only parts related to the embodiments of the present disclosure are shown. Refer to Figure 6 and the multi-terminal translation device 40 includes: a calculation module 41, a second confirmation module 42.

[0129] Among them, the calculation module 41 is configured to receive multiple first image information sent by multiple first devices, calculate the similarity between the multiple first image information; establish a target group for the multiple first devices according to the similarity, and send a first connection instruction to all first devices in the target group. The first connection instruction is used to instruct the first device to send a first confirmation instruction to the second device when receiving the first connection instruction sent by the second device.

[0130] A second confirmation module 42, configured to determine that a first device joins a target group in response to receiving a first confirmation instruction, translate audio information sent by a user other than the target user to obtain first translation information, and send the first translation information to the first device.

[0131] In an embodiment of the present disclosure, the calculation module 41 is specifically configured to:

[0132] Extract key feature points in each piece of first image information, convert the key feature points in each piece of first image information into point cloud data, and calculate the similarity between the point cloud data corresponding to each piece of first image information.

[0133] In an embodiment of the present disclosure, the calculation module 41 is specifically configured to:

[0134] Mark the first device corresponding to the first image information with a similarity greater than a preset threshold as a target device;

[0135] Establish a target group according to multiple target devices.

[0136] In an embodiment of the present disclosure, the multi-terminal translation device 40 further includes a second information recognition module;

[0137] The second information recognition module is configured to recognize the language type in the first audio information in response to receiving the first audio information sent by the first device, and determine the language type as the target language type of the target user.

[0138] In an embodiment of the present disclosure, the second confirmation module 42 is specifically configured to:

[0139] In response to receiving the second audio information sent by the first device, convert the second audio information into text information;

[0140] Synthesize the text information with the voice characteristics of the target user in the first audio information to obtain first translation information, and send the first translation information to a third device; the third device is the first device that has not sent the second audio information.

[0141] In an embodiment of the present disclosure, the second confirmation module 42 is specifically configured to:

[0142] In response to receiving the second audio information sent by the first device, convert the second audio information into text information;

[0143] Synthesize the text information with the communication characteristics of the user corresponding to the third device to obtain first translation information, and send the first translation information to the third device;

[0144] The third device is the first device that has not sent the second audio information.

[0145] SeeFigure 7 , Figure 7 is a schematic block diagram of a wearable device provided by an embodiment of the present disclosure. As Figure 7 shown, the wearable device 500 in this embodiment may include: one or more first processors 501, one or more first input devices 502, one or more first output devices 503, and one or more first memories 504. The above-mentioned first processor 501, first input device 502, first output device 503, and first memory 504 communicate with each other through a first communication bus 505. The first memory 504 is used to store a computer program, and the computer program includes program instructions. The first processor 501 is used to execute the program instructions stored in the first memory 504. Among them, the first processor 501 is configured to call the program instructions to execute the functions of each module / unit in the above device embodiments, for example Figure 5 the functions of the modules 31 to 33 shown.

[0146] It should be understood that in the embodiments of the present disclosure, the so-called first processor 501 may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.

[0147] The first input device 502 may include a touchpad, a fingerprint acquisition sensor (for acquiring the fingerprint information and the direction information of the fingerprint of the user), a microphone, etc., and the first output device 503 may include a display (such as an LCD), a speaker, etc.

[0148] The first memory 504 may include a read-only memory and a random access memory, and provide instructions and data to the first processor 301. A part of the first memory 504 may also include a non-volatile random access memory. For example, the first memory 504 may also store information about the device type.

[0149] In a specific implementation, the first processor 501, the first input device 502, and the first output device 503 described in the embodiments of the present disclosure may implement the implementation manners described in the first and second embodiments of the multi-terminal translation method provided by the embodiments of the present disclosure, and may also implement the implementation manner of the wearable device described in the embodiments of the present disclosure, which will not be elaborated herein.

[0150] See Figure 8 , Figure 8 which is a schematic block diagram of a terminal device provided by an embodiment of the present disclosure. As Figure 8 shown, the wearable device 600 in this embodiment may include: one or more second processors 601, one or more second input devices 602, one or more second output devices 603, and one or more second memories 604. The above-mentioned second processors 601, second input devices 602, second output devices 603, and second memories 604 communicate with each other through a second communication bus 605. The second memory 604 is used to store a computer program, and the computer program includes program instructions. The second processor 601 is used to execute the program instructions stored in the second memory 604. Among them, the second processor 601 is configured to call the program instructions to execute the functions of each module / unit in the above device embodiments, such as Figure 6 the functions of the modules 41 to 42 shown.

[0151] It should be understood that in the embodiments of the present disclosure, the so-called second processor 601 may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.

[0152] The second input device 602 may include a touchpad, a fingerprint acquisition sensor (for acquiring the fingerprint information and the fingerprint direction information of the user), a microphone, etc., and the second output device 603 may include a display (such as an LCD), a speaker, etc.

[0153] The second memory 604 may include a read-only memory and a random access memory, and provide instructions and data to the second processor 301. A part of the second memory 604 may also include a non-volatile random access memory. For example, the second memory 604 may also store information about the device type.

[0154] In a specific implementation, the second processor 601, the second input device 602, and the second output device 603 described in the embodiments of the present disclosure may implement the implementation manners described in the first and second embodiments of the multi-terminal translation method provided by the embodiments of the present disclosure, and may also implement the implementation manner of the wearable device described in the embodiments of the present disclosure, which will not be elaborated herein.

[0155] In another embodiment of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, all or part of the processes in the methods of the above embodiments are implemented. It can also be completed by instructing relevant hardware through the computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0156] The computer-readable storage medium may be an internal storage unit of the wearable device or the terminal device in any of the foregoing embodiments, such as the hard disk or memory of the terminal device. The computer-readable storage medium may also be an external storage device of the wearable device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the wearable device. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the wearable device. The computer-readable storage medium is used to store the computer program and other programs and data required by the wearable device. The computer-readable storage medium may also be used to temporarily store the data that has been output or will be output.

[0157] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this disclosure.

[0158] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the wearable device and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0159] In several embodiments provided in this application, it should be understood that the disclosed wearable device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces or units, and can also be electrical, mechanical, or other forms of connection.

[0160] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of this disclosure.

[0161] In addition, the functional units in various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0162] The above is only the specific implementation manner of this disclosure, but the protection scope of this disclosure is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or substitutions within the technical scope disclosed by this disclosure, and these modifications or substitutions should be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be subject to the protection scope of the claims.

Claims

1. A multi-terminal translation method, characterized in that: Applied to a first device, comprising: Sending first image information to a second device, where the first image information is image information of the environment in which the first device is located, and the first image information is used to instruct the second device to establish a target group according to the received multiple first image information, and the second device is a device connected to the first device; the second device uses image processing technology and user behavior analysis to identify a user group with similar environmental characteristics or a common activity area, and determines the first device corresponding to the user group as a device of the target group; In response to receiving a first connection instruction sent by the second device, sending a first confirmation instruction to the second device, where the first confirmation instruction is used to confirm that the first device joins the target group; the first connection instruction is an instruction generated by the second device based on the target group; In response to the first device being in a first response mode, receiving first translation information sent by the second device, where the first translation information is obtained by the second device translating audio information sent by a user other than a target user, where the target user is a user wearing the first device; Before sending the first image information to the second device, the multi-terminal translation method further includes: Collecting image information of a first area at a first collection frequency to obtain first image information, wherein the first area is a gaze area where a target user gazes for a duration greater than a first duration; Alternatively, image information of a second area is collected at a second collection frequency to obtain first image information, and the second area is an environment area where the target user is located.

2. The multi-terminal translation method according to claim 1, characterized in that: Also includes: In response to identifying the first audio information of the target user, sending the first audio information to the second device; The first audio information is used to instruct the second device to identify a language type in the first audio information and determine the language type as a target language type of a target user; The first audio information is further used to instruct the second device to extract sound features from the first audio information.

3. A multi-terminal translation method, characterized in that: Applied to the second device, comprising: receiving a plurality of first image information sent by a plurality of first devices, and calculating the similarity between the plurality of first image information; establishing a target group for the plurality of first devices according to the similarity, and sending a first connection instruction to all the first devices in the target group; the second device uses image processing technology and user behavior analysis to identify a user group with similar environmental features or a common activity area, and determines the first device corresponding to the user group as a device of the target group; The first connection instruction is used to instruct the first device to send a first confirmation instruction to the second device when receiving the first connection instruction sent by the second device; In response to receiving the first confirmation instruction, determining that the first device joins the target group, translating audio information sent by users other than the target user to obtain first translation information, and sending the first translation information to the first device; The first device is also used for: Collecting image information of a first area at a first collection frequency to obtain first image information, wherein the first area is a gaze area where a target user gazes for a duration greater than a first duration; Alternatively, image information of a second area is collected at a second collection frequency to obtain first image information, and the second area is an environment area where the target user is located.

4. The multi-terminal translation method according to claim 3, characterized in that: The calculating the similarity between the plurality of first image information comprises: Extract key feature points from each first image information, convert the key feature points in each first image information into point cloud data, and calculate the similarity between the point cloud data corresponding to each first image information.

5. The multi-terminal translation method according to claim 3, characterized in that: The establishing a target group for the plurality of first devices according to the similarity comprises: Marking a first device corresponding to the first image information whose similarity is greater than a preset threshold as a target device; Create target groups based on multiple target devices.

6. The multi-terminal translation method according to claim 3, characterized in that: Also includes: In response to receiving the first audio information sent by the first device, a language type in the first audio information is identified, and the language type is determined as a target language type of a target user.

7. The multi-terminal translation method according to claim 6, characterized in that: The translating the audio information sent by users other than the target user to obtain first translation information, and sending the first translation information to the first device includes: In response to receiving second audio information sent by the first device, converting the second audio information into text information; The text information and the voice feature of the target user in the first audio information are synthesized to obtain first translation information, and the first translation information is sent to a third device; the third device is the first device that does not send the second audio information.

8. The multi-terminal translation method according to claim 6, characterized in that: The translating the audio information sent by users other than the target user to obtain first translation information, and sending the first translation information to the first device includes: In response to receiving second audio information sent by the first device, converting the second audio information into text information; synthesizing the text information and the communication characteristics of the user corresponding to the third device to obtain first translation information, and sending the first translation information to the third device; The third device is the first device that does not send the second audio information.

9. A multi-terminal translation device applied to a first device, characterized in that: include: A first sending module, configured to send first image information to a second device, wherein the first image information is image information of an environment in which the first device is located, and the first image information is used to instruct the second device to establish a target group according to the received plurality of first image information, and the second device is a device connected to the first device; The second device is a device that uses image processing technology and user behavior analysis to identify a user group with similar environmental characteristics or a common activity area, and determines the first device corresponding to the user group as a target group device; A first confirmation module, configured to send a first confirmation instruction to the second device in response to receiving a first connection instruction sent by the second device, wherein the first confirmation instruction is used to confirm that the first device joins the target group; the first connection instruction is an instruction generated by the second device based on the target group; a first receiving module, configured to receive, in response to the first device being in a first response mode, first translation information sent by the second device, the first translation information being obtained by the second device translating audio information sent by a user other than a target user, the target user being a user wearing the first device; An image acquisition module, configured to acquire image information of a first area at a first acquisition frequency to obtain first image information, wherein the first area is a gaze area where a target user gazes for a duration greater than a first duration; Alternatively, image information of a second area is collected at a second collection frequency to obtain first image information, and the second area is an environment area where the target user is located.

10. A multi-terminal translation device applied to a second device, characterized in that: include: A calculation module, used for receiving a plurality of first image information sent by a plurality of first devices, and calculating similarities between the plurality of first image information; A target group is established for multiple first devices according to the similarity, and a first connection instruction is sent to all first devices in the target group; the first connection instruction is used to instruct the first device to send a first confirmation instruction to the second device when receiving the first connection instruction sent by the second device; the second device uses image processing technology and user behavior analysis to identify a user group with similar environmental characteristics or a common activity area, and determines the first device corresponding to the user group as a device of the target group; a second confirmation module, configured to, in response to receiving the first confirmation instruction, determine that the first device joins the target group, translate the audio information sent by users other than the target user to obtain first translation information, and send the first translation information to the first device; An image acquisition module, configured to acquire image information of a first area at a first acquisition frequency to obtain first image information, wherein the first area is a gaze area where a target user gazes for a duration greater than a first duration; Alternatively, image information of a second area is collected at a second collection frequency to obtain first image information, and the second area is an environment area where the target user is located.

11. A wearable device, comprising a first memory, a first processor, and a computer program stored in the first memory and running on the first processor, characterized in that: When the first processor executes the computer program, the steps of the method according to any one of claims 1 to 2 are implemented.

12. A terminal device comprising a second memory, a second processor, and a computer program stored in the second memory and running on the second processor, characterized in that: When the second processor executes the computer program, the steps of the method according to any one of claims 3 to 8 are implemented.

13. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Simultaneous interpretation method and device

    CN107992485A

  • Dialogue translation method and device, storage medium and electronic equipment

    CN111985252A

  • Group joining method and equipment based on head-mounted display equipment and medium

    CN114785752A