Smart wearable device and multi-party call method and system
Through a layered architecture that integrates smart wearable devices with cloud collaboration, multilingual multi-party calls between cross-language teams were enabled, solving the problem of fragmented communication in existing technologies and improving collaboration efficiency and communication fluency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GODOX PHOTO EQUIPMENT CO LTD
- Filing Date
- 2026-02-09
- Publication Date
- 2026-06-09
AI Technical Summary
Existing intelligent translation devices cannot achieve real-time, accurate, and broadcast-style distribution across language teams in multi-person collaborative scenarios, resulting in fragmented communication and low efficiency.
It adopts a layered architecture consisting of a smart wearable master device and multiple slave devices. The master device uniformly collects and recognizes voice signals, performs cross-language translation through the cloud, and completes voice synthesis and directional playback by the master device to realize multi-party calls.
It enables efficient collaboration through multilingual, multi-party calls, improving the fluency of communication and collaboration efficiency for cross-language teams. It is compatible with low-power, lightweight hardware and avoids network resource consumption and latency issues.
Smart Images

Figure CN122179689A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of multi-party calling technology, and in particular to a smart wearable device and a multi-party calling method and system. Background Technology
[0002] Simultaneous interpreting is a form of language service that continuously translates spoken content into real time and transmits it to the audience without interrupting the original speaker's expression. Its core goal is to achieve high-precision and low-latency cross-language speech conversion, and it can widely support scenarios with stringent requirements for smooth communication, such as international conferences and business negotiations.
[0003] With the evolution of large-scale artificial intelligence models, smart headphones equipped with real-time translation functions have gradually become available, and can replace human simultaneous interpretation to some extent. However, the current mainstream solutions are still limited to a two-person interaction mode: the user wears headphones, the device only captures the surrounding speech and translates it before pushing it one-way to the wearer, which is equivalent to each person carrying an independent translation terminal.
[0004] This model faces significant bottlenecks on international film and advertising shooting sets. Such teams often consist of directors, cinematographers, lighting technicians, actors, and performance coaches from different countries, with diverse language backgrounds and personnel distributed throughout the location. When everyone can only passively receive translated content and cannot broadcast or share the translation with specific individuals or groups, communication becomes fragmented.
[0005] On-site collaboration is not a linear dialogue, but a dynamic network of multiple roles operating in parallel, interacting frequently, and providing real-time feedback. For example, the director uses Chinese instructions to adjust the Japanese cinematographer's position, while the American lighting technician explains the coordination of lighting and movement to the Korean actors, and the acting director coordinates with the special effects team in French. If the translated content cannot be distributed in real-time, accurately, and broadcast across the audience, the collaboration chain will quickly break down, and overall efficiency will be significantly reduced. Summary of the Invention
[0006] To address at least one of the aforementioned problems, this application provides a smart wearable device and a multi-party calling method and system, aiming to solve the real-time multilingual communication challenges faced by cross-language teams in decentralized, dynamic, and multi-threaded collaboration.
[0007] According to a first aspect of the embodiments of this application, a multi-party calling method is provided, applied to a multi-party calling system, the multi-party calling system including multiple smart wearable devices and a cloud, the multiple smart wearable devices including at least one smart wearable master device and multiple smart wearable slave devices, the smart wearable master device being wirelessly connected to the multiple smart wearable slave devices, and the smart wearable master device also being communicatively connected to the cloud, the method comprising:
[0008] The smart wearable main device collects the corresponding user voice signals; The multiple smart wearable slave devices respectively collect corresponding user voice signals and send the collected user voice signals to the smart wearable master device; The smart wearable master device identifies the target language corresponding to the smart wearable master device and the plurality of smart wearable slave devices respectively; The smart wearable master device sends all the user's voice signals, as well as the target languages corresponding to the smart wearable master device and the multiple smart wearable slave devices, to the cloud. The cloud-based voice translation processing includes translating user voice signals (excluding those from the main smart wearable device itself) into the target language voice signals corresponding to the main smart wearable device, based on the target languages corresponding to the main smart wearable device and the multiple smart wearable slave devices, and translating user voice signals (excluding those from the slave smart wearable devices themselves) into the target language voice signals corresponding to the slave smart wearable devices. The smart wearable master device receives translated target language speech signals sent from the cloud and performs speech synthesis processing. The speech synthesis processing includes synthesizing the translated target language speech signals corresponding to the smart wearable master device to obtain a synthesized translation speech signal; and synthesizing the translated target language speech signals corresponding to the target smart wearable slave device to obtain a target synthesized translation speech signal. The target smart wearable slave device is any one of the plurality of smart wearable slave devices. The smart wearable master device plays the synthesized translated speech signal and sends the target synthesized translated speech signal to the target smart wearable slave device, so that the target smart wearable slave device receives the target synthesized translated speech signal and plays it.
[0009] According to a second aspect of the embodiments of this application, a multi-party calling method is provided, applied to a multi-party calling system, the multi-party calling system including multiple smart wearable devices and a cloud, the multiple smart wearable devices including at least one smart wearable master device and multiple smart wearable slave devices, the smart wearable master device being wirelessly connected to the multiple smart wearable slave devices, and the smart wearable master device also being wirelessly connected to the cloud, the method comprising: The smart wearable main device collects the corresponding user voice signals; The multiple smart wearable slave devices respectively collect corresponding user voice signals and send the collected user voice signals to the smart wearable master device; The smart wearable master device identifies the target language corresponding to the smart wearable master device and the plurality of smart wearable slave devices respectively; The smart wearable master device sends all the user's voice signals, as well as the target languages corresponding to the smart wearable master device and the multiple smart wearable slave devices, to the cloud. The cloud-based system performs speech translation and speech synthesis processing. The speech translation processing includes translating user speech signals (excluding those from the main smart wearable device itself) into speech signals in the target language corresponding to the main smart wearable device, and translating user speech signals (excluding those from the slave smart wearable devices themselves) into speech signals in the target language corresponding to the slave smart wearable devices. The speech synthesis processing includes synthesizing the translated speech signals from the main smart wearable device to obtain a synthesized translated speech signal, and synthesizing the translated speech signals from the target slave smart wearable device to obtain a target synthesized translated speech signal. The target slave smart wearable device can be any one of the multiple slave smart wearable devices. The smart wearable master device receives and plays the synthesized translated speech signal, and the smart wearable master device receives the target synthesized translated speech signal and sends the target synthesized translated speech signal to the target smart wearable device, so that the target smart wearable slave device receives and plays the target synthesized translated speech signal.
[0010] According to a third aspect of the embodiments of this application, a multi-party calling method is provided, applied to a smart wearable master device, wherein the smart wearable master device is wirelessly connected to multiple smart wearable slave devices, and the smart wearable master device is also connected to a cloud, the method comprising: The smart wearable master device collects the corresponding user voice signal and receives the user voice signals collected by the multiple smart wearable slave devices respectively; Identify the target language corresponding to the smart wearable master device and the plurality of smart wearable slave devices respectively; All user voice signals, as well as the target languages corresponding to the smart wearable master device and the multiple smart wearable slave devices, are sent to the cloud so that the cloud can perform voice translation processing. The voice translation processing includes translating each user voice signal except for the smart wearable master device itself into the target language voice signal corresponding to the smart wearable master device, and translating each user voice signal except for the smart wearable slave device itself into the target language voice signal corresponding to the smart wearable slave device. The system receives translated target language speech signals sent from the cloud and performs speech synthesis processing. The speech synthesis processing includes synthesizing the translated target language speech signals corresponding to the smart wearable master device to obtain a synthesized translated speech signal; and synthesizing the translated target language speech signals corresponding to the target smart wearable slave device to obtain a target synthesized translated speech signal. The target smart wearable slave device is any one of the plurality of smart wearable slave devices. Play the synthesized translated speech signal and send the target synthesized translated speech signal to the target smart wearable slave device, so that the target smart wearable slave device receives the target synthesized translated speech signal and plays it.
[0011] According to a fourth aspect of the embodiments of this application, a multi-party call method is provided, applied to a smart wearable master device, wherein the smart wearable master device is wirelessly connected to multiple smart wearable slave devices, and the smart wearable master device is also connected to a cloud, the method comprising: The smart wearable master device collects the corresponding user voice signal and receives the user voice signals collected by the multiple smart wearable slave devices respectively; Identify the target language corresponding to the smart wearable master device and the plurality of smart wearable slave devices respectively; All user voice signals, along with the target languages corresponding to the smart wearable master device and the plurality of smart wearable slave devices, are sent to the cloud. The cloud then performs voice translation and voice synthesis processing. The voice translation processing includes translating all user voice signals (excluding those from the smart wearable master device itself) into the target language voice signal corresponding to the smart wearable master device, and translating all user voice signals (excluding those from the smart wearable slave devices themselves) into the target language voice signal corresponding to the smart wearable slave device. The voice synthesis processing includes synthesizing the translated target language voice signals from the smart wearable master device to obtain a synthesized translated voice signal, and synthesizing the translated target language voice signals from the target smart wearable slave device to obtain a target synthesized translated voice signal. The target smart wearable slave device can be any one of the plurality of smart wearable slave devices. The device receives and plays the synthesized translated speech signal, and receives the target synthesized translated speech signal and sends the target synthesized translated speech signal to the target smart wearable device, so that the target smart wearable device receives and plays the target synthesized translated speech signal.
[0012] According to a fifth aspect of the embodiments of this application, a multi-party call method is provided, applied to smart wearable slave devices, wherein a plurality of smart wearable slave devices are wirelessly connected to a smart wearable master device, and the smart wearable master device is also connected to a cloud, the method comprising: Multiple smart wearable slave devices respectively collect corresponding user voice signals and send the collected user voice signals to the smart wearable master device, so that the smart wearable master device sends its own collected user voice signal, the user voice signals collected by the multiple smart wearable slave devices respectively, and the target language corresponding to the smart wearable master device and the multiple smart wearable slave devices respectively to the cloud, so that the cloud performs voice translation processing. The voice translation processing includes translating each user voice signal other than that of the smart wearable slave device itself into the target language voice signal corresponding to the smart wearable slave device based on the target language corresponding to the smart wearable master device and the multiple smart wearable slave devices respectively. The system receives and plays the target synthesized translated speech signal sent by the smart wearable master device. The target synthesized translated speech signal is obtained by the smart wearable master device receiving translated speech signals of various target languages sent by the cloud and performing speech synthesis processing. The speech synthesis processing includes synthesizing speech signals of various target languages corresponding to the target smart wearable slave device. The target smart wearable slave device is any one of the multiple smart wearable slave devices.
[0013] According to a sixth aspect of the embodiments of this application, a multi-party call method is provided, applied to smart wearable slave devices, wherein a plurality of smart wearable slave devices are wirelessly connected to a smart wearable master device, and the smart wearable master device is also connected to a cloud, the method comprising: Multiple smart wearable slave devices respectively collect corresponding user voice signals and send the collected user voice signals to the smart wearable master device. The smart wearable master device then sends its own collected user voice signal, the user voice signals collected by the multiple smart wearable slave devices, and the target language corresponding to the smart wearable master device and the multiple smart wearable slave devices to the cloud. The cloud then performs speech translation processing and speech synthesis processing. The speech translation processing includes translating each user voice signal (excluding the smart wearable slave device itself) into a speech signal in the target language corresponding to the smart wearable slave device, based on the target language corresponding to the smart wearable master device and the multiple smart wearable slave devices. The speech synthesis processing includes synthesizing the translated target language speech signals corresponding to the target smart wearable slave device to obtain a target synthesized translation speech signal. The target smart wearable slave device can be any one of the multiple smart wearable slave devices. The system receives and plays the target synthesized translated speech signal sent by the smart wearable main device, wherein the target synthesized translated speech signal is obtained by the smart wearable main device from the cloud.
[0014] According to a seventh aspect of the present application, a multi-party calling system is provided for performing the methods described in the first and second aspects of the present application. The multi-party calling system includes multiple smart wearable devices and a cloud. The multiple smart wearable devices include at least one smart wearable master device and multiple smart wearable slave devices. The smart wearable master device is wirelessly connected to the multiple smart wearable slave devices, and the smart wearable master device is also wirelessly connected to the cloud.
[0015] According to an eighth aspect of the present application, a smart wearable main device is provided for use with the method described in the third aspect of the present application, the smart wearable main device comprising: A wireless communication module is provided, through which the smart wearable master device communicates and connects with multiple smart wearable slave devices and the cloud. A voice acquisition module, which is used to acquire user voice signals; Language recognition module, which is used to identify the target language corresponding to the smart wearable main device and each of the smart wearable slave devices respectively; The speech synthesis module is used to perform speech synthesis processing, which includes synthesizing the translated target language speech signals corresponding to the smart wearable master device to obtain a synthesized translation speech signal; and synthesizing the translated target language speech signals corresponding to the target smart wearable slave device to obtain a target synthesized translation speech signal, wherein the target smart wearable slave device is any one of the plurality of smart wearable slave devices; A voice broadcast module is used to play the synthesized translated voice signal.
[0016] According to a ninth aspect of the present application, a smart wearable main device is provided for performing the method described in the fourth aspect of the present application, the smart wearable main device comprising: A wireless communication module is provided, through which the smart wearable master device communicates and connects with multiple smart wearable slave devices and the cloud. A voice acquisition module, which is used to acquire user voice signals; The language recognition module is used to identify the target language corresponding to the smart wearable main device and the multiple smart wearable slave devices respectively; A voice broadcast module is used to play synthesized translated voice signals.
[0017] According to a tenth aspect of the embodiments of this application, a smart wearable slave device is provided for performing the methods described in the fifth and sixth aspects of the embodiments of this application, the smart wearable slave device comprising: A wireless communication unit, through which the smart wearable slave device communicates with the smart wearable master device; A voice acquisition unit, wherein the voice acquisition unit is used to acquire user voice signals; A voice broadcasting unit is used to play the target synthesized translated voice signal.
[0018] According to an eleventh aspect of the embodiments of this application, a multi-party call method is provided, applied to a multi-party call system, the multi-party call system including multiple smart wearable devices and a cloud, the multiple smart wearable devices being respectively communicatively connected to the cloud, the method including: Any one of the plurality of smart wearable devices initiates a multi-party translation call request to establish a multi-party translation call; The multiple smart wearable devices participating in the multi-party translation call each collect the corresponding user voice signals; The multiple smart wearable devices each identify their respective target languages; The multiple smart wearable devices will send the collected user voice signals and corresponding target languages to the cloud. The cloud platform translates multiple user voice signals (excluding the target smart wearable device) into voice signals in the target language corresponding to the target smart wearable device, based on the target language corresponding to each of the multiple smart wearable devices. The target smart wearable device receives translated speech signals from various target languages sent by the cloud, and then synthesizes the translated speech signals from various target languages before playing them.
[0019] According to a twelfth aspect of the present application, a multi-party call method is provided, applied to a multi-party call system, the multi-party call system including multiple smart wearable devices and a cloud, the multiple smart wearable devices being respectively communicatively connected to the cloud, the method comprising: Any one of the plurality of smart wearable devices initiates a multi-party translation call request to establish a multi-party translation call; The multiple smart wearable devices participating in the multi-party translation call each collect the corresponding user voice signals; The multiple smart wearable devices each identify their respective target languages; The multiple smart wearable devices will send the collected user voice signals and corresponding target languages to the cloud. The cloud platform translates multiple user voice signals (excluding the target smart wearable device) into voice signals in the target language corresponding to the target smart wearable device, based on the target language corresponding to each of the multiple smart wearable devices. The cloud platform performs speech synthesis on the translated speech signals of each target language corresponding to the target smart wearable device. The target smart wearable device receives and plays the synthesized speech signals of various target languages corresponding to the target smart wearable device.
[0020] According to a thirteenth aspect of the present application, a multi-party calling system is provided for performing the methods described in the eleventh and twelfth aspects of the present application. The multi-party calling system includes multiple smart wearable devices and a cloud, wherein the multiple smart wearable devices are respectively communicatively connected to the cloud.
[0021] The technical solutions provided by the embodiments of this application have at least the following beneficial effects: The scheme disclosed in this application discloses a multi-party calling system comprising multiple smart wearable devices and a cloud. The multiple smart wearable devices include at least one smart wearable master device and multiple smart wearable slave devices. The smart wearable master device is wirelessly connected to the multiple smart wearable slave devices, and the smart wearable master device is also connected to the cloud. The multi-party calling method involves the smart wearable master device uniformly collecting user voice signals from itself and each smart wearable slave device, identifying the target language of each device, and uniformly interacting with the cloud. The cloud performs targeted cross-language translation based on the target language of each device. Then, the smart wearable master device performs speech synthesis and enables local playback on the smart wearable master device and targeted transmission and playback on the smart wearable slave devices. On the one hand, the centralized processing capability of the smart wearable master device simplifies the communication link of the multi-party calling system, avoiding the network resource consumption, communication latency, and increased device power consumption problems caused by direct interaction between each smart wearable slave device and the cloud, and adapting to the low-power and lightweight hardware characteristics of smart wearable devices; on the other hand, the centralized processing capability of the smart wearable master device simplifies the communication link of the multi-party calling system, avoiding the problems of network resource occupation, communication latency, and increased device power consumption caused by direct interaction between each smart wearable slave device and the cloud, and adapting to the low-power and lightweight hardware characteristics of smart wearable devices; Through precise cloud-based multi-device language translation and targeted synthesis playback on the main device, accurate translation and efficient transmission of each user's voice to the corresponding target language can be achieved in multi-party calls. This allows users speaking different languages to complete barrier-free multi-party voice calls through smart wearable devices, significantly improving the user experience and adaptability of smart wearable devices in multi-party cross-language call scenarios. It can effectively break down communication barriers in multi-language multi-party collaboration, avoid communication fragmentation, ensure the smoothness of the collaboration process, and greatly improve the overall efficiency of multi-language team collaboration on-site. At the same time, it can achieve simultaneous interpretation without interrupting the original speaker's expression, meeting the stringent requirements for communication fluency in such scenarios.
[0022] It should be understood that the above general description and the following detailed description are merely exemplary and do not limit this application. Attached Figure Description
[0023] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the principles of this application.
[0024] Figure 1 This is a schematic diagram of the first architecture of a multi-party call system provided in an embodiment of this application.
[0025] Figure 2 This is a schematic diagram of the second architecture of a multi-party call system provided in one embodiment of this application.
[0026] Figure 3 This is a schematic diagram of the third architecture of a multi-party call system provided in one embodiment of this application.
[0027] Figure 4 This is a schematic block diagram of the first structure of a smart wearable device provided in an embodiment of this application.
[0028] Figure 5 This is a schematic block diagram of the second structure of a smart wearable device provided in an embodiment of this application.
[0029] Figure 6 This is a schematic block diagram of the first structure of a smart wearable main device provided in an embodiment of this application.
[0030] Figure 7 This is a schematic block diagram of the second structure of a smart wearable main device provided in an embodiment of this application.
[0031] Figure 8 This is a schematic block diagram of the third structure of a smart wearable main device provided in an embodiment of this application.
[0032] Figure 9 This is a first flowchart of a multi-party call method performed by a multi-party call system according to an embodiment of this application.
[0033] Figure 10 This is a second flowchart of a multi-party call method performed by a multi-party call system according to an embodiment of this application.
[0034] Figure 11 This is a third flowchart of a multi-party call method executed by a multi-party call system according to an embodiment of this application.
[0035] Figure 12 This is a fourth flowchart of a multi-party call method executed by a multi-party call system according to an embodiment of this application.
[0036] Figure 13 This is a first flowchart of a multi-party call method performed by a smart wearable master device according to an embodiment of this application.
[0037] Figure 14 This is a second flowchart of a multi-party call method performed by a smart wearable master device according to an embodiment of this application.
[0038] Figure 15 This is a fifth flowchart of a multi-party call method executed by a multi-party call system according to an embodiment of this application.
[0039] Figure 16 This is a sixth flowchart of a multi-party call method executed by a multi-party call system according to an embodiment of this application. Detailed Implementation
[0040] To make the objectives, implementation methods, and advantages of this application clearer, exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these exemplary embodiments are provided to make the description of this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art. It should be noted that the brief descriptions of terminology in this application are merely for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.
[0041] In the description of this application, it should be understood that the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first," "second," or "third" may explicitly or implicitly include one or more features.
[0042] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0043] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0044] Current multi-party communication technology faces fundamental bottlenecks in cross-language collaboration scenarios: existing intelligent translation devices generally adopt a "one-to-one" or "one-way reception" mode, with each terminal independently processing voice input and translation, resulting in fragmented communication links and the inability to share information in multi-user scenarios, forming "semantic silos"; traditional conference simultaneous interpretation systems rely on fixed facilities and professional personnel, which cannot adapt to mobile, decentralized, and highly dynamic on-site environments (such as film shooting and emergency rescue), and have high latency and rigid deployment; if AI translation functions are directly deployed on lightweight smart wearable devices, they face hard constraints such as insufficient computing power, increased power consumption, and drastically reduced battery life, making it difficult to support long-term operation; at the same time, communication architectures that rely on public networks have poor stability in complex electromagnetic environments or areas without coverage, and lack low-latency, high-reliability communication networks specifically designed for multilingual collaboration.
[0045] Based on this, this application proposes a smart wearable device and a multi-party calling method and system, aiming to solve the problem of real-time multilingual communication in distributed, dynamic, and multi-threaded collaboration among cross-language teams.
[0046] Reference Figure 1 , Figure 1 This is a schematic diagram of the first architecture of a multi-party call system provided in an embodiment of this application. Figure 1 As shown, the multi-party calling system includes multiple smart wearable devices and a cloud 200. Among the multiple smart wearable devices, there is at least one smart wearable master device 110 and multiple smart wearable slave devices 120. The smart wearable master device 110 is wirelessly connected to the multiple smart wearable slave devices 120, and the smart wearable master device 110 is also connected to the cloud 200.
[0047] This multi-party calling system consists of a core architecture of "smart wearable device cluster + cloud 200". The smart wearable device cluster adopts a "master-slave division of labor" mode: at least one smart wearable master device 110 acts as the system hub. On the one hand, it can establish stable connections with multiple smart wearable slave devices 120 through wireless communication modules in frequency bands such as 2.4GHz / 1.9GHz / 5.8GHz / 1.4GHz, Bluetooth / BLE, WiFi hotspots, and other local wireless communication technologies (slave devices only need to support local communication, and the hardware configuration is more lightweight). It is responsible for aggregating user voice signals collected by itself and all slave devices and identifying the target language of each device. On the other hand, the smart wearable master device 110 can also communicate with the cloud 200 through communication technologies such as WiFi, 3G / 4G / 5G, and other generations of mobile communication standard protocols. As the only node communicating with the cloud 200, the smart wearable master device 110 can upload multiple voice signals and language information of each device to the cloud 200, and receive the translation data returned by the cloud 200 from each device. After completing the speech synthesis, it distributes the translated speech to the corresponding slave devices. Alternatively, the smart wearable master device 110 can also receive the translated voice messages returned by the cloud 200 for each device, and then perform local playback and targeted transmission playback to the slave devices. The smart wearable slave device 120 only needs to perform the basic functions of "collecting voice messages and uploading them to the smart wearable master device" and "receiving and playing the translated voice messages sent by the smart wearable master device". Through this layered architecture of "centralized management of the master device + lightweight collaboration of slave devices + professional execution of voice translation and / or voice synthesis in the cloud", it can not only adapt to the low power consumption and small size hardware characteristics of smart wearable devices, but also achieve efficient collaboration and accurate translation of multilingual multi-party calls.
[0048] Reference Figure 2 , Figure 2 This is a schematic diagram of the second architecture of a multi-party call system provided in an embodiment of this application. Figure 2As shown, the multi-party calling system includes multiple smart wearable devices and a cloud platform 200. The multiple smart wearable devices are configured into at least one smart wearable device group. Each smart wearable device group includes a smart wearable master device 110 and multiple smart wearable slave devices 120. The smart wearable master device 110 in each smart wearable device group is wirelessly connected to the multiple smart wearable slave devices 120. Each smart wearable master device 110 in each smart wearable device group is also wirelessly connected to the cloud platform 200.
[0049] This multi-party calling system consists of multiple smart wearable devices and a cloud platform 200. The smart wearable devices achieve efficient collaboration through "group management": multiple smart wearable devices are divided into at least one smart wearable device group, with each group fixedly configured with one smart wearable master device 110 and multiple smart wearable slave devices 120, forming an independent local communication unit. The smart wearable master device 110 within the group acts as the local control core and can establish stable connections with the smart wearable slave devices 120 within the group via wireless communication modules in frequency bands such as Bluetooth / BLE, 2.4GHz / 1.9GHz / 5.8GHz / 1.4GHz, and WiFi hotspots. The smart wearable master device 110 can maintain the list of smart wearable slave devices and supports dynamic addition and deletion, while also possessing voice pickup, voice playback, and / or voice synthesis functions. The smart wearable slave devices 120 within the group have a lighter hardware configuration, only needing to complete voice acquisition and uploading to the smart wearable master device, and receiving and playing the translated text sent by the smart wearable master device. Furthermore, each smart wearable device group's main smart wearable device 110 independently establishes a communication connection with the cloud 200. This connection can be established via 3G / 4G / 5G public networks or WiFi access the internet. It can call the cloud 200's translation module and / or speech synthesis module (supporting roaming calls), and can also upload multiple voice data streams within the group along with the target language information of each device to the cloud 200. After receiving the translated data from each device after translation processing by the cloud 200, it completes synthesis and targeted distribution. Alternatively, the main smart wearable device 110 can also receive the translated speech returned by the cloud 200 for each device, and then perform local playback and targeted transmission playback from other devices. Ultimately, this enables multi-device parallel operation and accurate intra-group translation in multi-party calls, adapting to the needs of multinational teams and multi-scenario collaboration.
[0050] Reference Figure 3 , Figure 3 This is a schematic diagram of the third architecture of a multi-party call system provided in an embodiment of this application. Figure 3 As shown, the multi-party calling system includes multiple smart wearable devices 100 and a cloud platform 200. Each of the multiple smart wearable devices 100 is communicatively connected to the cloud platform 200.
[0051] This multi-party calling system is based on a core architecture of "100 smart wearable devices + 200 cloud-based devices," with each of the 100 smart wearable devices having the ability to independently communicate with the cloud-based device 200. This means each smart wearable device 100 can communicate separately with the cloud-based device 200 to transmit the collected user voice signals. The cloud-based device 200 then performs voice translation and / or voice synthesis, ensuring that each user only receives and plays translated content matching their language, restoring the authentic rhythm of the conversation and achieving an immersive communication experience of "understanding others without needing to switch." This system can meet the remote communication needs of multinational teams and is also suitable for local temporary team collaboration scenarios, enabling efficient cross-language calls in multiple scenarios.
[0052] The smart wearable device can switch between host and slave modes. When switched to host mode, it becomes a smart wearable master device 110; when switched to slave mode, it becomes a smart wearable slave device 120. This means the smart wearable device has bidirectional switching capabilities between host and slave modes, flexibly changing between either mode depending on the actual usage scenario. In host mode, the device functions as a "local communication base station + cloud interaction hub"; in slave mode, its functions are streamlined and lightweight, requiring only wireless communication with the smart wearable master device, without needing cloud connectivity. This mode switching can be manually operated via device buttons and leverages the master-slave integration feature supported by the Bluetooth protocol stack (meaning the same device can initiate and be connected to), ensuring rapid adaptation to the corresponding role's communication and functional logic after switching. This meets the need for flexible device role adjustments in various scenarios (e.g., the commander in a team switches to host mode, while ordinary members switch to slave mode).
[0053] Among them, smart wearable devices include headsets (or earphones, headphones), smart glasses, bracelets, watches, rings, helmets and other wearable or wearable devices. The specific form of smart wearable devices is not specifically limited in the embodiments of this application.
[0054] Figures 1-3In the multi-party call system shown, the cloud 200 can perform voice translation processing. Specifically, based on the target languages corresponding to the smart wearable master device 110 and multiple smart wearable slave devices 120, it translates the voice signals of all users except the smart wearable master device 110 into the target language voice signal corresponding to the smart wearable master device 110, and translates the voice signals of all users except the smart wearable slave devices 120 into the target language voice signal corresponding to the smart wearable slave devices 120. By leveraging the powerful computing power of the cloud to support parallel multilingual translation, it avoids the problems of insufficient computing power and increased power consumption caused by local deployment of translation functions on smart wearable devices, adapting to their low-power and lightweight hardware characteristics, while ensuring translation accuracy and real-time performance. Through targeted translation logic based on the target language of each device, the master device only receives the adapted language translations of the voices from all other slave devices, and each slave device only receives the adapted language translations of the voices from the master device and other slave devices. This completely breaks down the "semantic silos" of traditional "one-to-one" translation devices, achieving efficient multilingual multi-party translation. Meanwhile, in master-slave mode, the master device centrally uploads data, and the cloud uniformly translates and distributes it to the master device for targeted distribution. This reduces the number of concurrent paths and network resource consumption, saving hardware and software costs. It is suitable for local team collaboration and supports scenarios such as cross-border roaming and multi-device parallel operation, greatly improving the communication fluency and collaboration efficiency of multilingual teams in dynamic and distributed scenarios.
[0055] In some embodiments, Figures 1-3In the multi-party call system shown, the cloud 200 can perform speech translation and speech synthesis processing. The speech translation processing includes translating user speech signals (excluding those from the smart wearable master device 110 itself) into the target language speech signals corresponding to the smart wearable master device 110, and translating user speech signals (excluding those from the smart wearable slave devices 120 themselves) into the target language speech signals corresponding to the smart wearable slave devices 120. The speech synthesis processing includes synthesizing the translated target language speech signals from the smart wearable master device to obtain a synthesized translated speech signal; and synthesizing the translated target language speech signals from the target smart wearable slave device to obtain a target synthesized translated speech signal. The target smart wearable slave device can be any one of the multiple smart wearable slave devices. Leveraging the powerful computing capabilities of the cloud, multilingual targeted translation is centrally completed. This translates the voice signals of all slave devices except the main smart wearable device (110), and translates the voice signals of the main device and other slave devices for each slave device. This avoids the problems of insufficient computing power and increased power consumption caused by local deployment of translation functions on smart wearable devices, adapting to their low-power and lightweight hardware characteristics. It also ensures the accuracy and real-time performance of multilingual translation, supporting roaming and remote call scenarios. Simultaneously, targeted speech synthesis is performed in the cloud, generating synthesized speech that integrates all compatible language translations for the main device and generating exclusive target synthesized speech for each slave device. Combined with the centralized management and distribution architecture of the main device in master-slave mode, this reduces system concurrency and network resource consumption, saving hardware and software costs. It also completely breaks down the "semantic silos" of traditional "one-to-one" translation, enabling efficient multi-role and multi-language translation. This adapts to diverse scenarios such as local team collaboration, cross-border distributed communication, and multi-device parallel processing, significantly improving the fluency of communication and collaboration efficiency of multilingual teams in dynamically distributed scenarios.
[0056] The speech translation processing supports two paths: one is to directly translate speech into the target language, and the other is to first convert speech to text, then translate the text into the target language, and finally convert the target language text back into speech. Specifically, the Cloud 200 can use its AI translation model to translate user speech signals collected by other smart wearable devices (excluding the smart wearable device itself) into individual speech signals in the target language. Alternatively, the Cloud 200 can use its AI translation model to convert user speech signals collected by other smart wearable devices (excluding the smart wearable device itself) into individual text messages, then translate these text messages into individual text messages in the target language, and finally convert the translated target language text messages into individual speech signals in the target language. The Cloud 200, based on its built-in AI translation model, directly converts user speech signals collected by other smart wearable devices into speech signals in the target language. This approach eliminates the need for intermediate text conversion, directly learning the mapping relationship between source and target speech through an end-to-end speech translation model (such as a Transformer-based sequence-to-sequence model). This minimizes latency and is suitable for scenarios with high real-time requirements. The speech-text-speech translation path balances translation accuracy and text traceability. Both paths rely on high-performance AI translation models, capable of handling multilingual scenarios (such as simultaneous translation of English and Russian), and improving translation accuracy and fluency through model optimization. The translation languages can include languages from around the world, dialects, and mixtures of foreign languages and dialects.
[0057] Understandably, the cloud-based 200 can dynamically switch translation paths based on network conditions and speech complexity. For example, when the network is stable, it prioritizes the speech-text-speech path to improve translation accuracy; in scenarios with weak network conditions or high real-time requirements, it automatically switches to the direct speech-target language speech path to reduce latency.
[0058] It's important to note that when multiple smart wearable devices share the same target language (i.e., other smart wearable devices share the same target language as the target smart wearable device), the cloud-based 200's translation of the other smart wearable devices' languages into the target smart wearable device's target language is still considered translation. For example, if smart wearable devices A and B both target Chinese, smart wearable device C targets English, and smart wearable device D targets Russian, then for smart wearable device A, the cloud-based 200 will translate the Chinese speech from smart wearable device B, the English speech from smart wearable device C, and the Russian speech from smart wearable device D into Chinese speech respectively. Here, even though smart wearable device B uses Chinese speech, the cloud-based 200's process of converting smart wearable device B's Chinese speech into Chinese speech is also considered translation. That is, the Chinese speech from smart wearable device B is still input into the AI translation model, and the AI translation model, after recognizing it as Chinese, can translate it into Chinese and output it.
[0059] Reference Figure 4 , Figure 4 This is a schematic block diagram of the first structure of a smart wearable device according to an embodiment of this application. The smart wearable device 100 includes a wireless communication module 101. This module supports multiple network protocols and frequency bands to adapt to the connection needs, power consumption, and bandwidth requirements in different scenarios. Specifically, it may include various standard communication protocols, such as Wi-Fi, Bluetooth, 3G / 4G / 5G, and other generations of mobile communication standard protocols; as well as various proprietary / private communication protocols, such as wireless communication modules using frequency bands such as 5.8GHz, 2.4GHz, 1.9GHz, and 1.4GHz. In some embodiments, the wireless communication module 110 may also employ various generations of mobile communication standard protocols such as 3G / 4G / 5G.
[0060] The wireless communication module 110 can be configured to support a specific single network protocol and frequency band, or it can be a wireless communication module including two or more protocols or frequency bands. For example, the wireless communication module 110 can simultaneously include a Bluetooth module and a 2.4GHz module working together. The Bluetooth module uses the standard Bluetooth protocol for communication, while the 2.4GHz module uses a custom protocol (non-standard) for communication. Through collaborative management and frequency hopping algorithms, it can automatically switch between the two communication modules to achieve optimal communication. As another example, the wireless communication module 110 can be a dual-band module, including communication modules for both 2.4GHz and 1.9GHz frequency bands. Communication between the two frequency bands can be switched through the frequency band management module of the protocol stack or a dynamic frequency selection algorithm to avoid current interference bands and improve call signal quality. The two frequency bands can also work together simultaneously, forming two parallel communication links, one link responsible for uplink data transmission and the other for downlink data transmission, or the transmission of a portion of the data required by the system can be pre-defined by one of the wireless modules.
[0061] Wi-Fi is typically used to provide high-bandwidth, wide-range local area network (LAN) connectivity. In system architecture, it can be used for high-speed data backhaul between smart wearable devices, or for smart wearable devices to access the cloud via Wi-Fi.
[0062] Bluetooth is suitable for short-range, low-power direct connections between devices. For example, various smart wearable devices can form a network via Bluetooth.
[0063] 3G / 4G / 5G (cellular networks) provide wide-area mobile internet access. Modules built into smart wearable devices can directly connect to the internet via cellular networks and then actively establish communication connections with the cloud, enabling remote collaboration without the limitations of fixed networks.
[0064] Proprietary protocols (such as those for 2.4GHz) are crucial for ensuring system exclusivity, stability, and low latency. Proprietary protocols can be used to achieve automatic network configuration between devices (such as between master and slave devices). The advantage of using proprietary protocols (rather than public Wi-Fi) is less interference from general networks and more stable communication. For example, 2.4GHz is a common ISM band; customizing proprietary protocols on this band can optimize time slot allocation (such as uplink / downlink time slots) and reduce contention, thereby providing a predictable low-latency channel for multi-party real-time simultaneous interpretation.
[0065] Furthermore, the wireless communication module 101 can also support digital SIM cards (i.e., eSIM cards) directly embedded in the device chip, eliminating the need for a physical card and enabling flexible switching between networks in different countries. By embedding this chip in a smart wearable device and switching operators by downloading a configuration file, it can easily access local networks, thus making the smart wearable device compatible with networks in various countries.
[0066] The wireless communication module 101 can be fixed to the circuit board inside the smart wearable device 100, i.e., as standard built-in hardware at the factory, or the wireless communication module 101 can be installed in a pluggable / detachable manner on the dedicated wireless module interface of the smart wearable device. This design allows users to replace or upgrade the communication module according to the actual network environment (e.g., the best network type available may differ in different countries or regions), thus enhancing the adaptability and future compatibility of the device.
[0067] The wireless communication module 101 supports diverse network access capabilities, including Wi-Fi, Bluetooth, cellular networks, and proprietary protocols. Combined with fixed or pluggable hardware, it can collectively build a robust, adaptive, and low-latency wireless link foundation. This ensures that in complex multi-party, multilingual call scenarios, voice data can be reliably collected, uploaded to the cloud, and the translation results returned to each participant in real time and clearly.
[0068] When the smart wearable device 100 switches to host mode, that is, when the smart wearable device 100 operates as the smart wearable master device 110, on the one hand, the wireless communication module 101 is responsible for establishing and maintaining stable wireless links with multiple smart wearable slave devices 120. Specifically, the smart wearable master device 110 can network with the smart wearable slave devices 120 through the wireless communication module 101, which can be achieved through various standard communication protocols, such as Wi-Fi, Bluetooth, BLE, and other industry-known communication standard protocols. The smart wearable master device 110 can also network with the smart wearable slave devices 120 through the wireless communication module 101 through various proprietary / private communication protocols, such as wireless communication modules using frequency bands such as 5.8GHz, 2.4GHz, 1.9GHz, and 1.4GHz to achieve networking through custom protocols. On the other hand, the wireless communication module 101 is also responsible for establishing and maintaining stable wireless links with the cloud 200. Based on the wireless communication module 101, the smart wearable main device 110 has flexible public network access capabilities. It can access the Internet through various generations of mobile communication standards such as 3G / 4G / 5G, and then actively establish a communication connection with the cloud 200 to adapt to different regional network environments. Hardware-wise, it supports two adaptation schemes: one is a built-in eSIM card, meeting the needs of cross-border roaming and long-distance calls without requiring an additional SIM card; the other is a SIM card interface that can accept SIM cards of various regional / national standards (including ordinary SIM cards supporting voice, SMS, and data, or data SIM cards designed specifically for Internet applications), allowing flexible switching of network operators. In addition, the smart wearable main device 110 also supports accessing the Internet via WiFi and then actively establishing a communication connection with the cloud 200, further expanding network adaptation scenarios.
[0069] When the smart wearable device 100 switches to slave mode, that is, when the smart wearable device 100 works as a smart wearable slave device 120, the wireless communication module 101 is responsible for establishing and maintaining a stable wireless link between the smart wearable slave device 120 and the smart wearable master device 110.
[0070] Reference Figure 4 The smart wearable device 100 also includes a voice acquisition module 102. The voice acquisition module 102 is used to acquire the voice signals of the user wearing the smart wearable device 100.
[0071] Specifically, the core task of the voice acquisition module 102 is to extract clear and pure voice signals from the wearer's mouth in complex acoustic environments, while minimizing background noise, echoes, and interference from other people's conversations. The voice acquisition module 102 may include a microphone array, typically consisting of 2–4 miniature microphones in a specific geometric layout (such as linear or circular), supporting beamforming to dynamically focus on the wearer's mouth and suppress side and rear noise. The voice acquisition module 102 may also include an active noise cancellation (ANC) chip, which analyzes the ambient noise spectrum in real time and generates inverse sound waves for cancellation, performing particularly well under low-frequency interference such as wind noise and mechanical vibration. The voice acquisition module 102 may also include a voice activity detection (VAD) engine, which incorporates a lightweight AI model to determine whether the user is speaking, avoiding the uploading of meaningless data and saving bandwidth and computing power. The voice acquisition module 102 may also include a near-field voice enhancement algorithm. This algorithm utilizes the difference in spatial distance between voice and noise to strengthen the signal near the mouth, weaken far-field interference, and improve the signal-to-noise ratio (SNR) by more than 15 dB.
[0072] The original voice collected by the voice acquisition module 102 is preprocessed and only valid voice segments (triggered by VAD) are uploaded, rather than a continuous audio stream, which can significantly reduce network load.
[0073] When the voice acquisition module 102 outputs the user's voice signal, it will attach metadata tags, such as: acquisition timestamp, smart wearable device ID, environmental noise level, etc., so that the base station can dynamically adjust the translation parameters (such as enabling a stronger error correction mechanism in a high-noise environment) and process (including translation, synthesis and distribution) each voice signal according to the acquisition timestamp.
[0074] The acquisition timestamp appended by the voice acquisition module 102 is not simply a record of time, but the core time reference for the intelligent wearable main device 110 and / or the cloud 200 to achieve precise synchronous processing of multiple voices. It is precisely relying on this timestamp that the intelligent wearable main device 110 and / or the cloud 200 can uniformly schedule voice streams from different intelligent wearable devices, with different network latencies and in different environments, into a smooth and error-free cross-language conversation. Specifically, at the moment of acquisition of each voice data packet, the voice acquisition module 102 embeds a timestamp accurate to the millisecond level (e.g., 2026-01-20T10:32:15.423Z). This timestamp is based on the high-precision clock source of the device (such as GPS or NTP synchronization), ensuring a high degree of consistency in the time reference of all intelligent wearable devices. The intelligent wearable main device 110 and / or the cloud 200 do not rely on the local system time, but fully trust the timestamp at the acquisition end, thus eliminating voice misalignment caused by clock drift or network jitter of intelligent wearable devices. When voice packets of multiple users arrive at the cloud 200, the cloud 200 does not translate them immediately. Instead, it sorts them by timestamp, arranging all voice packets in ascending order of acquisition time to form a "voice event sequence". After waiting for a set duration (e.g., 200–500 ms) to ensure that all voices collected within the same time window (e.g., A says "Hello" and B immediately says "Good morning") have all arrived. Then, the sorted voice packets are translated in the acquisition order (e.g., sent to multiple AI translation engine instances simultaneously), ensuring that the translation results are strictly consistent with the original speaking order. Exemplarily, if user A speaks at 10:32:15.423 and user B speaks at 10:32:15.610, even if A's signal arrives 300 ms later due to network latency, the cloud 200 will still translate A's voice first and then process B's voice. That is to say, the translation order is determined by "who speaks first" rather than "who arrives first". The translated voices in the target language need to be mixed, and the starting point of the synthesized voice is recalibrated according to the original acquisition timestamp, ensuring that the playback time of the synthesized voice is aligned with the speaking moment of the original speaker. At the same time, an independent voice stream is generated for each user, and time offset compensation is performed according to its original acquisition timestamp, so that in the final multi-language mix output, the "speaking moment" of each language is completely synchronized with the real scenario. For example, when A finishes saying "Hello", the Chinese translation "你好" of B will naturally follow and play exactly 0.2 seconds later, just like a real conversation, rather than being mechanically spliced. Finally, the synthesized voice stream is distributed to each intelligent wearable device according to the network status of each intelligent wearable device and the original acquisition timestamp. For example, a dedicated playback window can be allocated for each intelligent wearable device to ensure that the voice is played at the correct time point. If an intelligent wearable device has a high network latency (such as a weak 5G signal), its corresponding voice packet can be sent in advance so that it can still be aligned with the global timeline when played on the terminal.
[0075] Reference Figure 4 The smart wearable device 100 also includes a voice broadcast module 103. The voice broadcast module 103 is used to play the mixed target language audio signals (i.e., the synthesized translated audio signals). The voice broadcast module 103 is the final output unit in the smart wearable device 100 that achieves a closed-loop cross-language real-time simultaneous interpretation. It accurately, synchronously, and with low latency plays the mixed target language audio stream sent from the cloud 200, enabling the wearer to obtain a natural and interference-free auditory experience in complex multilingual environments.
[0076] The voice broadcast module 103 receives translated, mixed, and timestamped multilingual audio streams transmitted from the wireless communication module 101. The voice broadcast module 103 is not a passive player, but rather an intelligent acoustic actuator that uses timestamp acquisition as its core, language channels as its framework, and user auditory experience as its ultimate goal. Specifically, the voice broadcast module 103 can have a built-in dedicated codec (such as a chip supporting LC3 encoding) to decode compressed digital audio streams into PCM format in real time, with a latency of less than 10ms, far superior to traditional SBC encoding (40–60ms). The voice broadcast module 103 can also receive data via the I2S digital audio bus, convert it into analog signals via a DAC (digital-to-analog converter), and drive MEMS speakers with a miniature power amplifier circuit to achieve high signal-to-noise ratio and low distortion output. The voice broadcast module 103 can also support dual-speaker arrays or bone conduction + air conduction hybrid output to enhance the sense of speech direction and help users distinguish different speakers.
[0077] The playback behavior of the voice broadcast module 103 is entirely driven by the acquisition timestamp added by the voice acquisition module 102, forming an end-to-end time-series closed loop of "acquisition → translation → playback". After receiving the translated and synthesized voice packet, the voice broadcast module 103 does not play it immediately, but calculates its theoretical playback time in the global timeline based on the timestamp. If network latency causes data to arrive late, the voice broadcast module 103 will start the playback buffer in advance to ensure that the voice is output on time at the "should appear time point". All language voice streams are played in alignment with the original speech timestamp. Even if the translation of a certain language is delayed, its playback time is still synchronized with the original speech, avoiding the sense of "translation not keeping up with speech".
[0078] The voice broadcast module 103 supports language-specific playback and speaker recognition, improving the clarity of interaction. Specifically, the voice broadcast module 103 can combine the speaker ID (i.e., smart wearable device ID) metadata sent from the cloud 200 with the voiceprint recognition engine built into the smart wearable device to allocate independent audio channels for different speakers, achieving auditory positioning of "who speaks, who plays".
[0079] In some embodiments, the voice broadcasting module 103 can broadcast the corresponding synthesized translated speech signal according to the user voiceprint corresponding to each smart wearable device. Specifically, after the voice acquisition module 102 acquires the user's voice, it can perform user voiceprint recognition based on the user's voice to determine the user voiceprint corresponding to each smart wearable device (with a unique ID), bind the user voiceprint with the smart wearable device (ID), and upload it to the cloud 200 along with the user's voice signal. After voice translation and / or voice synthesis, the cloud 200 sends the voice package (containing the ID and user voiceprint corresponding to each smart wearable device) to the smart wearable device 100, so that the voice broadcasting module 103 of the smart wearable device 100 can broadcast the corresponding synthesized translated speech signal according to the user voiceprint corresponding to each smart wearable device 100. By dynamically synthesizing and playing the translated speech with the real timbre of the target speaker, an immersive interactive experience of "hearing the voice as if seeing the person" can be achieved.
[0080] For example, User 1 (speaks Chinese) wears smart wearable main device A, and the corresponding user voiceprint is a; User 2 (speaks English) wears smart wearable slave device B, and the corresponding user voiceprint is b; User 3 (speaks Russian) wears smart wearable slave device C, and the corresponding user voiceprint is c. After receiving Chinese voice messages from User 1, English voice messages from User 2, and Russian voice messages from User 3, Cloud 200 can translate User 2's English voice messages and User 3's Russian voice messages into their corresponding Chinese voice messages, perform mixing processing, and then package the mixed and synthesized translated voice signal together with the IDs of smart wearable slave devices B and C, as well as the user's voiceprint, into voice packet 1. Simultaneously, it translates User 1's Chinese voice messages and User 3's Russian voice messages into their corresponding English voice messages, performs mixing processing, and then packages the mixed and synthesized translated voice signal together with the IDs of smart wearable master devices A and C, as well as the user's voiceprint, into voice packet 2. Finally, it translates User 1's Chinese voice messages and User 2's English voice messages into their corresponding Russian voice messages, performs mixing processing, and then packages the mixed and synthesized translated voice signal together with the IDs of smart wearable master devices A and B, as well as the user's voiceprint, into voice packet 3. Voice packets 1, 2, and 3 are then returned to the main smart wearable device A, enabling the voice playback module 103 of the main smart wearable device A to play the Chinese voice translated by user 2 using user voiceprint b, and to play the Chinese voice translated by user 3 using user voiceprint c. Simultaneously, the main smart wearable device A sends voice packet 2 to the secondary smart wearable device B, enabling the voice playback module of the secondary smart wearable device B to play the English voice translated by user 1 using user voiceprint a, and to play the English voice translated by user 3 using user voiceprint c. The main smart wearable device A then sends voice packet 3 to the secondary smart wearable device C, enabling the voice playback module of the secondary smart wearable device C to play the Russian voice translated by user 1 using user voiceprint a, and to play the Russian voice translated by user 2 using user voiceprint b.
[0081] In some embodiments, the voice broadcast module 103 can play synthesized translated speech signals according to a preset electronic voice. That is, the voice broadcast module 103 does not rely on the speaker's identity or voiceprint characteristics, but uniformly calls a preset electronic voice engine to play the target language content with the same timbre, intonation, and rhythm. In high-risk scenarios such as medical, aviation, and emergency command, voice broadcasts must be stable, clear, and free of emotional fluctuations. The preset voice avoids comprehension biases caused by emotions, fatigue, and accents in natural speech. Multiple electronic voices can be preset, and users can switch to their preferred electronic voice.
[0082] Reference Figure 5 , Figure 5This is a schematic block diagram of the second structure of a smart wearable device provided in one embodiment of this application. In addition to a wireless communication module 101, a voice acquisition module 102, and a voice broadcast module 103, the smart wearable device 100 also includes an NFC module 104. The NFC module 104 is used to perform NFC communication and trigger the establishment of a communication connection between the smart wearable device 100 and other smart wearable devices when the smart wearable device 100 is close to them. Specifically, when a user brings the smart wearable device 100 close to other smart wearable devices, the NFC module 104 is automatically activated, establishing a physical layer connection with the NFC antennas of other smart wearable devices using a 13.56MHz high-frequency electromagnetic field. At this time, the smart wearable device 100 sends a preset device identification identifier (such as a unique device ID, supported communication protocol version, and registered user permissions) to other smart wearable devices. The other smart wearable devices then verify the identity and return security key negotiation parameters. Both parties complete two-way authentication within milliseconds, without manual operation or Bluetooth pairing procedures. Other smart wearable devices then initiate a high-speed data channel connection to the smart wearable device 100 via Bluetooth Low Energy (BLE) or Wi-Fi Direct to transmit voice streams, control commands, etc., thereby achieving "instant connection and instant use".
[0083] The NFC module 104 can be built into the smart wearable device 100, or it can be used independently or as a pluggable external module along with other functional modules (such as a wireless communication module) to adapt to older devices that do not natively support NFC, enabling low-cost upgrades. The NFC module 104 is particularly well-suited for screenless or small-screen smart wearable devices. For example, when a user wears an earpiece, bracelet, or smartwatch, NFC becomes the only natural and intuitive interaction point. The smart wearable device does not need to display a pairing list, nor does it require the user to click a "connect" button; it can automatically connect to the base station simply by bringing the smart wearable device close to it.
[0084] By setting up the NFC module 104, the smart wearable device 100 can still be identified, authorized, and activated even when it is not powered on, has no interface, and is not operated, thus completely solving the experience gap of "wanting to connect but not knowing where to start" for smart wearable devices.
[0085] In some embodiments, the NFC module 104 is further configured to trigger the output of a target language setting interaction message after triggering other smart wearable devices to establish a communication connection with the smart wearable device 100, so as to receive the user's setting of the target language of the smart wearable device 100.
[0086] Specifically, after a successful NFC connection, the NFC module 104 can immediately alert the user through vibration feedback, LED flashing, or audio prompts (such as a soft "beep"), and simultaneously present a minimalist language selection interface to the user, driven by the device's built-in micro-interaction engine. For smart wearable devices 100 without a screen or with a very small screen, this interaction is not a visual menu but is achieved through multimodal light interaction. For example, with voice guidance, the smart wearable device plays a pre-recorded prompt: "Please select the language you wish to listen to: Chinese, English, Japanese?" The user can switch candidate languages by tapping the physical button on the side of the wearable device once, and long-press to confirm the selection. Alternatively, the smart wearable device can play a pre-recorded prompt: "Please speak a sentence as your translation target language," thus setting the corresponding language as the target language for the smart wearable device. The corresponding target language information can then be identified through the user's voice signal collected by the voice acquisition module 102. If the smart wearable device 100 has a screen, a language selection menu can pop up on the screen, allowing the user to select and set the target language. Therefore, the target language can be identified, set, and selected on the smart wearable device.
[0087] Continue to refer to Figure 5 The smart wearable device 100 may also include a power supply module 105. The power supply module 105 is electrically connected to each of the functional modules (wireless communication module 101, voice acquisition module 102, voice broadcasting module 103 and NFC module 104) to provide power to each of the functional modules.
[0088] The power supply module 105 can provide precise power to each functional module through an independent voltage regulation circuit. Specifically, the wireless communication module 101 (such as Bluetooth / Wi-Fi) requires instantaneous high current to support electromagnetic field activation during NFC triggering and continuous power supply during high-speed Bluetooth voice stream transmission. The power supply module 105 can have a built-in dynamic boost circuit to ensure no power loss during connection establishment. The microphone array and front-end ADC in the voice acquisition module 102 require extremely low noise power. The power supply module 105 can use an LDO voltage regulator and filter network to suppress electromagnetic interference and ensure the sound pickup signal-to-noise ratio. The NFC module 104 wakes up subsequent modules immediately after triggering pairing. The power supply module 105 can achieve time-sharing power supply through a "wake-up link," that is, first powering the NFC module 104 to complete authentication, and then activating the voice acquisition module 102 and wireless communication module 101 as needed, significantly reducing standby power consumption.
[0089] In some embodiments, continue to refer to Figure 5The smart wearable device 100 may also include a main control module 106. All functional modules (including the wireless communication module 101, voice acquisition module 102, voice broadcasting module 103, NFC module 104, and power supply module 105) are connected to the main control module 106. When the smart wearable device 100 has no screen or a very small screen, the main control module 106 undertakes the crucial task of high-intelligence scheduling under low power consumption. It monitors the status of each module in real time via a hardware bus. Specifically, when the NFC module 104 detects proximity to the cloud 200, the main control module 106 can immediately trigger a "wake-up chain": first, it activates the power supply module 105 to provide instantaneous peak current to the NFC module 104, and then simultaneously starts the voice broadcasting module 103 and the voice acquisition module 102, reserving a channel for possible voice interaction. After the user selects the target language via buttons or touch, the main control module 106 encrypts the configuration command and writes it into the local cache of the voice broadcast module 103, and simultaneously updates the voiceprint matching strategy. This means that whether to enable the user-specific female / male voice or switch to a unified electronic voice is dynamically decided by the main control module based on preset rules. Voltage fluctuations in the power supply module 105, signal strength of the wireless communication module 101, and ambient noise levels of the microphone are continuously sampled by the main control module 106, which dynamically adjusts power consumption allocation through a lightweight AI algorithm. For example, it automatically reduces voice acquisition gain in quiet environments to extend battery life; and prioritizes Bluetooth transmission bandwidth during noisy meetings.
[0090] In some embodiments, the smart wearable device 100 may further include a voiceprint recognition module (not shown in the figure). The voiceprint recognition module is used to identify the user voiceprint corresponding to each smart wearable device 100 based on the user voice signals collected by each smart wearable device 100, and then output a unique voiceprint tag. The voiceprint tag can be used by the voice broadcasting module 103 to broadcast the corresponding synthesized translated speech according to the user voiceprints of other smart wearable devices besides the smart wearable device itself.
[0091] Specifically, the voiceprint recognition module extracts features from the speech signal, focusing on capturing key parameters reflecting individual voiceprint differences, such as fundamental frequency, formant frequency, and spectral envelope. Next, the module extracts voiceprint features using speech signal processing techniques (such as LPC analysis and spectrogram analysis), and then uses AI algorithms (such as deep learning models) to construct a user voiceprint model. For example, a unique voiceprint template can be generated by learning features from the spectrogram using a convolutional neural network (CNN). The voiceprint recognition module then compares the real-time acquired speech features with the pre-stored voiceprint model (e.g., calculating similarity values), and outputs a voiceprint tag (e.g., smart wearable device ID + voiceprint feature value) after confirming the user's identity. This voiceprint tag can be associated with the user's identity information (i.e., smart wearable device ID) for matching the corresponding voiceprint during subsequent voice playback.
[0092] In multi-party calls with multiple speakers, voiceprint tags ensure that each speaker's translated speech retains its unique voiceprint, preventing users from being unable to distinguish the source of the speech after mixing. For example, if user A's English speech is translated into Chinese, it will still be played using user A's original voiceprint.
[0093] In some embodiments, the smart wearable device 100 further includes a language recognition module (not shown in the figure). The language recognition module is used to identify the target language corresponding to each smart wearable device 100. Specifically, after each smart wearable device (such as headphones or a wristband) enables voice acquisition, it can send a short speech segment (usually 500ms-2s) to the language recognition module in real time. The language recognition module can run a lightweight acoustic model to analyze the spectral features, phoneme distribution, and prosodic patterns of the speech, quickly match it with a pre-trained language feature library (such as Chinese, English, Japanese, Spanish, etc.), and then output the result as a target language label (such as "zh-CN"). This identifies the target language corresponding to each smart wearable device 100.
[0094] The language recognition module can identify the target language corresponding to each smart wearable device 100 based on the user's voice signals collected by each smart wearable device 100. Alternatively, the language recognition module can determine the corresponding target language in response to the user's language selection operation.
[0095] In some embodiments, when the smart wearable device 100 switches to slave mode, that is, when the smart wearable device 100 works as a smart wearable slave device 120, the smart wearable slave device 120 includes a wireless communication unit (corresponding to the wireless communication module 101 mentioned above), a voice acquisition unit (corresponding to the voice acquisition module 102 mentioned above), and a voice broadcasting unit (corresponding to the voice broadcasting module 103 mentioned above). The wireless communication unit supports local communication protocols and corresponding frequency bands such as Bluetooth / BLE and 2.4GHz. It does not require public network access capabilities and is only used to establish a stable wireless connection with the smart wearable main device 110 to realize voice data upload and synthesized translation voice signal reception. The voice acquisition unit is equipped with hardware such as MEMS microphones and supports echo cancellation (AEC) and noise reduction (ANC) functions. It can accurately acquire the voice signal of the wearer and upload it to the smart wearable main device 110 without undertaking complex preprocessing tasks. The voice broadcasting unit is equipped with speakers and adapted codecs. It is only used to receive and play the target language voice signal (i.e., synthesized translation voice signal) that has been directionally distributed by the smart wearable main device 110. It does not need to participate in translation or synthesis processing. Through this simplified and targeted unit configuration, the basic functional requirements of the device in multi-party calls are met, and it is in line with its low power consumption and lightweight hardware design positioning, while reducing the overall hardware and software overhead of the system.
[0096] In some embodiments, when the smart wearable device 100 switches to slave mode, that is, when the smart wearable device 100 operates as a smart wearable slave device 120, the smart wearable slave device 120 includes a wireless communication unit (corresponding to the wireless communication module 101 mentioned above), a voice acquisition unit (corresponding to the voice acquisition module 102 mentioned above), and a voice broadcast unit (corresponding to the voice broadcast module 103 mentioned above), and may also include an NFC unit (corresponding to the NFC module 104 mentioned above) and / or a power supply unit (corresponding to the power supply module 105 mentioned above). The NFC unit can support near-field communication in the 13.56MHz band, and can quickly complete identity authentication and join a session with the smart wearable master device 110 via a "tap," adapting to local temporary group communication scenarios. The power supply unit can use a removable rechargeable battery that supports Type-C or magnetic fast charging. It reduces power consumption through precise power supply and time-sharing power-on mechanism, while providing stable power to each unit to ensure the device's long-term operation and battery life. The overall configuration not only meets the basic functional requirements of the device, but also adapts to different scenarios through optional modules, taking into account both lightweight design and usage flexibility.
[0097] Reference Figure 6 , Figure 6 This is a schematic block diagram of the first structure of a smart wearable master device provided in an embodiment of this application. In some embodiments, when the smart wearable device 100 switches to host mode, that is, when the smart wearable device 100 works as a smart wearable master device 110, in addition to the wireless communication module 101, the voice acquisition module 102, and the voice broadcasting module 103, the smart wearable master device 110 also includes a language recognition module 107. The language recognition module 107 is used to identify the target language corresponding to each smart wearable device (including the smart wearable device 100 and multiple smart wearable slave devices 120). Specifically, after receiving the user voice signals from each smart wearable slave device 120 and the smart wearable master device itself, the smart wearable master device (such as headphones or a wristband) can send short voice segments (usually 500ms-2s) to the language recognition module 107 in real time. The language recognition module 107 can run a lightweight acoustic model to analyze the spectral features, phoneme distribution, and prosodic patterns of speech, quickly match a pre-trained language feature library (such as Chinese, English, Japanese, Spanish, etc.), and then output a corresponding target language label for each user's speech signal.
[0098] The language recognition module 107 can identify the target language corresponding to the smart wearable main device 110 and the multiple smart wearable slave devices 120 based on the user's voice signal collected by the smart wearable main device 110 and the user's voice signals collected by the multiple smart wearable slave devices 120 respectively. Alternatively, the language recognition module 107 can determine the target language corresponding to the smart wearable main device 110 in response to the user's language selection operation, and the language recognition module 107 can identify the target language corresponding to the multiple smart wearable slave devices 120 based on the target language information sent by the multiple smart wearable slave devices 120 respectively, wherein the target language information is obtained by the smart wearable slave devices 120 based on the collected user voice signal or in response to the user's language selection operation.
[0099] In other words, when each smart wearable slave device 120 lacks language recognition functionality (e.g., it lacks a language recognition module), the smart wearable master device 110 can identify the target language corresponding to each smart wearable slave device 120 based on the user voice signals collected by each smart wearable slave device 120. However, when each smart wearable slave device 120 possesses language recognition functionality (e.g., it has a language recognition module), each smart wearable slave device 120 can identify the target language information based on the collected user voice signals or in response to the user's language selection operation, and then send this target language information to the smart wearable master device 110. This allows the smart wearable master device to identify the target language corresponding to each smart wearable slave device 120 based on the target language information sent by each smart wearable slave device 120.
[0100] Reference Figure 7 , Figure 7 This is a schematic block diagram of the second structure of a smart wearable main device provided in one embodiment of this application. In some embodiments, when the smart wearable device 100 switches to host mode, that is, when the smart wearable device 100 works as a smart wearable main device 110, in addition to the wireless communication module 101, voice acquisition module 102, voice broadcasting module 103, and language recognition module 107, the smart wearable main device 110 may also include a speech synthesis module 108. That is, the speech synthesis module 108 is arranged at the end of the smart wearable main device 110 to perform speech synthesis processing. At this time, the cloud 200 only needs to perform speech translation processing, and then return the translated target language speech signals to the smart wearable main device 110, where the speech synthesis module 108 of the smart wearable main device 110 performs speech synthesis processing.
[0101] Specifically, the speech synthesis module 108 is used to merge the translated target language speech signals (such as the Chinese speech of user A and user B) corresponding to the smart wearable master device 110 into a single audio stream to obtain the synthesized translated speech signal; and to merge the translated target language speech signals corresponding to the target smart wearable slave device into a single audio stream to obtain the target synthesized translated speech signal; in order to avoid confusion caused by the smart wearable device receiving multi-track signals. The target smart wearable slave device can be any one of multiple smart wearable slave devices. Simultaneously, the speech synthesis module 108 can also solve potential delay problems in digital signal processing (such as differences in signal processing time for different paths), ensuring that each speech signal is synchronized on the time axis and avoiding phase cancellation or comb filter effects when superimposed and overlapping.
[0102] Specifically, the speech synthesis module 108 can adjust the gain (e.g., unify volume levels) and reduce noise for each translated speech signal to ensure the consistency of the input signal. The speech synthesis module 108 can also use sampling compensation (e.g., inserting buffer delays) to align the speech signals from different paths in time, avoiding image shifts or phase problems caused by differences in processing time. The speech synthesis module 108 can send multiple speech signals to a shared sub-mixing track for centralized reverb, equalization, and other effects processing, rather than loading effects for each speech signal individually. The ratio of the original speech (dry signal) to the processed speech (wet signal) can be adjusted using a send knob to balance clarity and spatial awareness; for example, more dry signal can be retained for dialogue scenarios to ensure speech intelligibility. By sharing a mixing processing module (e.g., a sub-mixing track), the system's computing power consumption can be reduced, avoiding the resource waste caused by loading effects plugins for each speech signal individually.
[0103] Reference Figure 8 , Figure 8This is a schematic block diagram of the third structure of a smart wearable master device provided in one embodiment of this application. In some embodiments, when the smart wearable device 100 switches to host mode, that is, when the smart wearable device 100 works as a smart wearable master device 110, in addition to the wireless communication module 101, voice acquisition module 102, voice broadcasting module 103, main control module 106, language recognition module 107, and voice synthesis module 108, the smart wearable master device 110 may also include an NFC module 104 and / or a power supply module 105. The NFC module 104 supports "tap-to-connect" quick networking, facilitating automatic network connection between the smart wearable master device 110 and each smart wearable slave device 120. The power supply module 105 can be equipped with a removable rechargeable battery, supporting Type-C or magnetic fast charging, and provides stable power to each module through low-power management. The overall configuration not only meets the full-function requirements of the master device as a multi-party call control center, but also adapts to different scenarios such as local temporary grouping and long-term operation through optional modules, balancing practicality and flexibility.
[0104] In some embodiments, the smart wearable master device 110 further includes a mode switching module (not shown in the figure). When the language recognition module 107 detects that the target languages corresponding to each smart wearable device (including the smart wearable master device 110 and multiple smart wearable slave devices 120) are different, the mode switching module automatically switches to the multi-party simultaneous interpretation working mode, that is, sends all user voice signals and the target languages corresponding to the smart wearable master device and multiple smart wearable slave devices to the cloud, so that the cloud performs voice translation processing and / or voice synthesis processing to ultimately achieve multi-party simultaneous interpretation. When the language recognition module 107 detects that the target languages corresponding to each smart wearable device (including the smart wearable master device 110 and multiple smart wearable slave devices 120) are the same, the mode switching module automatically switches to the multi-party call working mode, that is, directly controls the voice synthesis module 108 to mix and play the user voice signals collected by other smart wearable devices besides the smart wearable device itself, and transmit them in a targeted manner. For example, if the target languages of the various smart wearable devices are different (such as Chinese, English, and French), the main smart wearable device 110 automatically enters "multi-party simultaneous interpretation mode," allowing each participant to hear the translated speech in real time. If all members of the team use Chinese, the main smart wearable device 110 enters "multi-party call mode," directly mixing the original speech and transmitting it to the corresponding smart wearable device for playback, achieving efficient communication similar to traditional walkie-talkies.
[0105] In this embodiment, the mode switching module automatically switches modes based on language recognition results, requiring no manual intervention. This caters to both cross-language communication needs (simultaneous interpretation mode) and efficient same-language calls (normal call mode), reducing user operating costs. Skipping the translation step in same-language scenarios reduces system computational power consumption, lowers latency, and improves real-time performance.
[0106] Reference Figure 9 , Figure 9 This is a first flowchart of a multi-party call method executed by a multi-party call system according to an embodiment of this application, wherein the multi-party call system may be the aforementioned Figures 1-2 Any of the multi-party calling systems shown. The multi-party calling method includes, but is not limited to, steps S910 to S970.
[0107] Step S910: The smart wearable main device collects the corresponding user voice signal; In step S920, multiple smart wearable slave devices collect corresponding user voice signals and send the collected user voice signals to the smart wearable master device; Step S930: The smart wearable master device identifies the target language corresponding to the smart wearable master device and multiple smart wearable slave devices respectively; In step S940, the smart wearable master device sends all user voice signals and the target languages corresponding to the smart wearable master device and multiple smart wearable slave devices to the cloud. Step S950: The cloud performs voice translation processing. The voice translation processing includes translating the user voice signals (excluding those of the smart wearable main device itself) into the target language voice signals of the smart wearable main device, based on the target languages of the smart wearable main device and multiple smart wearable slave devices respectively, and translating the user voice signals (excluding those of the smart wearable slave devices themselves) into the target language voice signals of the smart wearable slave devices respectively. In step S960, the smart wearable master device receives the translated target language speech signals sent from the cloud and performs speech synthesis processing. The speech synthesis processing includes synthesizing the translated target language speech signals corresponding to the smart wearable master device to obtain a synthesized translated speech signal; and synthesizing the translated target language speech signals corresponding to the target smart wearable slave device to obtain a target synthesized translated speech signal. The target smart wearable slave device can be any one of multiple smart wearable slave devices. In step S970, the smart wearable master device plays the synthesized translated speech signal and sends the target synthesized translated speech signal to the target smart wearable slave device, so that the target smart wearable slave device receives the target synthesized translated speech signal and plays it.
[0108] In this embodiment of the application, the multi-party call method is adapted Figures 1-2 The multi-party communication system architecture shown supports a master-slave mode. Through centralized management by the smart wearable master device 110, professional cloud translation, and master-slave collaborative execution, it achieves real-time and accurate multilingual multi-party communication. The smart wearable master device 110, acting as both a "collection terminal" and a "control center," collects the user's voice signal through its built-in voice acquisition module 102. Simultaneously, multiple smart wearable slave devices 120 collect the corresponding user's voice signal through their own voice acquisition units, eliminating the need for complex preprocessing. After collection, the smart wearable slave devices 120 send the collected voice signal to the smart wearable master device 110 in real time, achieving "multi-terminal collection, one-terminal aggregation," avoiding signal confusion caused by direct transmission from multiple devices. Next, the smart wearable master device 110, through its built-in language recognition module 107, identifies two types of language information: the target language of the smart wearable master device 110 itself and the target languages of each smart wearable slave device 120. The smart wearable master device 110, acting as the sole node communicating with the cloud 200, packages "full voice data" (voice signals collected by itself + voice signals uploaded by all slave devices) and "complete language information" (master device + target language of each slave device) and sends them to the cloud. This "single-point upload" mode avoids excessive system concurrency, network resource waste, and communication latency caused by multiple devices simultaneously connecting to the cloud, significantly saving hardware and software costs. After receiving the data packets uploaded by the smart wearable master device 110, the cloud 200 can invoke a professional voice translation module to perform accurate translation processing based on the "device-language" mapping relationship. Specifically, for the smart wearable master device 110, the voice signals of all smart wearable slave devices 120, excluding the master device itself, can be uniformly translated into the target language voice signal corresponding to the master device 110 (e.g., if the target language of the master device is Chinese, then all English, Japanese, and other slave device voices are translated into Chinese), i.e., synthesized translated voice signals. For each smart wearable slave device, the voice signals of all other devices (master device + other slave devices) except for the slave device itself are translated into the target language voice signal corresponding to the slave device (e.g., if the target language of a slave device is English, then the Chinese voice of the master device and the Japanese voice of other slave devices are translated into English), i.e., the target synthesized translation voice signal. Finally, the smart wearable master device 110 plays the synthesized translation voice signal and sends the target synthesized translation voice signal to the target smart wearable slave device, so that the target smart wearable slave device receives and plays the target synthesized translation voice signal.
[0109] This embodiment of the application achieves an optimal balance between computing power and real-time performance by centrally processing speech translation in the cloud (200) and locally performing audio mixing and synthesis on the smart wearable main device (110). The cloud (200), relying on a powerful AI model, can accurately translate multiple different languages into the specified language for each target device, ensuring translation quality and broad language coverage. This meets the needs of multinational teams and multilingual collaboration scenarios, avoiding translation errors or language limitations caused by insufficient computing power on local devices. The smart wearable main device (110) focuses on low-latency audio synthesis and wireless scheduling. The smart wearable main device (110) performs audio mixing and synthesis locally, without waiting for the cloud (200) to return the synthesis results, reducing data transmission round-trip time. Simultaneously, the smart wearable main device focuses on wireless scheduling, quickly distributing the synthesized translated speech, reducing end-to-end latency. The smart wearable slave device (120) only undertakes basic functions of speech acquisition and playback, without participating in computationally intensive tasks such as translation or synthesis, significantly reducing power consumption.
[0110] In some embodiments, the speech translation processing performed by the cloud 200 includes: based on the target languages corresponding to the smart wearable master device 110 and multiple smart wearable slave devices 120, the cloud 200 converts each user speech signal (excluding the smart wearable master device 110 itself) into text information, translates each text information into text information in the target language corresponding to the smart wearable master device 110, and then converts the translated text information in the target language corresponding to the smart wearable master device 110 into speech signals in the target language corresponding to the smart wearable master device 110; and converting each user speech signal (excluding the smart wearable slave device 120 itself) into text information, translating each text information into text information in the target language corresponding to the smart wearable slave device 120, and then converting the translated text information in the target language corresponding to the smart wearable slave device 120 into speech signals in the target language corresponding to the smart wearable slave device 120.
[0111] Reference Figure 10 , Figure 10 This is a second flowchart of a multi-party call method executed by a multi-party call system according to an embodiment of this application, wherein the multi-party call system may be the aforementioned Figures 1-2 Any of the multi-party calling systems shown. The multi-party calling method includes, but is not limited to, steps S1010 to S1060.
[0112] Step S1010: The smart wearable main device collects the corresponding user voice signal; Step S1020: Multiple smart wearable slave devices collect corresponding user voice signals and send the collected user voice signals to the smart wearable master device; Step S1030: The smart wearable master device identifies the target language corresponding to the smart wearable master device and multiple smart wearable slave devices respectively; Step S1040: The smart wearable master device sends all user voice signals and the target languages corresponding to the smart wearable master device and multiple smart wearable slave devices to the cloud. Step S1050: The cloud performs speech translation processing and speech synthesis processing. Speech translation processing includes translating user speech signals (excluding those from the main smart wearable device itself) into speech signals in the target language corresponding to the main smart wearable device, and translating user speech signals (excluding those from the slave smart wearable devices themselves) into speech signals in the target language corresponding to the slave smart wearable devices. Speech synthesis processing includes synthesizing the translated target language speech signals from the main smart wearable device to obtain a synthesized translation speech signal, and synthesizing the translated target language speech signals from the target slave smart wearable device to obtain a target synthesized translation speech signal. The target slave smart wearable device can be any one of the multiple slave smart wearable devices. In step S1060, the smart wearable master device receives and plays the synthesized translated speech signal, and the smart wearable master device receives the target synthesized translated speech signal and sends the target synthesized translated speech signal to the target smart wearable device, so that the target smart wearable slave device receives and plays the target synthesized translated speech signal.
[0113] In this embodiment of the application, the multi-party call method is adapted Figures 1-2The multi-party calling system architecture shown supports a master-slave mode. Through centralized management by the smart wearable master device and integrated cloud processing, it enables efficient multi-language, multi-party communication. Specifically, the smart wearable master device 110 first collects the voice signals of its own users. Multiple smart wearable slave devices 120 then collect the voice signals of their respective users and upload them to the smart wearable master device 110, forming a complete voice dataset. The smart wearable master device 110 identifies the target language of itself and all smart wearable slave devices 120, then packages the complete voice signals with the complete language information and sends them uniformly to the cloud 200 to avoid resource waste caused by multiple devices connecting to the cloud. After receiving the data, the cloud 200 performs speech translation and speech synthesis processing. The translation process involves translating user voice signals (excluding those from the main smart wearable device 110) into their respective target language versions based on the target languages of the main smart wearable device 110 and multiple slave smart wearable devices 120. Similarly, it translates user voice signals from each slave device (excluding those from the main smart wearable device 110) into their target language versions. The synthesis process involves combining the translated target language voice signals from the main smart wearable device 110 to obtain a synthesized translated voice signal, and combining the translated target language voice signals from the target slave smart wearable device to obtain a target synthesized translated voice signal. The target slave smart wearable device can be any one of the multiple slave smart wearable devices 120. Finally, the main smart wearable device 110 receives and plays the synthesized translated voice signal, and simultaneously receives the target synthesized translated voice signal and sends it to the target slave smart wearable device for playback.
[0114] In this embodiment, the smart wearable master device 110 centrally collects voice data, identifies the language, and uploads it to the cloud 200. This avoids problems such as excessive system concurrency and wasted network resources caused by multiple devices directly connecting to the cloud, significantly saving hardware and software costs. Relying on the powerful computing power of the cloud to complete multilingual targeted translation and speech synthesis ensures both translation accuracy and language coverage, while also supporting roaming and remote call scenarios. Simultaneously, the master device acts as a control center for targeted distribution of translations, and with the help of local area networking technology, it ensures the stability and real-time performance of voice transmission. The slave devices only perform lightweight functions of voice acquisition and translation playback, significantly reducing power consumption and extending battery life.
[0115] In some embodiments, refer to Figure 11 , Figure 11 This is a third flowchart of a multi-party call method executed by a multi-party call system according to an embodiment of this application, wherein the multi-party call system may be the aforementioned Figures 1-2Any of the multi-party calling systems shown. The multi-party calling method includes, but is not limited to, steps S1110 to S1190.
[0116] Step S1110: The smart wearable main device collects the corresponding user voice signal; Step S1120: Multiple smart wearable slave devices collect corresponding user voice signals and send the collected user voice signals to the smart wearable master device; Step S1130: The smart wearable master device identifies the target language corresponding to the smart wearable master device and multiple smart wearable slave devices respectively; Step S1140: If the smart wearable master device recognizes that the target languages corresponding to the smart wearable master device and multiple smart wearable slave devices are different, the smart wearable master device executes the multi-party simultaneous interpretation working mode to send all user voice signals and the target languages corresponding to the smart wearable master device and multiple smart wearable slave devices to the cloud. Step S1150: The cloud performs voice translation processing. The voice translation processing includes translating the user voice signals (excluding those of the smart wearable main device) into the target language voice signals of the smart wearable main device, based on the target languages of the smart wearable main device and multiple smart wearable slave devices, and translating the user voice signals (excluding those of the smart wearable slave devices) into the target language voice signals of the smart wearable slave devices. In step S1160, the smart wearable master device receives the translated target language speech signals sent from the cloud and performs speech synthesis processing. The speech synthesis processing includes synthesizing the translated target language speech signals corresponding to the smart wearable master device to obtain a synthesized translated speech signal; and synthesizing the translated target language speech signals corresponding to the target smart wearable slave device to obtain a target synthesized translated speech signal. The target smart wearable slave device can be any one of multiple smart wearable slave devices. In step S1170, the smart wearable master device plays the synthesized translated speech signal and sends the target synthesized translated speech signal to the target smart wearable slave device, so that the target smart wearable slave device receives the target synthesized translated speech signal and plays it. Step S1180: If the smart wearable master device recognizes that the target language of the smart wearable master device and multiple smart wearable slave devices is the same, the smart wearable master device switches to the multi-party call working mode to synthesize the user voice signals other than the smart wearable master device into a synthesized voice signal based on the unique device identifiers of the smart wearable master device and multiple smart wearable slave devices, and synthesizes the voice signals other than the target smart wearable slave device into a target synthesized voice signal. In step S1190, the smart wearable master device plays the synthesized speech signal and sends the target synthesized speech signal to the target smart wearable slave device, so that the target smart wearable slave device receives the target synthesized speech signal and plays it.
[0117] In this embodiment, a dual-mode automatic switching mechanism based on language recognition enables multiple optimizations in resource utilization efficiency, scenario adaptability, and device battery life. Specifically, when the smart wearable master device detects differences in the target language between the master device and multiple slave devices, it automatically switches to a multi-party simultaneous interpretation mode. The master device centrally uploads all voice and language information to the cloud for targeted translation, and then synthesizes and distributes the translation locally. This approach leverages the powerful computing capabilities of the cloud to ensure the accuracy and coverage of multilingual translation while avoiding excessive concurrent paths and resource waste caused by multiple devices directly connecting to the cloud. When the main smart wearable device recognizes that the target language of the main device and multiple slave devices is consistent, it automatically switches to multi-party call mode. Without calling cloud translation resources, the main device directly completes speech synthesis and distribution based on the device's unique identifier. This significantly reduces data transmission links and cloud computing power consumption, and significantly reduces communication latency and device power consumption, making it suitable for efficient collaboration scenarios of teams speaking the same language. This on-demand switching mode design requires no manual intervention from the user. It covers diverse scenarios such as cross-border multilingual collaboration and local same-language communication. Furthermore, the centralized management architecture of the main device optimizes the power consumption performance of smart wearable devices, extends battery life, and ensures the real-time and smoothness of voice interaction in both modes, thus improving the overall user experience.
[0118] In some embodiments, refer to Figure 12 , Figure 12 This is a fourth flowchart of a multi-party call method executed by a multi-party call system according to an embodiment of this application, wherein the multi-party call system may be the aforementioned Figures 1-2 Any of the multi-party calling systems shown. The multi-party calling method includes, but is not limited to, steps S1210 to S1280.
[0119] Step S1210: The smart wearable main device collects the corresponding user voice signal; Step S1220: Multiple smart wearable slave devices collect corresponding user voice signals and send the collected user voice signals to the smart wearable master device; Step S1230: The smart wearable master device identifies the target language corresponding to the smart wearable master device and multiple smart wearable slave devices respectively; Step S1240: If the smart wearable master device recognizes that the target languages corresponding to the smart wearable master device and multiple smart wearable slave devices are different, the smart wearable master device executes the multi-party simultaneous interpretation working mode to send all user voice signals and the target languages corresponding to the smart wearable master device and multiple smart wearable slave devices to the cloud. Step S1250: The cloud performs speech translation processing and speech synthesis processing. Speech translation processing includes translating user speech signals (excluding those from the main smart wearable device itself) into speech signals in the target language corresponding to the main smart wearable device, and translating user speech signals (excluding those from the slave smart wearable devices themselves) into speech signals in the target language corresponding to the slave smart wearable devices. Speech synthesis processing includes synthesizing the translated target language speech signals from the main smart wearable device to obtain a synthesized translation speech signal, and synthesizing the translated target language speech signals from the target slave smart wearable device to obtain a target synthesized translation speech signal. The target slave smart wearable device can be any one of the multiple slave smart wearable devices. In step S1260, the smart wearable master device receives and plays the synthesized translated speech signal, and the smart wearable master device receives the target synthesized translated speech signal and sends the target synthesized translated speech signal to the target smart wearable device, so that the target smart wearable device receives and plays the target synthesized translated speech signal from the device. Step S1270: If the smart wearable master device recognizes that the target language of the smart wearable master device and multiple smart wearable slave devices is the same, the smart wearable master device switches to the multi-party call working mode to synthesize the user voice signals other than the smart wearable master device into a synthesized voice signal based on the unique device identifiers of the smart wearable master device and multiple smart wearable slave devices, and synthesizes the voice signals other than the target smart wearable slave device into a target synthesized voice signal. In step S1280, the smart wearable master device plays the synthesized speech signal and sends the target synthesized speech signal to the target smart wearable slave device, so that the target smart wearable slave device receives the target synthesized speech signal and plays it.
[0120] In this embodiment, the language recognition-driven dual-mode intelligent switching design achieves multiple beneficial effects, including efficient resource utilization, flexible scenario adaptation, and enhanced user experience. Specifically, when the smart wearable master device detects a difference in target language between itself and multiple slave devices, it automatically initiates a multi-party simultaneous interpretation mode. The master device centrally uploads all voice and language information to the cloud, and leverages the powerful computing power of the cloud to simultaneously complete targeted translation and speech synthesis. The master device then receives and distributes the translated text, ensuring both the accuracy and coverage of multilingual translation while avoiding excessive concurrent paths and resource waste caused by direct cloud connections from multiple devices. This approach is suitable for cross-language communication scenarios such as multinational team collaboration. When the smart wearable master device detects that the target language is the same as that of multiple slave devices, it automatically switches to a multi-party call mode. Without accessing cloud resources, the master device directly completes speech synthesis and distribution based on the device's unique identifier. This significantly reduces data transmission links and computing power consumption, substantially lowers communication latency and device power consumption, and is suitable for daily collaboration scenarios among local teams speaking the same language. This automatic switching mechanism, which requires no manual user intervention, not only covers diverse communication needs, but also, combined with the anti-interference design of the main device's local area network and low-power hardware configuration, further extends the device's battery life and improves the overall smoothness and usability of the interaction.
[0121] Reference Figure 13 , Figure 13 This is a first flowchart of a multi-party call method executed by a smart wearable main device according to an embodiment of this application, wherein the smart wearable main device may be the aforementioned Figures 6-7 Any one of the smart wearable main devices shown. The multi-party calling method includes, but is not limited to, steps S1310 to S1350.
[0122] Step S1310: The smart wearable master device collects the corresponding user voice signal and receives user voice signals collected by multiple smart wearable slave devices respectively; Step S1320: Identify the target language corresponding to the smart wearable master device and multiple smart wearable slave devices respectively; Step S1330: Send all user voice signals and the target languages corresponding to the smart wearable master device and multiple smart wearable slave devices to the cloud so that the cloud can perform voice translation processing. The voice translation processing includes translating each user voice signal except for the smart wearable master device itself into the target language voice signal corresponding to the smart wearable master device, and translating each user voice signal except for the smart wearable slave device itself into the target language voice signal corresponding to the smart wearable slave device. Step S1340: Receive the translated target language speech signals sent from the cloud and perform speech synthesis processing. The speech synthesis processing includes synthesizing the translated target language speech signals corresponding to the smart wearable master device to obtain a synthesized translated speech signal; and synthesizing the translated target language speech signals corresponding to the target smart wearable slave device to obtain a target synthesized translated speech signal. The target smart wearable slave device can be any one of multiple smart wearable slave devices. Step S1350: Play the synthesized translated speech signal and send the target synthesized translated speech signal to the target smart wearable slave device so that the target smart wearable slave device receives and plays the target synthesized translated speech signal.
[0123] In this embodiment, a smart wearable master device undertakes the core functions of data aggregation, language recognition, and speech synthesis and distribution. Combined with cloud-based professional translation capabilities, it enables efficient collaboration in multilingual, multi-party calls. Specifically, the master device centrally collects its own and each slave device's voice signals and uniformly identifies the target language, avoiding information chaos caused by multiple devices operating independently. Simultaneously, by uploading data to the cloud from a single point for targeted translation, it significantly reduces the number of concurrent paths in the system, saving network resources and hardware / software overhead. The cloud relies on powerful computing capabilities to ensure the accuracy of multilingual translation, while the master device handles local speech synthesis and targeted distribution. This reduces the pressure on cloud data backhaul and leverages the low latency of local area networking to improve the real-time performance of translated text playback. Slave devices only need to complete lightweight tasks of voice acquisition and translated text playback. Combined with low-power hardware configuration, this significantly extends battery life, adapting to long-term operation scenarios. Furthermore, the master device, acting as a control center, can dynamically add or remove slave devices, flexibly adapting to diverse scenarios such as local team collaboration and cross-border roaming communication, completely breaking down the "semantic silos" of cross-language collaboration and improving the fluency and efficiency of multi-role communication.
[0124] Reference Figure 14 , Figure 14 This is a second flowchart of a multi-party call method executed by a smart wearable main device according to an embodiment of this application, wherein the smart wearable main device may be the aforementioned Figures 6-7 Any one of the smart wearable main devices shown. The multi-party calling method includes, but is not limited to, steps S1410 to S1440.
[0125] Step S1410: The smart wearable master device collects the corresponding user voice signal and receives user voice signals collected by multiple smart wearable slave devices respectively. Step S1420: Identify the target language corresponding to the smart wearable master device and multiple smart wearable slave devices respectively; Step S1430: All user voice signals and the target languages corresponding to the smart wearable master device and multiple smart wearable slave devices are sent to the cloud, so that the cloud can perform voice translation processing and voice synthesis processing. The voice translation processing includes translating each user voice signal (excluding the smart wearable master device itself) into the target language voice signal corresponding to the smart wearable master device, and translating each user voice signal (excluding the smart wearable slave devices themselves) into the target language voice signal corresponding to the smart wearable slave devices. The voice synthesis processing includes synthesizing the translated target language voice signals corresponding to the smart wearable master device to obtain a synthesized translated voice signal, and synthesizing the translated target language voice signals corresponding to the target smart wearable slave device to obtain a target synthesized translated voice signal. The target smart wearable slave device can be any one of the multiple smart wearable slave devices. Step S1440: Receive and play the synthesized translated speech signal, and receive the target synthesized translated speech signal and send the target synthesized translated speech signal to the target smart wearable device, so that the target smart wearable device receives and plays the target synthesized translated speech signal from the device.
[0126] In this embodiment, the smart wearable master device coordinates the entire process of voice acquisition, language recognition, data upload, and translation distribution, and synchronously completes translation and synthesis tasks via the cloud. The master device centrally aggregates its own and all slave devices' voice signals and uniformly identifies the target language. Single-point upload to the cloud avoids the problems of excessive concurrency and resource waste caused by direct cloud connections from multiple devices, significantly reducing hardware and software costs. The cloud, with its powerful computing capabilities, synchronously completes multilingual targeted translation and speech synthesis, ensuring both translation accuracy and language coverage, while also supporting roaming and remote calling scenarios, adapting to the needs of cross-border team collaboration. The master device only needs to receive the synthesized translation from the cloud and distribute it to the corresponding slave devices, without bearing the computing power consumption of local synthesis, effectively reducing device power consumption and extending battery life. Simultaneously, the slave devices only need to perform lightweight tasks of voice acquisition and translation playback. The overall architecture balances efficient cross-language communication with the hardware characteristics of smart wearable devices, achieving a smooth and stable multi-party call experience in multiple scenarios.
[0127] Reference Figure 15 , Figure 15 This is a fifth flowchart of a multi-party call method executed by a multi-party call system according to an embodiment of this application, wherein the multi-party call system may be the aforementioned Figure 3 Any of the multi-party calling systems shown. The multi-party calling method includes, but is not limited to, steps S1510 to S1560.
[0128] Step S1510: Any one of the multiple smart wearable devices initiates a multi-party translation call request to establish a multi-party translation call; Step S1520: Multiple smart wearable devices participating in the multi-party translation call collect the corresponding user voice signals respectively; Step S1530: Multiple smart wearable devices identify their respective target languages. In step S1540, multiple smart wearable devices send the collected user voice signals and corresponding target languages to the cloud. Step S1550: The cloud translates multiple user voice signals (excluding the target smart wearable device) into voice signals in the target language corresponding to the target smart wearable device, based on the target language corresponding to each of the multiple smart wearable devices. The target smart wearable device is any one of the multiple smart wearable devices. In step S1560, the target smart wearable device receives the translated voice signals of various target languages sent from the cloud, and then synthesizes the translated voice signals of various target languages before playing them.
[0129] In this embodiment, the multi-party call method is based on a non-master-slave architecture, where any smart wearable device initiates the conversation, and each device interacts independently with the cloud. This method supports any smart wearable device as the initiator to establish a multi-party translation call, and offers various convenient joining methods such as conversation ID and password, conversation group QR code, conversation push invitation, and NFC tap-to-join. It adapts to the remote collaboration needs of multinational teams and also meets the rapid communication needs of local temporary groups, significantly improving the system's flexibility and ease of use. Each device independently collects voice, identifies the target language, and uploads it to the cloud. The cloud then performs unified targeted translation, eliminating the need for centralized management by a master device and avoiding system paralysis due to single-device failure, thus improving communication stability. The target smart wearable device receives the translated voice signal from the cloud and synthesizes and plays it locally. This leverages the powerful computing power of the cloud to ensure the accuracy and coverage of multilingual translation, while distributed interaction reduces intermediate data transmission links, balancing translation quality and real-time communication. Furthermore, each device calls cloud resources on demand, effectively controlling power consumption and adapting to the lightweight and long-lasting hardware characteristics of smart wearable devices.
[0130] Reference Figure 16 , Figure 16 This is a sixth flowchart of a multi-party call method executed by a multi-party call system according to an embodiment of this application, wherein the multi-party call system may be the aforementioned Figure 3 Any of the multi-party calling systems shown. The multi-party calling method includes, but is not limited to, steps S1610 to S1670.
[0131] Step S1610: Any one of the multiple smart wearable devices initiates a multi-party translation call request to establish a multi-party translation call; Step S1620: Multiple smart wearable devices participating in the multi-party translation call collect the corresponding user voice signals respectively; Step S1630: Multiple smart wearable devices identify their respective target languages; In step S1640, multiple smart wearable devices send the collected user voice signals and corresponding target languages to the cloud. Step S1650: The cloud translates multiple user voice signals (excluding the target smart wearable device) into voice signals in the target language corresponding to the target smart wearable device, based on the target language corresponding to each of the multiple smart wearable devices. The target smart wearable device is any one of the multiple smart wearable devices. Step S1660: The cloud performs speech synthesis on the translated speech signals of each target language corresponding to the target smart wearable device. In step S1670, the target smart wearable device receives and plays the speech signals of each target language after speech synthesis corresponding to the target smart wearable device.
[0132] In this embodiment, the multi-party calling method is based on a non-master-slave architecture, where any smart wearable device initiates the conversation, and each device interacts independently with the cloud. This method supports any device initiating a multi-party translation call, and offers various convenient joining methods such as conversation ID and password, conversation group QR code, conversation push invitation, and NFC tap-to-join. It adapts to the remote collaboration needs of multinational teams and also meets the rapid communication needs of local temporary groups, significantly improving the system's flexibility and ease of use. Each device independently collects voice, identifies the target language, and uploads it to the cloud. The cloud centrally performs targeted translation and speech synthesis, eliminating the need for centralized management by a master device and avoiding the risk of system paralysis due to single-device failure. Simultaneously, the powerful computing power of the cloud ensures the accuracy of multilingual translation and the smoothness of synthesis, supporting roaming calls across different locations. The target smart wearable device directly receives and plays the synthesized voice signal from the cloud, which reduces intermediate data transmission links, ensures real-time communication, and allows each device to access cloud resources on demand. Combined with low-power hardware configuration, it effectively controls power consumption, extends battery life, and adapts to the lightweight and long-term operation characteristics of smart wearable devices, achieving an efficient, stable, and low-cost multilingual multi-party call experience.
[0133] In this process, any one of the multiple smart wearable devices initiates a multi-party translation call request to establish a multi-party translation call. This includes: any one of the multiple smart wearable devices initiates a multi-party translation call request and provides a session joining method, which may include session ID and password, session group QR code, session push invitation, or NFC module near-field touch to join; other smart wearable devices join the session based on the session joining method to establish a multi-party translation call.
[0134] Specifically, multiple smart wearable devices can initiate multi-party translation call requests without requiring a designated host. Each device can request a dedicated multi-party translation call session from the cloud, which assigns a unique identifier to ensure the session's independence and security. Joining methods include: the initiating device generating a session ID and password, which other users can manually enter to join. This method is suitable for scenarios where face-to-face interaction is inconvenient, such as remote cross-border collaboration, requiring only the exchange of digital information to form a team. The initiating device can also generate a corresponding session group QR code, which other devices can scan to automatically identify and join the session. This method is intuitive and convenient, suitable for small teams quickly forming call groups offline. Alternatively, the initiating device can select a target contact from the companion app's friend list and send a push notification for the session, which the recipient can join by clicking or responding with voice. This method is suitable for familiar scenarios with clearly defined communication partners, allowing for precise invitations to specific individuals. The initiating device can also automatically pair with other devices by touching them in close proximity using NFC, allowing other devices to join the session with a single click, without manual input or scanning. This method is well-suited for emergency communication scenarios involving temporary local team formation, enabling near-instant team building. Regardless of the joining method, all devices will enter the same multi-party translation call session after successful joining. The cloud will perform targeted translation and / or synthesis of multiple voice messages based on the target language settings of each device. The entire process does not rely on fixed base stations or main equipment, improving the system's flexibility and ease of use while adapting to the communication needs of diverse scenarios such as remote collaboration and local emergency response.
Claims
1. A multi-party call method, applied to a multi-party call system, characterized in that, The multi-party calling system includes multiple smart wearable devices and a cloud platform. The multiple smart wearable devices include at least one smart wearable master device and multiple smart wearable slave devices. The smart wearable master device is wirelessly connected to the multiple smart wearable slave devices, and the smart wearable master device is also connected to the cloud platform. The method includes: The smart wearable main device collects the corresponding user voice signals; The multiple smart wearable slave devices respectively collect corresponding user voice signals and send the collected user voice signals to the smart wearable master device; The smart wearable master device identifies the target language corresponding to the smart wearable master device and the plurality of smart wearable slave devices respectively; The smart wearable master device sends all the user's voice signals, as well as the target languages corresponding to the smart wearable master device and the multiple smart wearable slave devices, to the cloud. The cloud-based voice translation processing includes translating user voice signals (excluding those from the main smart wearable device itself) into the target language voice signals corresponding to the main smart wearable device, based on the target languages corresponding to the main smart wearable device and the multiple smart wearable slave devices, and translating user voice signals (excluding those from the slave smart wearable devices themselves) into the target language voice signals corresponding to the slave smart wearable devices. The smart wearable master device receives translated target language speech signals sent from the cloud and performs speech synthesis processing. The speech synthesis processing includes synthesizing the translated target language speech signals corresponding to the smart wearable master device to obtain a synthesized translation speech signal; and synthesizing the translated target language speech signals corresponding to the target smart wearable slave device to obtain a target synthesized translation speech signal. The target smart wearable slave device is any one of the plurality of smart wearable slave devices. The smart wearable master device plays the synthesized translated speech signal and sends the target synthesized translated speech signal to the target smart wearable slave device, so that the target smart wearable slave device receives the target synthesized translated speech signal and plays it.
2. A multi-party call method, applied to a multi-party call system, characterized in that, The multi-party calling system includes multiple smart wearable devices and a cloud platform. The multiple smart wearable devices include at least one smart wearable master device and multiple smart wearable slave devices. The smart wearable master device is wirelessly connected to the multiple smart wearable slave devices, and the smart wearable master device is also connected to the cloud platform. The method includes: The smart wearable main device collects the corresponding user voice signals; The multiple smart wearable slave devices respectively collect corresponding user voice signals and send the collected user voice signals to the smart wearable master device; The smart wearable master device identifies the target language corresponding to the smart wearable master device and the plurality of smart wearable slave devices respectively; The smart wearable master device sends all the user's voice signals, as well as the target languages corresponding to the smart wearable master device and the multiple smart wearable slave devices, to the cloud. The cloud-based system performs speech translation and speech synthesis processing. The speech translation processing includes translating user speech signals (excluding those from the main smart wearable device itself) into speech signals in the target language corresponding to the main smart wearable device, and translating user speech signals (excluding those from the slave smart wearable devices themselves) into speech signals in the target language corresponding to the slave smart wearable devices. The speech synthesis processing includes synthesizing the translated speech signals from the main smart wearable device to obtain a synthesized translated speech signal, and synthesizing the translated speech signals from the target slave smart wearable device to obtain a target synthesized translated speech signal. The target slave smart wearable device can be any one of the multiple slave smart wearable devices. The smart wearable master device receives and plays the synthesized translated speech signal, and the smart wearable master device receives the target synthesized translated speech signal and sends the target synthesized translated speech signal to the target smart wearable device, so that the target smart wearable slave device receives and plays the target synthesized translated speech signal.
3. The method according to any one of claims 1 or 2, characterized in that, The smart wearable master device identifies the target language corresponding to the smart wearable master device and the plurality of smart wearable slave devices, including: The smart wearable master device identifies the target language corresponding to the smart wearable master device and the multiple smart wearable slave devices based on the user voice signals collected by the smart wearable master device and the user voice signals collected by the multiple smart wearable slave devices respectively; Alternatively, the smart wearable master device may determine the target language corresponding to the smart wearable master device in response to the user's language selection operation, and the smart wearable master device may identify the target language corresponding to the multiple smart wearable slave devices based on the target language information sent by the multiple smart wearable slave devices respectively, wherein the target language information is obtained by the smart wearable slave devices based on the collected user voice signal or in response to the user's language selection operation.
4. The method according to any one of claims 1 or 2, characterized in that, After the smart wearable master device identifies the target language corresponding to the smart wearable master device and the plurality of smart wearable slave devices respectively, the method includes: If the smart wearable master device detects that the target languages corresponding to the smart wearable master device and the multiple smart wearable slave devices are different, the smart wearable master device executes a multi-party simultaneous interpretation working mode to perform the step of sending all the user voice signals and the target languages corresponding to the smart wearable master device and the multiple smart wearable slave devices to the cloud. If the smart wearable master device identifies that the target language corresponding to the smart wearable master device and the multiple smart wearable slave devices is the same, the smart wearable master device switches to a multi-party call working mode to synthesize the user voice signals (excluding those of the smart wearable master device) into a synthesized voice signal based on the unique device identifiers corresponding to the smart wearable master device and the multiple smart wearable slave devices, and to synthesize the voice signals (excluding those of the target smart wearable slave device) into a target synthesized voice signal. The smart wearable master device plays the synthesized speech signal and sends the target synthesized speech signal to the target smart wearable slave device, so that the target smart wearable slave device receives the target synthesized speech signal and plays it.
5. The method according to any one of claims 1 or 2, characterized in that, The cloud-based speech translation processing includes: The cloud platform, based on the target languages corresponding to the smart wearable master device and the multiple smart wearable slave devices, converts all user voice signals (excluding those from the smart wearable master device itself) into text information, then translates the text information into text information in the target language corresponding to the smart wearable master device, and finally converts the translated text information into voice signals in the target language corresponding to the smart wearable master device; and also converts all user voice signals (excluding those from the smart wearable slave devices themselves) into text information, then translates the text information into text information in the target language corresponding to the smart wearable slave devices, and finally converts the translated text information into voice signals in the target language corresponding to the smart wearable slave devices.
6. The method according to any one of claims 1 or 2, characterized in that, The smart wearable main device plays the synthesized translated speech signal, including: The smart wearable master device broadcasts the synthesized translated speech signal according to the user voiceprints corresponding to the multiple smart wearable slave devices respectively, or the smart wearable master device broadcasts the synthesized translated speech signal according to a preset electronic voice. Correspondingly, the process of receiving and playing the target synthesized translated speech signal from the device includes: The target smart wearable slave device broadcasts the target synthesized translated speech signal according to the user voiceprints corresponding to the plurality of smart wearable slave devices other than the target smart wearable slave device and the smart wearable master device, or the target smart wearable slave device broadcasts the target synthesized translated speech signal according to a preset electronic voice.
7. A multi-party calling method, applied to a smart wearable main device, characterized in that, The smart wearable master device is wirelessly connected to multiple smart wearable slave devices, and the smart wearable master device is also connected to the cloud. The method includes: The smart wearable master device collects the corresponding user voice signal and receives the user voice signals collected by the multiple smart wearable slave devices respectively; Identify the target language corresponding to the smart wearable master device and the plurality of smart wearable slave devices respectively; All user voice signals, as well as the target languages corresponding to the smart wearable master device and the multiple smart wearable slave devices, are sent to the cloud so that the cloud can perform voice translation processing. The voice translation processing includes translating each user voice signal except for the smart wearable master device itself into the target language voice signal corresponding to the smart wearable master device, and translating each user voice signal except for the smart wearable slave device itself into the target language voice signal corresponding to the smart wearable slave device. The system receives translated target language speech signals sent from the cloud and performs speech synthesis processing. The speech synthesis processing includes synthesizing the translated target language speech signals corresponding to the smart wearable master device to obtain a synthesized translated speech signal; and synthesizing the translated target language speech signals corresponding to the target smart wearable slave device to obtain a target synthesized translated speech signal. The target smart wearable slave device is any one of the plurality of smart wearable slave devices. Play the synthesized translated speech signal and send the target synthesized translated speech signal to the target smart wearable slave device, so that the target smart wearable slave device receives the target synthesized translated speech signal and plays it.
8. A multi-party calling method, applied to a smart wearable main device, characterized in that, The smart wearable master device is wirelessly connected to multiple smart wearable slave devices, and the smart wearable master device is also connected to the cloud. The method includes: The smart wearable master device collects the corresponding user voice signal and receives the user voice signals collected by the multiple smart wearable slave devices respectively; Identify the target language corresponding to the smart wearable master device and the plurality of smart wearable slave devices respectively; All user voice signals, along with the target languages corresponding to the smart wearable master device and the plurality of smart wearable slave devices, are sent to the cloud. The cloud then performs voice translation and voice synthesis processing. The voice translation processing includes translating all user voice signals (excluding those from the smart wearable master device itself) into the target language voice signal corresponding to the smart wearable master device, and translating all user voice signals (excluding those from the smart wearable slave devices themselves) into the target language voice signal corresponding to the smart wearable slave device. The voice synthesis processing includes synthesizing the translated target language voice signals from the smart wearable master device to obtain a synthesized translated voice signal, and synthesizing the translated target language voice signals from the target smart wearable slave device to obtain a target synthesized translated voice signal. The target smart wearable slave device can be any one of the plurality of smart wearable slave devices. The device receives and plays the synthesized translated speech signal, and receives the target synthesized translated speech signal and sends the target synthesized translated speech signal to the target smart wearable device, so that the target smart wearable device receives and plays the target synthesized translated speech signal.
9. A multi-party calling method, applied to a smart wearable slave device, characterized in that, Multiple smart wearable slave devices are wirelessly connected to a smart wearable master device, and the smart wearable master device is also connected to a cloud. The method includes: Multiple smart wearable slave devices respectively collect corresponding user voice signals and send the collected user voice signals to the smart wearable master device, so that the smart wearable master device sends its own collected user voice signal, the user voice signals collected by the multiple smart wearable slave devices respectively, and the target language corresponding to the smart wearable master device and the multiple smart wearable slave devices respectively to the cloud, so that the cloud performs voice translation processing. The voice translation processing includes translating each user voice signal other than that of the smart wearable slave device itself into the target language voice signal corresponding to the smart wearable slave device based on the target language corresponding to the smart wearable master device and the multiple smart wearable slave devices respectively. The system receives and plays the target synthesized translated speech signal sent by the smart wearable master device. The target synthesized translated speech signal is obtained by the smart wearable master device receiving translated speech signals of various target languages sent by the cloud and performing speech synthesis processing. The speech synthesis processing includes synthesizing speech signals of various target languages corresponding to the target smart wearable slave device. The target smart wearable slave device is any one of the multiple smart wearable slave devices.
10. A multi-party calling method, applied to a smart wearable slave device, characterized in that, Multiple smart wearable slave devices are wirelessly connected to a smart wearable master device, and the smart wearable master device is also connected to a cloud. The method includes: Multiple smart wearable slave devices respectively collect corresponding user voice signals and send the collected user voice signals to the smart wearable master device. The smart wearable master device then sends its own collected user voice signal, the user voice signals collected by the multiple smart wearable slave devices, and the target language corresponding to the smart wearable master device and the multiple smart wearable slave devices to the cloud. The cloud then performs speech translation processing and speech synthesis processing. The speech translation processing includes translating each user voice signal (excluding the smart wearable slave device itself) into a speech signal in the target language corresponding to the smart wearable slave device, based on the target language corresponding to the smart wearable master device and the multiple smart wearable slave devices. The speech synthesis processing includes synthesizing the translated target language speech signals corresponding to the target smart wearable slave device to obtain a target synthesized translation speech signal. The target smart wearable slave device can be any one of the multiple smart wearable slave devices. The system receives and plays the target synthesized translated speech signal sent by the smart wearable main device, wherein the target synthesized translated speech signal is obtained by the smart wearable main device from the cloud.
11. A multi-party communication system for performing the method according to any one of claims 1-6, characterized in that, The multi-party calling system includes multiple smart wearable devices and a cloud. The multiple smart wearable devices include at least one smart wearable master device and multiple smart wearable slave devices. The smart wearable master device is wirelessly connected to the multiple smart wearable slave devices, and the smart wearable master device is also connected to the cloud.
12. The multi-party calling system according to claim 11, characterized in that, The plurality of smart wearable devices are configured into at least one smart wearable device group, each smart wearable device group including one smart wearable master device and a plurality of smart wearable slave devices, the smart wearable master device in the smart wearable device group is wirelessly connected to the plurality of smart wearable slave devices, and each smart wearable master device in each smart wearable device group is communicatively connected to the cloud.
13. The multi-party calling system according to claim 11 or 12, characterized in that, The smart wearable device can switch between host mode and slave mode. When the smart wearable device switches to host mode, it becomes the smart wearable master device. When the smart wearable device switches to slave mode, it becomes the smart wearable slave device.
14. The multi-party calling system according to claim 11 or 12, characterized in that, The smart wearable main device includes: The smart wearable master device communicates with each of the smart wearable slave devices and the cloud through the wireless communication module. A voice acquisition module, which is used to acquire user voice signals; The language recognition module is used to identify the target language corresponding to the smart wearable main device and the plurality of smart wearable slave devices respectively; And / or a speech synthesis module, wherein the speech synthesis module is used to perform speech synthesis processing, the speech synthesis processing including synthesizing the translated target language speech signals corresponding to the smart wearable master device to obtain a synthesized translated speech signal; and synthesizing the translated target language speech signals corresponding to the target smart wearable slave device to obtain a target synthesized translated speech signal, wherein the target smart wearable slave device is any one of the plurality of smart wearable slave devices; A voice broadcast module is used to play the synthesized translated voice signal.
15. The multi-party calling system according to claim 14, characterized in that, The smart wearable main device also includes: The NFC module is used to perform NFC communication and trigger the smart wearable slave device to establish a communication connection with the smart wearable master device when the smart wearable master device is close to the smart wearable slave device; and / or a power supply module, wherein the power supply module is connected to the wireless communication module, the voice acquisition module, the language recognition module, the voice broadcasting module, the NFC module and / or the voice synthesis module respectively, to supply power to the wireless communication module, the voice acquisition module, the language recognition module, the voice broadcasting module, the NFC module and / or the voice synthesis module.
16. The multi-party calling system according to claim 14, characterized in that, The smart wearable slave device includes: A wireless communication unit, through which the smart wearable slave device communicates with the smart wearable master device; A voice acquisition unit, wherein the voice acquisition unit is used to acquire user voice signals; A voice broadcasting unit is used to play the target synthesized translated voice signal.
17. The multi-party calling system according to claim 16, characterized in that, The smart wearable device also includes: The NFC unit is used to perform NFC communication and trigger the establishment of a communication connection between the smart wearable slave device and the smart wearable master device when the smart wearable slave device is close to the smart wearable master device; and / or a power supply module, wherein the power supply module is connected to the wireless communication unit, the voice acquisition unit, the voice broadcasting unit, and the NFC unit respectively, to provide power to the wireless communication unit, the voice acquisition unit, the voice broadcasting unit, and the NFC unit.
18. A smart wearable master device for performing the method of claim 7, characterized in that, The smart wearable main device includes: A wireless communication module is provided, through which the smart wearable master device communicates and connects with multiple smart wearable slave devices and the cloud. A voice acquisition module, which is used to acquire user voice signals; Language recognition module, which is used to identify the target language corresponding to the smart wearable main device and each of the smart wearable slave devices respectively; The speech synthesis module is used to perform speech synthesis processing, which includes synthesizing the translated target language speech signals corresponding to the smart wearable master device to obtain a synthesized translation speech signal; and synthesizing the translated target language speech signals corresponding to the target smart wearable slave device to obtain a target synthesized translation speech signal, wherein the target smart wearable slave device is any one of the plurality of smart wearable slave devices; A voice broadcast module is used to play the synthesized translated voice signal.
19. A smart wearable master device for performing the method of claim 8, characterized in that, The smart wearable main device includes: A wireless communication module is provided, through which the smart wearable master device communicates and connects with multiple smart wearable slave devices and the cloud. A voice acquisition module, which is used to acquire user voice signals; The language recognition module is used to identify the target language corresponding to the smart wearable main device and the multiple smart wearable slave devices respectively; A voice broadcast module is used to play synthesized translated voice signals.
20. A smart wearable slave device for performing the method of any one of claims 9 and 10, characterized in that, The smart wearable slave device includes: A wireless communication unit, through which the smart wearable slave device communicates with the smart wearable master device; A voice acquisition unit, wherein the voice acquisition unit is used to acquire user voice signals; A voice broadcasting unit is used to play the target synthesized translated voice signal.
21. A multi-party calling method, applied to a multi-party calling system, the multi-party calling system comprising multiple smart wearable devices and a cloud, wherein the multiple smart wearable devices are respectively communicatively connected to the cloud, characterized in that, The method includes: Any one of the plurality of smart wearable devices initiates a multi-party translation call request to establish a multi-party translation call; The multiple smart wearable devices participating in the multi-party translation call each collect the corresponding user voice signals; The multiple smart wearable devices each identify their respective target languages; The multiple smart wearable devices will send the collected user voice signals and corresponding target languages to the cloud. The cloud platform translates multiple user voice signals (excluding the target smart wearable device) into voice signals in the target language corresponding to the target smart wearable device, based on the target language corresponding to each of the multiple smart wearable devices. The target smart wearable device receives translated speech signals from various target languages sent by the cloud, and then synthesizes the translated speech signals from various target languages before playing them.
22. A multi-party calling method, applied to a multi-party calling system, the multi-party calling system comprising multiple smart wearable devices and a cloud, wherein the multiple smart wearable devices are respectively communicatively connected to the cloud, characterized in that, The method includes: Any one of the plurality of smart wearable devices initiates a multi-party translation call request to establish a multi-party translation call; The multiple smart wearable devices participating in the multi-party translation call each collect the corresponding user voice signals; The multiple smart wearable devices each identify their respective target languages; The multiple smart wearable devices will send the collected user voice signals and corresponding target languages to the cloud. The cloud platform translates multiple user voice signals (excluding the target smart wearable device) into voice signals in the target language corresponding to the target smart wearable device, based on the target language corresponding to each of the multiple smart wearable devices. The cloud platform performs speech synthesis on the translated speech signals of each target language corresponding to the target smart wearable device. The target smart wearable device receives and plays the synthesized speech signals of various target languages corresponding to the target smart wearable device.
23. The method according to claim 21 or 22, characterized in that, Any one of the plurality of smart wearable devices initiates a multi-party translation call request to establish a multi-party translation call, including: Any one of the multiple smart wearable devices initiates a multi-party translation call request and provides a method for joining the session, including session ID and password, session group QR code, session push invitation, or joining by NFC module near-field touch. Other smart wearable devices join the session based on the aforementioned session joining method to establish a multi-party translation call.
24. A multi-party communication system for performing the method according to any one of claims 21-23, characterized in that, The multi-party calling system includes multiple smart wearable devices and a cloud platform, with each of the multiple smart wearable devices communicating with the cloud platform.
25. The multi-party calling system according to claim 24, characterized in that, The smart wearable device includes: A wireless communication module, through which the smart wearable device communicates with the cloud; A voice acquisition module, which is used to acquire user voice signals; A language recognition module is used to identify the target language corresponding to each of the smart wearable devices. and / or a speech synthesis module, wherein the speech synthesis module is used to synthesize the translated speech signals of each target language corresponding to the target smart wearable device, wherein the target smart wearable device is any one of the plurality of smart wearable devices; The voice broadcast module is used to play the synthesized voice signals of various target languages corresponding to the target smart wearable device.