Conversation translation method and related apparatus

By automatically detecting the user's speech and sentence breaks through the headphones, dialogue translation can be achieved without manual operation, improving translation accuracy and user experience.

WO2026108918A1PCT designated stage Publication Date: 2026-05-28HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2025-11-20
Publication Date
2026-05-28

AI Technical Summary

Technical Problem

Existing dialogue translation technologies require users to manually operate their phones to alternately collect voice messages, resulting in cumbersome operation and a poor user experience.

Method used

The system detects the user's speech and pauses through headphones, automatically collects the voice, and sends it to the terminal device for translation. The terminal device then performs an accurate translation based on the pause detection, reducing the need for user intervention.

Benefits of technology

It enables automatic voice capture and translation without manual operation, improving translation accuracy and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025136393_28052026_PF_FP_ABST
    Figure CN2025136393_28052026_PF_FP_ABST
Patent Text Reader

Abstract

The present application discloses a conversation translation method and a related apparatus. The method comprises: when an earphone detects that an earphone-side user speaks, collecting a first speech; the earphone sending the first speech to a terminal device; and when the earphone detects the earphone-side user segmenting a sentence, sending a first identifier to the terminal device, and on the basis of the first identifier, the terminal device translating the first speech to obtain a translation result of the first speech. In this way, a user wearing the earphone for conversation translation can complete speech reception, sentence segmentation, and translation of the user without manual operations, thereby effectively reducing user operations during conversation translation of the user, improving the translation accuracy, and improving user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Dialogue translation method and related device

[0001] The present application claims priority to the Chinese patent application No. 202411701905.1, filed on November 25, 2024, with the State Intellectual Property Office of China, and the Chinese patent application No. 202411701905.1 has the title of "Dialogue translation method and related device", the whole content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] The present application relates to the field of electronic technology, in particular to a dialogue translation method and related device. BACKGROUND

[0003] With the development of electronic technology, terminal devices such as mobile phones can provide real-time bilingual dialogue translation for conversations between users. For example, a mobile phone installs a translation application, which can collect the user's voice through the microphone of the mobile phone, obtain and display the translation result of the voice.

[0004] Currently, in actual use scenarios, different users of dialogue translation need to constantly alternate the use of mobile phones to complete real-time bilingual dialogue translation. For example, user 1 first operates the mobile phone, selects the translated language 2, and manually controls the mobile phone to collect the voice of user 1, translates and displays the translation result of language 2; then, user 2 operates the above mobile phone, selects the translated language 1, and manually controls the mobile phone to collect the voice of user 2, translates and displays the translation result of language 1. In this way, the mobile phone is alternately used to manually trigger voice collection, and the microphone of the mobile phone is alternately used to collect the voice of the user, which is cumbersome to operate and has poor user experience. SUMMARY

[0005] The embodiments of the present application provide a dialogue translation method and related device, which can complete the functions of receiving, punctuating, and translating for a user wearing a headset without manual operation, effectively reducing the user operation of the user during dialogue translation, improving the accuracy of translation, and improving the user experience.

[0006] In a first aspect, the present application provides a dialogue translation method applied to a dialogue translation system, the dialogue translation system comprising a terminal device and a headset, the terminal device and the headset establishing a communication connection, the method comprising: the headset detecting that the headset-side user is speaking, and collecting a first voice; the headset sending the first voice to the terminal device; the headset detecting that the headset-side user punctuates, and sending a first identifier to the terminal device; and based on the first identifier, the terminal device translates the first voice to obtain a translation result of the first voice.

[0007] The earphone automatically collects the voice of the user when detecting that the user on the earphone side speaks, and sends the voice to the terminal device. The earphone automatically transmits an end identifier to the terminal device when detecting that the user pauses, to instruct the terminal device to start translation. In this way, for the user on the earphone side, the user can complete the functions of voice collection, pausing, and translation without manual operation, effectively reducing the user operation when the user's speech is translated. Based on the end identifier of the pausing detection, the terminal device can accurately pause the voice of the user on the earphone side, and then translate by sentence, effectively improving the accuracy of translation and improving the user experience.

[0008] In an implementation manner, the terminal device translates the first voice based on the first identifier, including: in a case where a first application of the terminal device runs in a foreground, the first application translates the first voice based on the first identifier; or, in a case where a first interface of the first application of the terminal device runs in the foreground, the first application translates the first voice based on the first identifier, and the first interface is used for conversation translation. It can be understood that the first application / first interface is not running in the foreground, indicating that the user is not interested in the voice of the conversation translation. In the embodiment of the application, when the first application / first interface is not running in the foreground, the voice of the conversation translation does not need to be translated. In this way, unnecessary resource waste can be avoided.

[0009] In an implementation manner, when the first voice is collected, the method further includes: the earphone performs pausing detection when detecting that the user on the earphone side stops speaking; and the earphone determines that the user on the earphone side pauses when detecting that the duration for which the user on the earphone side stops speaking reaches a first duration. In the embodiment of the application, when the duration for which the user on the earphone side stops speaking reaches the first duration, it is determined that the user on the earphone side pauses, to instruct the terminal device to pause and translate the voice of the user on the earphone side. In this way, the terminal device can accurately pause the voice of the user on the earphone side, and then translate by sentence, effectively improving the accuracy of translation.

[0010] In an implementation manner, when the first voice is collected, the method further includes: the earphone stops collecting the first voice when detecting that the user on the earphone side stops speaking; or, the earphone stops collecting the first voice when detecting that the user on the earphone side pauses. In the embodiment of the application, the earphone stops collecting the audio stream on the earphone side when the user stops speaking. In this way, the collection and processing of the audio stream without voice can be avoided, and the energy consumption of the device is saved.

[0011] In an implementation manner, before the earphone detects that the user speaks and collects the first voice, the method further includes: the terminal device receives a first user operation; in response to the first user operation, the terminal device displays a first interface and sends a first message to the earphone; the first interface is used for conversation translation; based on the first message, the earphone enters a conversation translation mode; the earphone detects that the earphone-side user speaks and collects the first voice, including: in the conversation translation mode, the earphone detects that the earphone-side user speaks and collects the first voice; the earphone detects that the earphone-side user pauses and sends a first identifier to the terminal device, including: in the conversation translation mode, the earphone detects that the earphone-side user pauses and sends the first identifier to the terminal device. By implementing the embodiment of the application, when the terminal device enters the interface for conversation translation, the earphone is instructed to enter the conversation translation mode; in the conversation translation mode, the earphone can detect whether the earphone-side user speaks and pauses, and then realize the listening, pausing and translation of the user, effectively reducing the user operation of the user in the conversation translation, effectively improving the accuracy of translation, and improving the user experience.

[0012] In an implementation manner, after the earphone enters the conversation translation mode, the method further includes: the terminal device receives a second user operation; in response to the second user operation, the terminal device collects a second voice and sends a second message to the earphone; based on the second message, the earphone exits the conversation translation mode. By implementing the embodiment of the application, the priority of the terminal device listening is higher than that of the earphone listening, and when the terminal device listens, the earphone can exit the conversation translation mode; so as to avoid the confusion of conversation translation caused by the two users speaking at the same time.

[0013] In an implementation manner, after the earphone exits the conversation translation mode, the method further includes: the terminal device receives a third user operation; in response to the third user operation, the terminal device stops collecting the second voice and sends a third message to the earphone; based on the third message, the earphone enters the conversation translation mode. By implementing the embodiment of the application, after the earphone exits the conversation translation mode, when the terminal device stops listening, the earphone can enter the conversation translation mode again; in this way, the terminal device and the earphone alternately listen, and real-time bilingual conversation translation can be realized.

[0014] In an implementation manner, after the earphone exits the conversation translation mode, the method further including: the terminal device detects that the terminal-side user pauses and sends a third message to the earphone; based on the third message, the earphone enter the conversation translation mode.

[0015] In an implementation manner, after the second voice is collected, the method further includes: the terminal device detects that the terminal-side user pauses and stops collecting the second voice. By implementing the embodiment of the application, when the earphone-side user pauses, the earphone can stop collecting the earphone-side audio stream; in this way, the collection and processing of the audio stream without voice can be avoided, and the device energy consumption is saved.

[0016] In a second aspect, the present application provides a translation method applied to a headset, the method comprising: collecting a first voice when the headset detects that a user on the headset side speaks; sending the first voice to a terminal device; and sending a first identifier to the terminal device when the headset detects that the user on the headset side pauses, the first identifier being used to instruct the terminal device to translate the first voice.

[0017] By implementing the embodiments of the present application, the headset automatically collects the voice of the user when detecting that the user on the headset side speaks, and sends the voice to the terminal device; the headset automatically transmits an end identifier to the terminal device when detecting that the user pauses, to instruct the terminal device to start translation. In this way, for the user on the headset side, the user can complete the operations of receiving, pausing and translating the voice of the user without manual operation, effectively reducing the user operation when the user is in a conversation; based on the end identifier of the pause detection, the terminal device can accurately pause the voice of the user on the headset side, and then translate by sentence, effectively improving the accuracy of translation and improving the user experience.

[0018] In an implementation manner, when the first voice is collected, the method further comprises: detecting the pause when the headset detects that the user on the headset side stops speaking; and determining that the user pauses when the headset detects that the duration that the user on the headset side stops speaking reaches a first duration. By implementing the embodiments of the present application, when the duration that the user on the headset side stops speaking reaches the first duration, it is determined that the user on the headset side pauses, to instruct the terminal device to pause and translate the voice of the user on the headset side; in this way, the terminal device can accurately pause the voice of the user on the headset side, and then translate by sentence, effectively improving the accuracy of translation.

[0019] In an implementation manner, when the first voice is collected, the method further comprises: stopping collecting the first voice when the headset detects that the user on the headset side stops speaking; or stopping collecting the first voice when the headset detects that the user on the headset side pauses. By implementing the embodiments of the present application, the headset stops collecting the audio stream on the headset side when the user stops speaking; in this way, it can be avoided to collect and process the audio stream without voice, and then the energy consumption of the device is saved.

[0020] In an implementation manner, before the earphone collects the first voice when the earphone detects that the user on the earphone side speaks, the method further includes: the earphone receiving a first message sent by the terminal device; based on the first message, the earphone entering a conversation translation mode; the earphone collecting the first voice when the earphone detects that the user on the earphone side speaks, including: in the conversation translation mode, the earphone collecting the first voice when the earphone detects that the user on the earphone side speaks; the earphone sending a first identifier to the terminal device when the earphone detects that the user pauses, including: in the conversation translation mode, the earphone sending the first identifier to the terminal device when the earphone detects that the user on the earphone side pauses. By implementing the embodiment of the application, when the terminal device enters an interface for conversation translation, the earphone is instructed to enter the conversation translation mode; in the conversation translation mode, the earphone can detect whether the user on the earphone side speaks and pauses, thereby realizing the listening, pausing and translation of the user, effectively reducing the user operation of the user in the conversation translation, effectively improving the accuracy of translation, and improving the user experience.

[0021] In an implementation manner, after the earphone enters the conversation translation mode, the method further includes: the earphone receiving a second message sent by the terminal device; the second message is sent by the terminal device when the terminal device detects a user operation of triggering the terminal device to start listening; based on the second message, the earphone exiting the conversation translation mode. By implementing the embodiment of the application, the priority of the terminal device listening is higher than that of the earphone listening, and when the terminal device listens, the earphone can exit the conversation translation mode; so as to avoid the conversation translation being confused due to two users speaking at the same time.

[0022] In an implementation manner, after the earphone exits the conversation translation mode, the method further includes: the earphone receiving a third message sent by the terminal device; the third message is sent by the terminal device when the terminal device detects a user operation of triggering the terminal device to stop listening, or the third message is sent by the terminal device when the terminal device detects that the user on the terminal side pauses; based on the third message, the earphone entering the conversation translation mode. By implementing the embodiment of the application, after the earphone exits the conversation translation mode, when the terminal device stops listening, the earphone can enter the conversation translation mode again; in this way, the terminal device and the earphone alternately listen, and real-time bilingual conversation translation can be realized.

[0023] In an implementation manner, after the earphone sends the first identifier to the terminal device, the method further includes: the earphone receiving a translation result of the first voice sent by the terminal device; and the earphone voice broadcasting the translation result of the first voice.

[0024] In an implementation manner, the earphone includes a VPU, and a bone conduction signal collected by the VPU is used to detect whether the user wearing the earphone speaks; the earphone collecting the first voice when the earphone detects that the user on the ear phone side speaks, including: the earphone collecting the first voice according to the bone conduction signal collected by the VPU when the earphone detects that the user on the earphone side speaks.

[0025] In a third aspect, the present application provides a method for dialogue translation, applied to a terminal device, the method comprising: receiving, by the terminal device, a first voice sent by a headset; receiving, by the terminal device, a first identifier sent by the headset, the first identifier indicating that a user on the headset side has paused; and based on the first identifier, translating, by the terminal device, the first voice to obtain a translation result of the first voice.

[0026] According to the embodiments of the present application, when the headset detects that the user on the headset side is speaking, the headset automatically collects the voice of the user and sends the voice to the terminal device; and when the headset detects that the user has paused, the headset automatically transmits an end identifier to the terminal device to instruct the terminal device to start translation. In this way, for the user on the headset side, the user can complete the voice collection, pausing and translation of the user without manual operation, effectively reducing the user operation during dialogue translation; based on the end identifier of the pausing detection, the terminal device can accurately pause the voice of the user on the headset side and then translate the voice by sentence, effectively improving the accuracy of translation and improving the user experience.

[0027] In an implementation manner, the step of translating, by the terminal device, the first voice based on the first identifier comprises: in a case where a first application of the terminal device is running in the foreground, translating, by the first application, the first voice based on the first identifier; or in a case where a first interface of the first application of the terminal device is running in the foreground, translating, by the first application, the first voice based on the first identifier, the first interface being used for dialogue translation. It can be understood that the first application / first interface is not running in the foreground, indicating that the user is not interested in the voice for dialogue translation. According to the embodiments of the present application, when the first application / first interface is not running in the foreground, the voice for dialogue translation does not need to be translated. In this way, unnecessary resource waste can be avoided.

[0028] In an implementation manner, before the step of receiving, by the terminal device, the first voice sent by the headset, the method further comprises: receiving, by the terminal device, a first user operation; in response to the first user operation, displaying, by the terminal device, a first interface and sending, by the terminal device, a first message to the headset, the first interface being used for dialogue translation and the first message being used for instructing the headset to enter a dialogue translation mode. According to the embodiments of the present application, when the terminal device enters the interface for dialogue translation, the headset is instructed to enter the dialogue translation mode; in the dialogue translation mode, the headset can detect whether the user on the headset side is speaking and pausing, and then realize the voice collection, pausing and translation of the user, effectively reducing the user operation during dialogue translation of the user, effectively improving the accuracy of translation and improving the user experience.

[0029] In an implementation manner, after the first message is sent to the earphone, the method further includes: receiving a second user operation by the terminal device; and in response to the second user operation, collecting a second voice by the terminal device, and sending a second message to the earphone, the second message being used to instruct the earphone to exit the conversation translation mode. By implementing the embodiment of the application, the priority of the terminal device to collect the sound is higher than that of the earphone, and when the terminal device collects the sound, the earphone can exit the conversation translation mode, so as to avoid the confusion of the conversation translation caused by the simultaneous speech of the two users.

[0030] In an implementation manner, after the second message is sent to the earphone, the method further includes: receiving a third user operation by the terminal device; and in response to the third user operation, stopping collecting the second voice by the terminal device, and sending a third message to the earphone, the third message being used to instruct the earphone to enter the conversation translation mode. By implementing the embodiment of the application, after the earphone exits the conversation translation mode, when the terminal device stops collecting the sound, the earphone can enter the conversation translation mode again, so that the terminal device and the earphone can collect the sound alternately, and the real-time bilingual conversation translation can be realized.

[0031] In an implementation manner, after the second voice is collected, the method further includes: sending a third message to the earphone when it is detected that the terminal-side user pauses, the third message being used to instruct the earphone to enter the conversation translation mode.

[0032] In an implementation manner, after the second voice is collected, the method further includes: stopping collecting the second voice when it is detected that the terminal-side user pauses. By implementing the embodiment of the application, when the terminal-side user pauses, the terminal device can stop collecting the terminal-side audio stream, so that the audio stream without voice can be avoided to be collected and processed, and the device energy consumption can be saved.

[0033] In a fourth aspect, an embodiment of the application provides an electronic device, which includes: a processor and a memory, the memory being coupled with the processor, and the memory being used to store computer program codes, the computer program codes including computer instructions, and the computer instructions being read by the processor from the memory, so that the electronic device executes the conversation translation method in the second aspect.

[0034] In a fifth aspect, an embodiment of the application provides an earphone, which includes: a processor and a memory, the memory being coupled with the processor, and the memory being used to store computer program codes, the computer program codes including computer instructions, and the computer instructions being read by the processor from the memory, so that the earphone executes the conversation translation method in the third aspect.

[0035] In a sixth aspect, an embodiment of the present application provides a computer storage medium, including computer instructions, when the computer instructions are executed on an electronic device, causing the electronic device to perform the dialogue translation method in any possible implementation manner of any one of the aspects.

[0036] In a seventh aspect, an embodiment of the present application provides a computer program product, when the computer program product is executed on a computer, causing the computer to perform the dialogue translation method in any possible implementation manner of any one of the aspects. BRIEF DESCRIPTION OF DRAWINGS

[0037] FIG. 1 is a schematic diagram of a system architecture of a communication system according to an embodiment of the present application;

[0038] FIG. 2 is a schematic diagram of a structure of a terminal device according to an embodiment of the present application;

[0039] FIG. 3 is a schematic diagram of a structure of an earphone according to an embodiment of the present application;

[0040] FIG. 4 is a schematic diagram of a software architecture according to an embodiment of the present application;

[0041] FIG. 5 is a schematic diagram of a dialogue translation scenario according to an embodiment of the present application;

[0042] FIGS. 6A to 6F are interactive interfaces of dialogue translation according to an embodiment of the present application;

[0043] FIGS. 7A to 7G are another interactive interfaces of dialogue translation according to an embodiment of the present application;

[0044] FIG. 8 is a state machine of dialogue translation according to an embodiment of the present application;

[0045] FIGS. 9A and 9B are method flow diagrams of a dialogue translation method according to an embodiment of the present application;

[0046] FIG. 9C is an interactive flow diagram of a dialogue translation method according to an embodiment of the present application;

[0047] FIG. 10 is a schematic diagram of another dialogue translation scenario according to an embodiment of the present application;

[0048] FIGS. 11A to 11E are another interactive interfaces of dialogue translation according to an embodiment of the present application;

[0049] FIG. 12 is a method flow diagram of a dialogue translation method according to an embodiment of the present application. DETAILED DESCRIPTION

[0050] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings. In the description of the embodiments of the present application, unless otherwise specified, " / " represents the meaning of or, for example, A / B can represent A or B; the "and / or" in the text only describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent: A alone, A and B together, and B alone, in addition, in the description of the embodiments of the present application, "multiple" means two or more than two.

[0051] Hereinafter, the terms "first" and "second" are only used for description purposes, and cannot be understood as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first" and "second" can explicitly or implicitly include one or more of the features, and in the description of the embodiments of the present application, unless otherwise specified, the meaning of "multiple" is two or more than two.

[0052] The term "user interface (UI)" in the following embodiments of the present application is a medium interface for interaction and information exchange between an application program or an operating system and a user, which realizes the conversion between the internal form of information and the form that the user can accept. The user interface is source code written in a specific computer language such as Java, extensible markup language (XML), etc. The interface source code is parsed, rendered and finally presented as content that can be recognized by the user on the electronic device. The commonly used form of user interface is graphic user interface (GUI), which refers to a user interface related to computer operation displayed in a graphical manner. It can be a visual interface element such as text, icon, button, menu, tab, text box, dialog box, status bar, navigation bar, Widget, etc. displayed in the display screen of the electronic device. For example, FIG. 1 shows a system architecture of a communication system 10 provided in an embodiment of the present application.

[0053] As shown in FIG. 1, the communication system 10 includes a terminal device 100 (such as a mobile phone 100) and a headset 200.

[0054] Terminal device 100 and headset 200 can establish a communication connection; terminal device 100 is equipped with a translation application (APP), which provides dialogue translation functionality. For example, the translation APP can provide real-time bilingual dialogue translation for two users speaking different languages; wherein at least one user performing the dialogue translation wears headset 200 and receives audio through headset 200. In this embodiment, headset 200 receiving audio means that headset 200 uses audio acquisition elements (e.g., microphone, voice pickup sensor (VPU)) to acquire the audio stream on the headset side; if the user on the headset side is speaking when headset 200 receives audio, then the audio stream on the headset side includes the voice of the user on the headset side.

[0055] In this embodiment, the earphone 200 can detect the speaking state of the wearer (i.e., the user on the earphone side). When the earphone 200 detects that the user on the earphone side is "speaking," it automatically collects the user's voice and sends it to the terminal device 100. When the earphone 200 detects that the user on the earphone side is "not speaking," it performs sentence segmentation detection. When a sentence segmentation is detected, it automatically sends an end marker to the terminal device 100 to instruct the terminal device 100 to translate the voice collected by the earphone 200 from the user on the earphone side. The terminal device 100 can display the translation result of the voice. Specifically, when the user is detected to be "not speaking" or to have a sentence segmentation, the earphone 200 can automatically stop collecting the user's voice. Thus, for the user wearing the earphone 200, no manual operation is required; the recording and translation of the user's voice are automatically completed, achieving more convenient dialogue translation and effectively improving the user experience. Furthermore, through the end marker, the terminal device 100 can accurately segment the user's voice on the earphone side and then translate it sentence by sentence, effectively improving the accuracy of the translation.

[0056] In some embodiments, upon receiving the aforementioned end marker, if the terminal device 100 detects that the translation application / its dialogue translation interface is running in the foreground, it translates the speech sent by the headset 200; otherwise, no translation is required. It can be understood that if the translation application / its dialogue translation interface is not running in the foreground, it indicates that the current user is not paying attention to dialogue translation. This avoids the waste of resources caused by "translating speech that the user is not paying attention to."

[0057] In some embodiments, the terminal device 100 displays the translation result of the user's voice on the headset side, which may include one or more of the following: displaying the translation result as text on the display screen of the terminal device 100; broadcasting the translation result aloud through the speaker of the terminal device 100; sending the translation result / the voice message corresponding to the translation result to the headset 200, instructing the headset 200 to broadcast the translation result aloud.

[0058] In some embodiments, the headset 200 includes a voice pickup sensor (VPU) and a voice activity detection (VAD) module. The VAD module can detect the user's speaking state and sentence breaks based on the bone conduction signals collected by the VPU. The speaking state includes two types: "speaking" and "not speaking". The headset 200 is not limited to detecting speaking state and sentence breaks through the VPU; it can also detect speaking state and sentence breaks through other methods. This application embodiment does not specifically limit these methods.

[0059] In some embodiments, as shown in FIG1, the communication system 10 further includes a translation server 300, and the terminal device 100 and the translation server 300 can establish a communication connection; the terminal device 100 can upload the speech / text corresponding to the speech to the translation server 300; the translation server 300 can obtain the translation result of the speech and send it to the terminal device 100. In some embodiments, the communication system 10 does not need to include a translation server 300, and the terminal device 100 itself has offline translation capabilities and can translate the speech itself.

[0060] In this embodiment, the earphone 200 may contain only a single earphone body or a pair of earphone bodies (i.e., a left earphone worn on the left ear and a right earphone worn on the right ear). In one implementation, when the earphone 200 contains a pair of earphone bodies, the pair of earphone bodies may be divided into a main earphone and a secondary earphone; the main earphone can establish a communication connection with the terminal device 100 and communicate directly, while the secondary earphone needs to communicate indirectly with the terminal device 100 through the main earphone. In another implementation, when the earphone 200 contains a pair of earphone bodies, both the left and right earphones can establish a communication connection with the terminal device 100, and the left and right earphones are not distinguished as main and secondary.

[0061] In some embodiments, when the earphone 200 includes a pair of earpieces, the earphone 200 can receive sound, detect the user's speaking state, and pause sentences through the left and / or right earpieces. In one implementation, the earphone 200 has only one earpiece, which has the ability to receive sound, detect speaking state, and pause sentences. In another implementation, both earpieces of the earphone 200 have the ability to receive sound, detect speaking state, and pause sentences; when a user wears both earpieces of the earphone 200, the earphone 200 can comprehensively determine the user's speaking state and pause sentences based on the detection results of the two earpieces; for example, when at least one earpiece detects that the user is speaking, the earphone 200 determines that the user is "speaking".

[0062] The earphone 200 can be a wireless earphone, such as a TWS earphone or a neckband earphone. The earphone 200 can establish a wireless communication connection with the terminal device 100 through near-field communication technology, and then transmit the voice, end markers, and translation results collected by the earphone 200 to the terminal device 100 through this wireless communication connection. This application does not specifically limit the aforementioned near-field communication technology. For example, the aforementioned near-field communication technology can be Wireless Fidelity (WiFi) communication technology, Bluetooth communication technology (such as Classic Bluetooth (Basic Rate / Enhanced Data Rate, BR / EDR) or Bluetooth Low Energy (BLE)), Ultra Wide Band (UWB) communication technology, Near Field Communication (NFC) technology, or ZigBee communication technology, etc. The embodiments of this application do not specifically limit this.

[0063] The headset 200 can also be a wired headset. The headset 200 can establish a wired communication connection with the terminal device 100 through the headset jack, and then transmit the voice, end markers, and translation results collected by the headset 200 to the terminal device 100 through this wired communication connection. For example, the headset jack of the terminal device 100 can be a USB interface 130, or a 3.5mm open mobile electronic device platform standard interface, or a Cellular Telecommunications Industry Association standard interface.

[0064] Terminal device 100 can be a mobile phone, tablet computer, desktop computer, laptop computer, handheld computer, notebook computer, ultra-mobile personal computer (UMPC), netbook, cellular phone, personal digital assistant (PDA), augmented reality (AR) device, virtual reality (VR) device, artificial intelligence (AI) device, wearable device (e.g., smart bracelet), in-vehicle device, smart home device (e.g., smart TV, smart screen, large screen device, etc.) and / or smart city device. This application embodiment does not impose special limitations on the specific type of terminal device 100.

[0065] The translation server 300 can be a single server, a server cluster consisting of multiple servers, or a cloud computing center. The translation server 300 involved in this application embodiment can also be referred to as a cloud server, cloud-based, or cloud-side. The translation server 300 can establish a communication connection with the terminal device 100 through a communication network. This communication network can be implemented using any known network communication protocol, which can be various wired or wireless communication protocols, such as Ethernet, Universal Serial Bus (USB), FireWire, Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), Bluetooth, Wi-Fi, NFC, Voice over Internet Protocol (VoIP), communication protocols supporting network slicing architecture, or any other suitable communication protocol.

[0066] It should be understood that Figure 1 is merely a schematic diagram of the system structure of the communication system provided in the embodiment of this application, and does not constitute a specific limitation on the communication system 10. The communication system 10 may include more or fewer devices than shown in the figure.

[0067] The application scenarios of the dialogue translation method provided in this application include dialogue translation scenario one and dialogue translation scenario two.

[0068] In the dialogue translation scenario one, two users are translating a dialogue. One user receives audio through a terminal device 100, and the other user wears an earphone 200 and receives audio through the earphone 200. The terminal device 100 can translate the speech of the user on the earphone side captured by the earphone 200, obtain and display the corresponding translation result for the user on the earphone side, and can also translate the speech of the user on the terminal side captured by the terminal device 100, obtain and display the corresponding translation result for the user on the terminal side.

[0069] Wherein, "receiving sound from terminal device 100" means that terminal device 100 uses a microphone to collect audio streams from the terminal side; if the terminal user is speaking when terminal device 100 is receiving sound, then the terminal audio stream includes the terminal user's voice.

[0070] In the dialogue translation scenario two, two users are translating the dialogue. One user wears the left earpiece of headset 200 and receives audio through the left earpiece, while the other user wears the right earpiece of headset 200 and receives audio through the right earpiece. Terminal device 100 can translate the speech collected by the left earpiece and the right earpiece respectively, and obtain and display the translation results.

[0071] The following embodiments will provide a detailed description of the dialogue translation methods in two application scenarios.

[0072] The structure of a terminal device 100 provided in this application embodiment is described below. Figure 2 shows a schematic diagram of the structure of the terminal device 100.

[0073] Terminal device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0074] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the terminal device 100. In other embodiments of this application, the terminal device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0075] Processor 110 may include one or more processing units, such as application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU). These different processing units may be independent devices or integrated into one or more processors.

[0076] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.

[0077] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0078] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0079] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple I2C buses. The processor 110 can couple to the touch sensor 180K, charger, flash, camera 193, etc., through different I2C bus interfaces. For example, the processor 110 can couple to the touch sensor 180K through the I2C interface, enabling the processor 110 and the touch sensor 180K to communicate through the I2C bus interface, thereby realizing the touch function of the terminal device 100.

[0080] The I2S interface can be used for audio communication. In some embodiments, the processor 110 may include multiple I2S buses. The processor 110 can be coupled to the audio module 170 via the I2S bus to enable communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the I2S interface to enable the function of answering phone calls through a Bluetooth headset.

[0081] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled via the PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 via the PCM interface, enabling the function of answering phone calls through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.

[0082] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the UART interface to enable music playback through Bluetooth headphones.

[0083] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI) and a display serial interface (DSI). In some embodiments, the processor 110 and the camera 193 communicate via the CSI interface to enable the shooting function of the terminal device 100. The processor 110 and the display screen 194 communicate via the DSI interface to enable the display function of the terminal device 100.

[0084] The GPIO interface can be configured via software. It can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to a camera 193, a display screen 194, a wireless communication module 160, an audio module 170, a sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.

[0085] USB port 130 is a USB standard compliant interface, specifically a Mini USB port, Micro USB port, or USB Type-C port. USB port 130 can be used to connect a charger to charge terminal device 100, and can also be used for data transfer between terminal device 100 and peripheral devices. It can also be used to connect headphones for audio playback. This interface can also be used to connect other electronic devices, such as AR devices.

[0086] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the terminal device 100. In other embodiments of this application, the terminal device 100 may also adopt different interface connection methods or a combination of multiple interface connection methods as described in the above embodiments.

[0087] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the terminal device 100. While charging the battery 142, the charging management module 140 can also supply power to the terminal device 100 via the power management module 141.

[0088] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.

[0089] The wireless communication function of the terminal device 100 can be implemented through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.

[0090] Antennas 1 and 2 are used to transmit and receive electromagnetic wave signals. Each antenna in terminal device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with a tuning switch.

[0091] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the terminal device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.

[0092] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through an audio device (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.

[0093] The wireless communication module 160 can provide solutions for wireless communication applications on the terminal device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, demodulates and filters the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, frequency modulate and amplify them, and then convert them into electromagnetic waves for radiation via antenna 2.

[0094] In some embodiments, antenna 1 of terminal device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling terminal device 100 to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS).

[0095] Terminal device 100 implements display functions through a GPU, display screen 194, and application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0096] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a minimized LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the terminal device 100 may include one or N display screens 194, where N is a positive integer greater than 1.

[0097] Terminal device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.

[0098] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, converting it into an image visible to the naked eye. The ISP can also perform algorithmic optimization on image noise and brightness. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0099] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the terminal device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0100] A digital signal processor (DSP) is used to process digital signals. Besides digital image signals, it can also process other digital signals. For example, when terminal device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.

[0101] Video codecs are used to compress or decompress digital video. Terminal device 100 may support one or more video codecs. Thus, terminal device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, etc.

[0102] NPU stands for Neural Network (NN) Computing Processor. By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in terminal devices, such as image recognition, facial recognition, speech recognition, and text understanding.

[0103] Internal memory 121 may include one or more random access memory (RAM) and one or more non-volatile memory (NVM).

[0104] Random access memory can include static random-access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), and double data rate synchronous dynamic random access memory (DDR SDRAM, such as fifth-generation DDR SDRAM, which is generally called DDR5 SDRAM). Non-volatile memory can include disk storage devices and flash memory.

[0105] Flash memory can be classified according to its operating principle, including NOR FLASH, NAND FLASH, 3D NAND FLASH, etc.; according to the level of the storage cell, including single-level cell (SLC), multi-level cell (MLC), triple-level cell (TLC), quad-level cell (QLC), etc.; and according to the storage specification, including universal flash storage (UFS) and embedded multimedia card (eMMC), etc.

[0106] The random access memory can be directly read and written by the processor 110. It can be used to store executable programs (such as machine instructions) of the operating system or other running programs, as well as user and application data.

[0107] Non-volatile memory can also store executable programs and user and application data, and can be pre-loaded into random access memory for direct reading and writing by the processor 110.

[0108] The external memory interface 120 can be used to connect to external non-volatile memory, thereby expanding the storage capacity of the terminal device 100. The external non-volatile memory communicates with the processor 110 through the external memory interface 120 to perform data storage functions. For example, music, video, and other files can be stored in the external non-volatile memory.

[0109] Terminal device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.

[0110] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.

[0111] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The terminal device 100 can listen to music or make hands-free calls through the speaker 170A.

[0112] The receiver 170B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the terminal device 100 answers a phone call or voice message, the receiver 170B can be brought close to the listener's ear to hear the voice.

[0113] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Terminal device 100 may be equipped with at least one microphone 170C. In some embodiments, terminal device 100 may be equipped with two microphones 170C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, terminal device 100 may be equipped with three, four, or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.

[0114] In some embodiments, user 1 manually triggers terminal device 100 to pick up audio; when the user operation triggering terminal device 100 to pick up audio is detected, microphone 170C can collect a speech stream, which includes user 1's speech; terminal device 100 can translate user 1's speech in language 1 to obtain a translation result in language 2. In some embodiments, terminal device 100 converts the translation result into speech and plays the speech through speaker 170A.

[0115] The 170D headphone jack is used to connect wired headphones.

[0116] The pressure sensor 180A is used to sense pressure signals and can convert pressure signals into electrical signals.

[0117] The gyroscope sensor 180B can be used to determine the motion attitude of the terminal device 100. In some embodiments, the angular velocity of the terminal device 100 about three axes (i.e., the x, y, and z axes) can be determined by the gyroscope sensor 180B.

[0118] The 180C barometric pressure sensor is used to measure barometric pressure.

[0119] The magnetic sensor 180D includes a Hall sensor. The terminal device 100 can use the magnetic sensor 180D to detect the opening and closing of the flip cover.

[0120] The accelerometer 180E can detect the magnitude of acceleration of the terminal device 100 in various directions (generally three axes).

[0121] Distance sensor 180F is used to measure distance. Terminal device 100 can measure distance via infrared or laser.

[0122] The proximity light sensor 180G may include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode.

[0123] The 180L ambient light sensor is used to detect ambient light intensity.

[0124] The fingerprint sensor 180H is used to collect fingerprints.

[0125] The 180J temperature sensor is used to detect temperature.

[0126] Touch sensor 180K, also known as a "touch device," can be located on display screen 194. The touch sensor 180K and display screen 194 together form a touchscreen, also known as a "touchscreen." Touch sensor 180K detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180K may also be located on the surface of terminal device 100, in a different position than display screen 194.

[0127] The bone conduction sensor 180M can acquire vibration signals.

[0128] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Terminal device 100 can receive button input and generate key signal inputs related to user settings and function control of terminal device 100.

[0129] Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback.

[0130] Indicator 192 can be an indicator light, which can be used to indicate charging status, power changes, messages, missed calls, notifications, etc.

[0131] The SIM card interface 195 is used to connect the SIM card.

[0132] The structure of an earphone 200 provided in the embodiments of this application is described below.

[0133] Figure 3 shows a schematic diagram of the structure of the headset 200. The headset 200 may include a processor 210, a memory 220, a communication module 230, an audio module 240, a speaker 240A, a microphone 240B, a sensor module 250, etc. If the headset 200 is a wireless headset, it may also include a power module 260, a battery 260A, a charging interface 260C, and a power management module 260B.

[0134] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the terminal device 100. In other embodiments of this application, the terminal device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0135] The memory 220 can be used to store program code, such as program code for establishing virtual or physical connections between the headset and multiple electronic devices (e.g., terminal device 100), switching physical connections with electronic devices, handling dialogue translation services of electronic devices, and program code for charging the wireless headset and wireless pairing connections between the wireless headset and other electronic devices.

[0136] The processor 210 can be used to execute the above-described application code and call relevant modules to implement the functions of the headset 200 in this embodiment. For example, it can enable a virtual or physical connection between the headset 200 and the terminal device 100, detect the user's speaking state, collect the speech stream, detect sentence breaks, and play the audio corresponding to the translation results.

[0137] Processor 210 may include one or more processing units (e.g., Digital Signal Processor (DSP) 210A), which may be independent devices or integrated into one or more processors 210. Specifically, processor 210 may be an integrated control chip or may consist of circuitry including various active and / or passive components, and the circuitry is configured to perform the functions belonging to processor 210 as described in the embodiments of this application.

[0138] In some embodiments, the DSP 210A may include a VAD module, a preprocessing module, and an encoding module. VAD technology is primarily used for speech coding and speech recognition, simplifying speech processing and removing non-speech segments during audio conversations to capture the user's speech. The VAD module can be used to detect the user's speaking state and sentence breaks. The preprocessing module preprocesses the audio stream acquired by the headset 200, filtering out interference signals such as environmental noise. The encoding module encodes the speech acquired by the headset 200, which can then send the encoded speech data to the terminal device 100.

[0139] The communication module 230 can be a wired communication module and / or a wireless communication module. The wireless communication module can be used to support data exchange between the headset 200 and other electronic devices or the headset itself, including data exchange via wireless communication technologies such as BT, WLAN (such as WiFi), Zigbee, FM, NFC, IR, or general 2.4G / 5G wireless communication technologies.

[0140] In some embodiments, the wireless communication module can be a Bluetooth chip. The headset 200 can use this Bluetooth chip to pair with and establish a wireless connection with the Bluetooth chips of other electronic devices (e.g., terminal device 100) to achieve wireless communication and business processing with these other electronic devices. This wireless connection can be a physical connection or a virtual connection. Typically, the Bluetooth chip can support basic rate (BR) / enhanced data rate (EDR) Bluetooth and BLE, for example, it can send / receive page messages, send / receive BLE broadcast messages, etc.

[0141] In addition, the wireless communication module may also include an antenna. The wireless communication module receives electromagnetic waves through the antenna, modulates and filters the electromagnetic wave signals, and sends the processed signal to the processor 210. The wireless communication module can also receive signals to be transmitted from the processor 210, modulate and amplify them, and then convert them into electromagnetic waves for radiation through the antenna.

[0142] The audio module 240 can be used to manage audio data, enabling the input and output of audio signals to the headset 200, and allowing functions such as making and receiving calls, playing music, adjusting volume, acquiring audio streams for dialogue translation, and broadcasting translation results via the headset 200. The audio module 240 may include components such as a speaker 240A (or earpiece, receiver) and a microphone 240B (or microphone, transducer) for outputting audio signals. The speaker 240A can be used to convert audio electrical signals into sound signals and play them. The microphone 240B can be used to convert acquired sound signals into audio electrical signals (i.e., audio data).

[0143] The sensor module 250 can detect the necessary information according to different functions of the headphones 200. In some embodiments, the sensor module 250 may include a wear sensor 250A, which can detect whether the user is wearing the headphones. Optionally, the wear sensor 250A may be specifically implemented as one or more of the following: a proximity sensor, and motion sensors (e.g., accelerometers, gyroscopes, magnetometers, vibration sensors, etc.), capacitive sensors, distance sensors, temperature sensors, etc. In some embodiments, the sensor module 250 may include a VPU 250B, which is used to acquire bone conduction signals accompanying the headphone wearer's speech by picking up head bone vibrations and / or vocal cord vibrations; based on the bone conduction signals acquired by the VPU 250B, the headphones 200 can detect the user's speech state and phrasing.

[0144] When the earphone 200 is a wireless earphone, the power module 260 can provide system power to the wireless earphone, powering each module of the wireless earphone; and support the wireless earphone receiving charging input, etc. The power module 260 may include a battery 260A, a power management unit (PMU) 260B, and a charging interface 260C. The power management unit can receive external charging input; transform the electrical signal input to the charging circuit and provide it to the battery for charging; it can also transform the electrical signal provided by the battery and provide it to other modules such as the wireless communication module, audio module 240, and sensor module 250; and it can prevent the battery from overcharging, over-discharging, short-circuiting, or experiencing overcurrent. In some embodiments, the power module 260 may also include a wireless charging coil for wirelessly charging the wireless earphone. Additionally, the power management module can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). The charging interface 260C can provide a wired connection for charging or communication between the wireless earphone and the earphone case. In some embodiments, this input / output interface can be a USB interface. In other embodiments, the charging interface 260C can be an earphone electrical connector. When the wireless earphones are placed in the earphone case, the wireless earphones can establish an electrical connection with the electrical connector in the earphone case through the earphone electrical connector, thereby charging the battery in the wireless earphones. In other embodiments, after the electrical connection is established, the wireless earphones can also communicate data with the earphone case, for example, they can receive pairing commands from the earphone case.

[0145] The terminal device 100 provided in this application embodiment can run an operating system (OS). This operating system can be various operating systems used in industry, such as operating systems developed based on OpenHarmony, like HarmonyOS; or other operating systems such as Android™, iOS mobile operating systems; it can also be various open-source operating systems or their derivatives, such as Linux OS, and other embedded operating systems; or it can be a future new operating system, such as an AI operating system based on artificial intelligence. An operating system is a set of interconnected system software programs that manage and control the operation of the terminal device 100, utilize and run hardware and software resources, and provide public services to organize user interaction. In the terminal device 100, the operating system connects downwards to the physical devices at the hardware layer and provides a runtime environment for application software upwards.

[0146] An operating system typically includes a kernel layer, a middleware layer, and an application layer. The application layer includes applications, which can include system applications and third-party applications. The middleware layer includes a suite of software providing various services to application developers, or frameworks providing services such as databases, multimedia, and graphics, or capabilities such as distributed scheduling and system scaling. For example, the middleware layer may include a framework layer and / or a system service layer. The framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The system service layer includes the system's core capabilities, providing services to applications through the framework layer. The kernel layer is the layer between hardware and software. The kernel layer may include hardware drivers and the operating system kernel. In addition to providing hardware drivers, the kernel layer also supports functions such as memory management and system process management.

[0147] The electronic devices we use in our daily lives come in various types and forms, and are applied in a wide range of scenarios. Therefore, based on the different forms and functions of electronic devices, different application scenarios, and different user needs, the operating systems used in these devices may also differ. The basic functions implemented by the terminal device 100 provided in this application can be implemented using a general-purpose operating system or a dedicated operating system. To more clearly illustrate the implementation of the embodiments of this application under a specific operating system, the architecture of HarmonyOS is shown below. Those skilled in the art can deduce the implementation of the embodiments of this application under other specific operating systems, such as Android™.

[0148] For example, as shown in Figure 4, the software architecture of the terminal device 100 can be divided into several layers. In some embodiments, from bottom to top, these layers are: kernel layer, system service layer, framework layer, and application layer. Layers communicate with each other through software interfaces. System functions can be tailored, added, or combined at the subsystem granularity in different device deployment scenarios, and each subsystem can also be tailored, added, or combined at the functional granularity.

[0149] The kernel layer can include: Kernel Abstract Layer (KAL), kernel subsystem, driver subsystem, etc.

[0150] Kernel Abstraction Layer: By shielding the differences between multiple kernels, it provides basic kernel capabilities to the upper layers, including but not limited to process / thread management, memory management, file system, network management, and peripheral device management.

[0151] Kernel Subsystem: Supports the selection of a suitable OS kernel for different resource-constrained devices, including but not limited to Linux kernel, HarmonyOS kernel, LiteOS (Lite Operating System), etc.

[0152] Driver Subsystem: The driver framework is the foundation for the open system hardware ecosystem, providing unified peripheral access capabilities and a framework for driver development and management. The driver framework includes: display drivers, camera drivers, audio drivers, Bluetooth drivers, sensor drivers, etc.

[0153] The system service layer comprises the core capabilities of the system, providing services to applications through the framework layer. This layer includes, but is not limited to, the following subsystems:

[0154] The system's basic capability subsystem set provides fundamental capabilities for the operation, scheduling, and migration of distributed applications across multiple devices. This set may include distributed soft bus, distributed data management, distributed task scheduling, and Ark multi-language runtime; it may also include multi-modal input subsystem, graphics subsystem, security subsystem, and AI subsystem.

[0155] The basic software service subsystem set provides common and general software services. This set may include an event notification subsystem, a telephone service subsystem, a multimedia subsystem, and an audio framework system. The multimedia subsystem provides a simple and easy-to-understand interface, allowing applications to access and use the system's media resources. The audio framework system manages various audio sources on terminal devices 100, receives audio-related commands from the application layer, and calls the corresponding interfaces to process those commands. For example, it receives a command from a translation app, calls the speaker interface based on that command, and plays the translation result through the speaker.

[0156] Enhanced software service subsystem suite: Provides differentiated enhanced software services for different devices; the enhanced software service subsystem suite may include smart screen proprietary business subsystem, wearable proprietary business subsystem, IoT proprietary business subsystem, etc.

[0157] Hardware service subsystem set: Provides hardware services; the hardware service subsystem set may include location service subsystem, user IAM (Identity and Access Management) subsystem, wearable proprietary hardware service subsystem, biometric identification, IoT proprietary hardware service subsystem, etc.

[0158] Distributed task scheduling enables distributed service management (discovery, synchronization, registration, and invocation), supporting remote startup, remote invocation, remote connection, and migration of applications across devices.

[0159] Distributed data management enables data synchronization, data storage, data sharing, and data access across all scenarios and devices.

[0160] The distributed soft bus provides communication-related capabilities for seamless interconnection between multiple devices, including: WLAN service capabilities, Bluetooth service capabilities, soft bus, inter-process communication RPC (Remote Procedure Call), and StarFlash communication capabilities.

[0161] Ark Multilingual Runtime is a unified compilation runtime platform designed to support the joint compilation and execution of multiple programming languages ​​and multiple chip platforms.

[0162] The framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The framework layer includes: the ArkUI framework (which provides a complete infrastructure for UI development of system applications, including UI functions such as components, layouts, animations, and interactive events, as well as a real-time interface preview tool), the user application framework, and the Ability framework (an Ability is a lightweight application; the Ability framework schedules and manages the operation and lifecycle of Abilities). Different devices may have different operating systems, and therefore support different APIs.

[0163] The HarmonyOS API is a series of open capabilities provided to support HarmonyOS application development. The HarmonyOS API can be set at the framework layer or independently of the framework layer. The HarmonyOS API includes the Audio API (audio service), Push API (push service), and Account API (account service), among others.

[0164] Applications can include system apps and extended / third-party apps. System apps can include the desktop, control bar, settings, contacts, phone, camera, audio manager, etc., while extended / third-party apps can include translation apps, travel apps, etc.

[0165] The audio manager, which can be a system application, extended app, or third-party app on the terminal device 100, can be used to manage the headset 200 connected to the terminal device 100; it can also be used to sense whether the terminal device 100 has entered dialogue translation mode; it can also be used to sense the working status of dialogue translation on the terminal side: "idle", "receiving sound", "sentence segmentation detection"; and it can also be used to sense preemption events of the terminal device 100. A preemption event of the terminal device 100 can refer to: after the headset 200 enters dialogue translation mode, if the terminal device 100 enters the sound receiving state, it instructs the headset 200 to exit dialogue translation mode and prioritize the sound receiving and translation of the terminal device 100. In some embodiments, the audio manager does not need to be installed; the functions of the audio manager can also be integrated into other layers of the software system, such as the framework layer, which is not specifically limited here.

[0166] The translation app provides conversational translation functionality. It can capture the user's voice through a terminal device 100, obtain and display the translation results; it can also capture the user's voice through an earphone 200, obtain and display the translation results.

[0167] The following embodiments will be illustrated by taking a mobile phone 100 as an example of a terminal device 100.

[0168] For example, Figure 5 shows a schematic diagram of dialogue translation scenario one.

[0169] As shown in Figure 5, the translation app on mobile phone 100 enters dialogue translation mode, and mobile phone 100 instructs earphone 200 to also enter dialogue translation mode. The users performing the dialogue translation include user 1 and user 2. User 1 receives audio through mobile phone 100, and user 2 wears earphone 200 and receives audio through earphone 200. For example, mobile phone 100 receives audio and collects the speech of user 1 (language 1), and mobile phone 100 can obtain and display the translation result of that speech in language 2. Earphone 200 receives audio and collects the speech of user 2 (language 2) on the earphone side, and sends that speech to mobile phone 100; mobile phone 100 obtains and displays the translation result of that speech in language 1.

[0170] The following describes two interface implementations for Dialogue Translation Scenario 1: Interface Implementation 1 and Interface Implementation 2.

[0171] Example 1: The translation app displays a microphone control on the phone side and a microphone control on the headphone side.

[0172] For example, Figures 6A to 6F show an interactive interface for dialogue translation.

[0173] Figure 6A illustrates a user interface 11 for displaying installed applications. The user interface 11 may include: a status bar 101, a tray 102 with icons for frequently used applications, icons for other applications 103, and a page indicator 104.

[0174] The status bar 101 may include: a mobile communication signal (also known as a cellular signal) indicator, a battery status indicator, a time indicator, a Wi-Fi signal indicator, and a Bluetooth signal indicator. Other application icons 103 include an icon 103A for a translation app. These other application icons may be distributed across multiple pages, and a page indicator 104 can be used to indicate which page the user is currently viewing.

[0175] As shown in Figures 6A and 6B, after detecting an input operation (e.g., a touch operation) applied to the icon 103A of the translation app, the mobile phone 100 displays the user interface of the translation app. This user interface may include a menu bar 201, which may include various function options of the translation app, such as text translation option 201A, translation option for images captured by the camera, conversation translation option 201B, and simultaneous interpretation option. When one function option in the menu bar 201 is selected, the mobile phone 100 can display the user interface corresponding to the selected function option, allowing the user to switch between the selected options. For example, as shown in Figure 6B, when conversation translation option 201B is selected, the mobile phone 100 displays the conversation translation interface 12 corresponding to option 201B. The conversation translation interface 12 may include: controls 202 and 203, a bilingual settings bar 204, a mobile phone-side microphone control 205, and a headphone-side microphone control 206. Control 202 is used to divide the dialogue translation interface 12 into two windows vertically, with the two windows displaying in opposite directions. One window is used to display the dialog box of the mobile phone user, and the other window is used to display the dialog box of the mobile phone user. Control 203 is used to display more content related to dialogue translation.

[0176] The bilingual setting bar 204 may include: a device identifier 204A on the mobile phone side, the language of the mobile phone user 204B (e.g., Chinese), a drop-down control 204C for the mobile phone side language, a device identifier 204D on the headset side, the language of the headset user 204E (e.g., English), a drop-down control 204F for the headset user language, and a language switching control 204G. The drop-down control 204C is used to switch the language of the mobile phone user, and the drop-down control 204F is used to switch the language of the headset user. Taking the drop-down control 204C as an example, after detecting a user operation (e.g., touch operation) on the drop-down control 204C, the mobile phone 100 can display multiple language options for the user to choose from. The mobile phone 100 can update the mobile phone user's language 204B according to other languages ​​selected by the user. The language switching control 204G is used to swap the mobile phone user's language 204B and the headset user's language 204E.

[0177] The headphone-side radio control 206 is used to control the headphone 200 to start and stop radio reception.

[0178] As shown in Figures 6B and 6C, in response to a user operation (e.g., a click) on the microphone control 206, the mobile phone 100 instructs the earphone 200 to start acquiring the earphone-side audio stream. The earphone-side audio stream may include the voice of the user on the earphone side. The microphone control 206 is switched from its initial state to a microphone state to indicate that the earphone 200 is recording audio. This embodiment does not limit the specific implementation of the initial state and microphone state of the microphone control 206. The state change between the initial state and the microphone state may include changes in one or more display parameters such as the displayed text, color, brightness, and shape. For example, the earphone-side microphone control 206 may display the language set by the current user on the earphone side (e.g., English) in the initial state, and may display "Recording" in the microphone state. As shown in Figure 6C, the mobile phone 100 may also display a prompt message 207 on the right side of the dialog box display area to indicate that audio is being recorded from the user on the earphone side.

[0179] As shown in Figures 6C and 6D, in response to another user operation (e.g., a click) on the microphone control 206, the mobile phone 100 instructs the headset 200 to stop acquiring the headset-side audio stream. The mobile phone 100 obtains a translation result based on the acquired voice from the headset-side user, displays the translation result in dialog box 208, and switches the microphone control 206 back to its initial state to instruct the headset 200 to stop acquiring audio. Dialog box 208 may include the English text 208A corresponding to the headset-side user's voice and the translated Chinese text 208B.

[0180] The mobile phone side radio control 205 is used to control the mobile phone 100 to start and stop radio reception.

[0181] As shown in Figures 6D and 6E, in response to a user operation (e.g., a click) on the mobile phone-side audio control 205, the mobile phone 100 begins to acquire the mobile phone-side audio stream, which includes the voice of the mobile phone-side user. The audio control 205 is then switched from its initial state to a recording state to indicate that the mobile phone 100 is recording. Similar to the audio control 206 described above, this embodiment does not limit the specific implementation of the initial state and recording state of the audio control 205. For example, the audio control 205 may display the language set by the current mobile phone-side user (e.g., Chinese) in its initial state, and may display "Recording" in its recording state. As shown in Figure 6E, the mobile phone 100 may also display a prompt message 209 on the left side of the dialog box display area to indicate that it is recording the voice of the mobile phone-side user.

[0182] As shown in Figures 6E and 6F, in response to another user operation (e.g., a click) on the microphone control 205, the mobile phone 100 stops acquiring the audio stream from the mobile phone side, obtains a translation result based on the acquired user's speech, displays the translation result in dialog box 210, and switches the microphone control 205 back to its initial state to instruct the mobile phone 100 to stop acquiring audio. Dialog box 210 may include the Chinese text 210A corresponding to the user's speech and the translated English text 210B.

[0183] Referring to Figures 6B to 6F, and so on, alternating between the microphone control on the mobile phone side and the microphone control on the headset side can enable real-time bilingual dialogue translation between two users.

[0184] In the first embodiment of the interface, two users engaged in face-to-face dialogue translation use a mobile phone 100 and an earphone 200 respectively for audio reception; compared to both using the mobile phone 100, this operation is more convenient. However, in this embodiment, users not only need to operate the microphone control on the mobile phone side to start and stop the microphone on the mobile phone 100 to match their own speech, but also need to operate the microphone control on the earphone side to start and stop the microphone on the earphone 200 to match the other person's speech. This back-and-forth operation of two microphone controls results in low interaction efficiency and can easily cause inconvenience when the two people do not speak the same language.

[0185] Example 2: The translation app displays the microphone control on the phone side, eliminating the need to display the microphone control on the headphone side.

[0186] For example, Figures 7A to 7G show another interactive interface for dialogue translation.

[0187] Figure 7A shows another dialogue translation interface 13. Compared to dialogue translation interface 12, dialogue translation interface 13 only displays the microphone control 301 on the mobile phone side, without displaying the microphone control on the headset side. The headset 200 automatically controls the start, stop, and translation of the microphone by detecting the user's speaking state and sentence breaks. For the specific implementation of the same interface elements in dialogue translation interface 13 and dialogue translation interface 12, please refer to the relevant description of dialogue translation interface 12, which will not be repeated here.

[0188] In some embodiments, when the earphone 200 detects the user speaking, it automatically starts to collect the audio stream on the earphone side and sends the voice of the user on the earphone side to the mobile phone 100. As shown in FIG7B, when the voice of the user on the earphone side is received, the mobile phone 100 automatically displays a prompt message 302 on the left side of the dialog box display area of ​​the dialogue translation interface 13 to indicate that it is collecting the voice of the user on the earphone side.

[0189] In some embodiments, when the headset 200 detects that the user has stopped speaking, it performs a sentence segmentation detection. When a sentence segmentation is detected, for example, when the duration of the user's silence reaches 2 seconds, the headset 200 sends an end flag to the mobile phone 100. As shown in FIG7C, in response to the end flag, the mobile phone 100 translates the user's speech on the headset side and displays the translation result in a dialog box 303. The dialog box 303 may include the English text 303A corresponding to the user's speech on the headset side, the translated Chinese text 303B, and a voice playback control 303C. In one implementation, in response to a user operation (e.g., a click operation) on the voice playback control 303C, the mobile phone 100 broadcasts the translated Chinese text 303B through its speaker, and / or instructs the headset 200 to broadcast the Chinese text 303B. In one implementation, when the mobile phone 100 displays the dialog box 303, it automatically broadcasts the Chinese text 303B through its speaker, and / or automatically instructs the headset 200 to broadcast the Chinese text 303B.

[0190] In some embodiments, when the earphone 200 detects that the user has stopped speaking or detects a pause in speech, it stops acquiring the audio stream on the earphone side.

[0191] In some embodiments, in response to the aforementioned end marker, if the mobile phone 100 detects that the translation app / dialogue translation interface 13 is running in the foreground, it obtains the translation result of the user's voice on the headset side and displays the translation result in the dialog box 303; otherwise, it is not necessary to obtain the translation result. It can be understood that if the translation app / dialogue translation interface 13 is not running in the foreground, it indicates that the user is not interested in the user's voice on the headset side, therefore there is no need to translate the user's voice on the headset side to avoid unnecessary resource waste.

[0192] Similar to the mobile phone side radio control 205, the mobile phone side radio control 301 is used to control the mobile phone 100 to start and stop radio reception.

[0193] As shown in Figures 7C and 7D, in response to a user operation (e.g., a click) on the microphone control 301, the mobile phone 100 begins to acquire the audio stream from the mobile phone side, which includes the voice of the user on the mobile phone side. The microphone control 301 is switched from its initial state to a microphone state to indicate that the mobile phone 100 is currently recording audio. Similar to the aforementioned microphone control 206, this embodiment does not limit the specific implementation of the initial state and microphone state of the microphone control 301. For example, the microphone control 301 may display "Start Speaking" in the initial state and "Recording" in the microphone state. As shown in Figure 7D, the mobile phone 100 may also display a prompt message 304 on the left side of the dialog box display area to indicate that it is recording audio from the user on the mobile phone 100 side.

[0194] As shown in Figures 7D and 7E, in response to another user operation (e.g., a click) on the microphone control 301, the mobile phone 100 stops acquiring the audio stream from the mobile phone side, translates the acquired voice of the user on the mobile phone side, displays the translation result in the dialog box 305, and switches the microphone control 301 to its initial state to instruct the mobile phone 100 to stop acquiring audio. The dialog box 305 may include Chinese text 305A corresponding to the user's voice, translated English text 305B, and a voice playback control 305C. In one implementation, in response to a user operation (e.g., a click) on the voice playback control 305C, the mobile phone 100 verbally plays the translated English text 305B, and / or instructs the headset 200 to verbally play the translated English text 305B. In another implementation, when the mobile phone 100 displays the dialog box 305, the mobile phone 100 automatically verbally plays the translated English text 305B, and / or automatically instructs the headset 200 to verbally play the translated English text 305B.

[0195] In some embodiments, mobile phone audio reception takes priority over headset audio reception, and mobile phone audio reception can preempt headset audio reception. As shown in Figures 7F and 7G, mobile phone 100 displays a prompt message 302 to indicate that audio is being received from the user on the headset side. If a user operation (e.g., a click operation) is detected on the audio control 301, mobile phone 100 stops displaying the prompt message 302, pauses receiving voice from headset 200, instructs headset 200 to stop acquiring headset-side audio stream, acquires mobile phone-side audio stream, and displays a prompt message 306 to indicate that audio is being received from the user on the mobile phone side.

[0196] Referring to Figures 7A to 7G, and so on, real-time bilingual dialogue translation between two users can be achieved through the microphone control on the mobile phone side and the automatic start / stop of microphone on the earphone side.

[0197] In the dialogue translation method described in Interface Embodiment 2, the two users in a face-to-face dialogue translation use a mobile phone 100 and an earphone 200 respectively for audio recording. Compared to both users using the mobile phone 100, the operation is more convenient. For the user wearing the earphone 200, no manual operation is required; the audio recording and translation for that user are completed automatically, achieving more convenient dialogue translation and effectively improving the user experience. Furthermore, Interface Embodiment 1 only uses audio recording controls to stop audio recording and translation without performing sentence segmentation detection. In contrast, in Interface Embodiment 1, the terminal device 100 can accurately segment the user's speech on the earphone side using the end marker after sentence segmentation detection by the earphone 200, and then translate sentence by sentence, effectively improving the accuracy of the translation.

[0198] For example, regarding the dialogue translation method involved in the second embodiment of the interface, the state machine shown in Figure 8 illustrates the state changes of the headset 200 and the mobile phone 100 in the above-mentioned dialogue translation method. When the mobile phone 100 displays the dialogue translation interface of the translation APP (such as the aforementioned dialogue translation interface 13), the mobile phone 100 enters the dialogue translation mode and instructs the headset 200 to also enter the dialogue translation mode. The working states of the headset 200 in the dialogue translation mode include: idle, headset reception, and sentence segmentation detection; the working states of the mobile phone 100 in the dialogue translation mode include: idle, sentence segmentation detection, mobile phone reception, and headset reception.

[0199] In some embodiments, as shown in Figure 8, after the headset 200 enters the dialogue translation mode, it will continuously detect the user's speaking status in all three states; when the headset 200 enters the dialogue translation mode, it will first be in the "idle" state.

[0200] When the headset 200 is in "idle" state and detects that the user is speaking (i.e., VAD = True), the headset 200 switches to "headset radio" state.

[0201] When the earphone 200 is in "earphone radio" mode, it collects the audio stream from the earphone side. When it detects that the user is not speaking (i.e., VAD = False), the earphone 200 switches to "sentence segmentation detection" mode.

[0202] When the earphone 200 is in the "sentence segmentation detection" state, if it detects that the duration of the user's silence exceeds 1 second (e.g., 2 seconds), the earphone 200 switches to the "idle" state; when it detects that the user is speaking (i.e., VAD = True), the earphone 200 switches back to the "earphone radio" state.

[0203] In some embodiments, as shown in Figure 8, when the mobile phone 100 enters the dialogue translation mode, it will first be in an "idle" state.

[0204] When mobile phone 100 is in "idle" state and detects user interaction with a radio control (such as the aforementioned radio control 301) to start radio reception, mobile phone 100 switches to "mobile phone radio" state. When receiving headphone-side audio stream / voice from headphone 200 in "idle" state, mobile phone 100 can switch to "headphone radio" state. The aforementioned headphone-side audio stream includes the voice of the headphone-side user.

[0205] When mobile phone 100 is in "headphone radio" mode, if it receives an end signal from headset 200, mobile phone 100 can translate the received voice from the user on the headset side. When mobile phone 100 is in "headphone radio" mode, if it detects that the user manipulates the radio control used to start radio reception (such as the aforementioned radio control 301), mobile phone 100 switches to "mobile phone radio" mode, that is, "mobile phone radio" mode can preempt "headphone radio" mode.

[0206] When the phone 100 is in "phone radio" mode, and detects that the user clicks the radio control used to stop radio reception (such as the aforementioned radio control 301), the phone 100 can switch back to "idle" mode. In "phone radio" mode, the phone 100 collects the audio stream from the phone side and detects the user's speaking status; when it detects that the user is not speaking, the phone 100 switches to "sentence segmentation detection" mode.

[0207] When the phone is in "sentence segmentation detection" mode, it can switch back to "phone radio" mode when it detects that the user is speaking; when it detects that the user has not spoken for more than 1 second (e.g., 2 seconds), it can switch to "idle" mode.

[0208] Furthermore, when the "phone radio" state preempts the "headset radio" state, the phone 100 indicates a preemption event to the headset 200. Based on this preemption event, the headset 200 pauses its dialogue translation mode, i.e., exits the dialogue translation mode. When the phone 100 switches from the "phone radio" and "sentence segmentation check" states back to the "idle" state, the phone 100 indicates an idle event to the headset 200. Based on the idle event, the headset 200 re-enters the dialogue translation mode.

[0209] Based on the aforementioned communication system, hardware structure, interface embodiment two, and state machine of the dialogue translation mode, the method flow of the dialogue translation method provided in this application embodiment will be described in detail below.

[0210] For example, Figures 9A and 9B illustrate the method flow of the dialogue translation method, which includes some or all of steps S101 to S120; Figure 9C illustrates a device interaction flow of the above-described dialogue translation method. The specific implementation of steps S101 to S120 is described below.

[0211] S101, Mobile Phone 100 receives user input when entering the dialogue translation interface.

[0212] S102. In response to the above user operation, mobile phone 100 displays a dialogue translation interface.

[0213] For example, referring to the relevant descriptions in Figures 6A and 7A, after detecting that the user clicks the translation app icon 103A, the mobile phone 100 displays the dialogue translation interface 13, that is, enters the dialogue translation mode of the translation app. In some embodiments, after launching the translation app, the mobile phone 100 may first display the user interface corresponding to the text translation option 201A; when the user selects the dialogue translation option 201B, the mobile phone 100 then displays the dialogue translation interface 13. In this embodiment, the dialogue translation interface 13 of the translation app on the mobile phone 100 can also be accessed in other ways, and the user operation for entering the dialogue translation mode of the translation app is not specifically limited here.

[0214] S103, Mobile Phone 100 indicates that Earphone 200 has entered dialogue translation mode.

[0215] In some implementations, as shown in Figure 9B, mobile phone 100 includes an audio manager, and step S103 specifically includes steps S103A and S103B. S103A: The translation app on mobile phone 100 instructs the audio manager to enter dialogue translation mode. S103B: The audio manager on mobile phone 100 instructs the headset 200 to enter dialogue translation mode.

[0216] S104, the headset 200 enters the dialogue translation mode; in the dialogue translation mode, the headset 200 executes steps S105 to S108.

[0217] S105, when the earphone 200 detects that the user is speaking, it collects the audio stream on the earphone side, which includes the user's voice 1 on the earphone side.

[0218] In some embodiments, referring to step (1) of FIG9C, in dialogue translation mode, the earphone 200 can use the VPU to acquire the bone conduction signal accompanying the earphone wearer's speech; based on the bone conduction signal acquired by the VPU, the earphone 200 can use the VAD module to determine whether the earphone wearer is speaking. In this way, by using the bone conduction signal acquired by the VPU, it is more accurate to detect whether the wearer of the earphone 200 (i.e., the aforementioned earphone-side user) is speaking.

[0219] In some embodiments, in dialogue translation mode, the headset 200 uses a microphone to capture the audio stream on the headset side; the VAD module of the headset 200 can identify whether the audio stream on the headset side contains user speech, and thus determine whether the user is speaking. This application embodiment does not specifically limit the implementation method of the headset 200 detecting whether the user is speaking.

[0220] In some embodiments, the earphone 200 acquires an earphone-side audio stream, including: audio acquired via a microphone, and / or bone conduction signals acquired via a VPU. It should be noted that when a user speaks, the bone vibration signals acquired by the VPU on the earphone 200 can also be considered as a form of user voice data. For example, referring to step (2) of Figure 9C, the earphone 200 acquires an earphone-side audio stream via a microphone and a VPU.

[0221] S106, Earphone 200 sends voice message 1 to mobile phone 100.

[0222] In dialogue translation mode, when the headset 200 detects that the user is speaking, it collects the user's voice 1 in real time and sends it to the translation app on the mobile phone 100. In some implementations, as shown in Figure 9B, step S106 specifically includes steps S106A and S106B. S106A: The headset 200 sends voice 1 to the audio manager of the mobile phone 100. S106B: The audio manager of the mobile phone 100 sends voice 1 to the translation app.

[0223] In some embodiments, the earphone 200 sends voice 1 to the mobile phone 100, including: the earphone 200 first preprocesses the earphone-side audio stream acquired by the earphone 200 to remove non-speech segments such as environmental noise in the earphone-side audio stream, obtains voice 1 in the earphone-side audio stream, and then sends voice 1 to the terminal device 100; or, the earphone 200 obtains voice 1 in the earphone-side audio stream and sends the text corresponding to voice 1 to the terminal device 100; or, the earphone 200 sends the audio stream acquired by the earphone 200 to the mobile phone 100, wherein the audio stream carries the user's voice 1, and the mobile phone 100 can preprocess the earphone-side audio stream to obtain voice 1.

[0224] This application does not specifically limit the communication method between the headset 200 and the mobile phone 100. In some embodiments, the headset 200 and the mobile phone 100 can communicate via near-field communication technology, such as Bluetooth communication.

[0225] For example, referring to steps (3) to (5) of Figure 9C, the earphone 200 sends voice 1 to the mobile phone 100, including: (3) The earphone 200 uses the preprocessing module to preprocess the earphone-side audio stream acquired by the earphone 200 to obtain voice 1 in the earphone-side audio stream; (4) The earphone 200 uses the encoding / decoding module to encode voice 1 to obtain the encoded voice 1; (5) The earphone 200 sends the encoded voice 1 to the mobile phone 100 through the Bluetooth module. Referring to steps (6) and (7) of Figure 9C, the translation APP of the mobile phone 100 receives voice 1 sent by the earphone 100, including: (6) The Bluetooth module of the mobile phone 100 sends the encoded voice 1 to the audio manager; (7) The decoding module of the audio manager decodes the encoded voice 1 to obtain voice 1 and sends voice 1 to the translation APP.

[0226] S107, when the earphone 200 detects that the user has stopped speaking, it stops collecting the audio stream on the earphone side and performs sentence segmentation detection.

[0227] How the headset 200 detects that the user is not speaking can be found in the description of step S105, which will not be repeated here. When the headset 200 detects that the user has stopped speaking, it stops collecting the audio stream on the headset side, which avoids collecting and processing audio streams without speech, thereby saving device power consumption.

[0228] When the S108 and earphone 200 detect a sentence break, they send an end marker to the translation app on the mobile phone 100.

[0229] In some implementations, as shown in Figure 9B, step S108 specifically includes steps S108A and S108B. S108A: When the earphone 200 detects a sentence break, it sends an end marker to the audio manager of the mobile phone 100. S108B: The audio manager of the mobile phone 100 sends an end marker to the translation app.

[0230] In some embodiments, the earphone 200 starts timing when it detects that the user has stopped speaking. When the timing of the user stopping speaking reaches a duration of 1 (e.g., 2 seconds), it determines that a sentence break has been detected and sends an end marker to the translation APP of the mobile phone 100 to instruct the translation APP to translate the sent voice 1.

[0231] In other embodiments, the earphone 200 stops acquiring the earphone-side audio stream only when detecting sentence breaks.

[0232] S109. In response to the above-mentioned end marker, the mobile phone 100 acquires and displays the translation result of the user's voice 1 on the headset side.

[0233] It is understandable that by using end markers, each sentence in the user's voice on the headset can be segmented and translated sentence by sentence; this can improve the accuracy of sentence segmentation, and thus improve the accuracy of translation.

[0234] In some embodiments, referring to step (8) of FIG9C, the translation app of mobile phone 100 can obtain the translation result of the user's voice 1 on the headset side through the translation server 300 in the cloud. In one implementation, the translation app of mobile phone 100 sends voice 1 to the translation server 300 in the cloud; the translation server 300 obtains the translation result of voice 1 and returns it to the translation app of mobile phone 100. In another implementation, the translation app of mobile phone 100 can convert voice 1 into text and send it to the translation server 300 in the cloud; the translation server 300 obtains the translation result of the text (i.e., the translation result of voice 1) and returns it to the translation app of mobile phone 100.

[0235] In some embodiments, the mobile phone 100 has translation capabilities and can translate speech 1 locally to obtain the translation result of speech 1.

[0236] In some embodiments, before step S109, the mobile phone 100 can set the languages ​​used by both parties in the dialogue translation. For example, referring to the aforementioned bilingual setting bar 204, the user can manually set the languages ​​used by both parties in the translation app on the mobile phone 100, or the translation app can set the languages ​​used by both parties in the dialogue translation by default, namely, language 1 (e.g., Chinese) for the mobile phone user and language 2 (e.g., English) for the headset user. In step S109, according to the languages ​​set by the mobile phone 100, the mobile phone 100 / translation server 300 can obtain the text of language 1 corresponding to the voice 1 of the headset user, and translate the text of language 1 into the text of language 2 (i.e., the translation result of voice 1).

[0237] In some embodiments, the mobile phone 100 sets two languages ​​(e.g., language 1 and language 2) used by both parties in the dialogue translation. The mobile phone 100 / translation server 300 can automatically identify the language used by the voice 1 as one of the two languages ​​(e.g., language 1) and translate the voice 1 into text in the other language (e.g., language 2).

[0238] In some embodiments, the mobile phone 100 can display the translation result of voice 1 on the screen via text; for example, referring to FIG7C, in the dialogue translation interface 13, the translation APP can display the Chinese text 303B translated by the user's voice on the headset side. The mobile phone 100 can also broadcast the translation result of voice 1 via its speaker; for example, referring to FIG7C, the voice playback control 303C is used to trigger the voice broadcast of the Chinese text 303B. The mobile phone 100 can also send the translation result of voice 1 / the corresponding voice to the headset 200, instructing the headset 200 to broadcast the above translation result.

[0239] For example, referring to steps (9) to (10) of Figure 9C, the mobile phone 100 can also broadcast the translation result of voice 1 through the speaker of the mobile phone 100, including: (9) the translation APP sends the voice corresponding to the translation result of voice 1 to the audio frame; (10) the audio frame sends the voice corresponding to the translation result to the speaker, and the speaker plays the above voice, that is, the speaker broadcasts the translation result of voice 1.

[0240] S110 and Mobile Phone 100 receive user commands to start the mobile phone radio.

[0241] S111. In response to the above user operation, mobile phone 100 collects the mobile phone side audio stream, which includes voice 2.

[0242] In some embodiments, the translation app on mobile phone 100 displays a microphone control, which is used to trigger mobile phone 100 to start recording audio. For example, referring to FIG7C, in response to a touch operation on the microphone control 301, mobile phone 100 starts recording audio, i.e., it acquires the audio stream from the mobile phone side. Not limited to the microphone control, this embodiment of the application can also trigger mobile phone 100 to record audio in other ways, which are not specifically limited here; for example, if mobile phone 100 does not display a microphone control, the user can trigger mobile phone 100 to start recording audio via voice command.

[0243] For example, referring to steps (11) and (12) of Figure 9C, the mobile phone 100 collects the audio stream on the mobile phone side, and the translation APP obtains the voice 2, including: (11) the mobile phone 100 collects the audio stream on the mobile phone side through the microphone; (12) the mobile phone 100 uses the preprocessing module to preprocess the audio stream on the mobile phone side to remove non-speech segments such as environmental noise in the audio stream on the mobile phone side, obtains the voice 2 in the audio stream on the mobile phone side, and sends it to the translation APP.

[0244] S112. In response to the above user operation, the mobile phone 100 indicates a mobile phone preemption event to the headset 200.

[0245] In some implementations, as shown in Figure 9B, step S112 specifically includes steps S112A and S112B. S112A: In response to the aforementioned user operation, the translation app on mobile phone 100 indicates a phone preemption event to the audio manager. S112B: The audio manager on mobile phone 100 indicates a phone preemption event to the headset 200.

[0246] S113, In response to the phone preemption incident, the headset 200 exits the dialogue translation mode.

[0247] In some embodiments, when the mobile phone 100 starts receiving audio, the mobile phone 100 also sends a message to the headset 200 to indicate a preemption event, that is, the mobile phone 100 has preempted the right to speak in the dialogue translation. Based on the preemption event, the headset 200 suspends the working state of the dialogue translation mode of the mobile phone 100, that is, exits the dialogue translation mode.

[0248] In some embodiments, when the mobile phone 100 is receiving audio, the mobile phone takes over the conversation translation function. The mobile phone 100 no longer receives and processes the voice from the headset 200 for conversation translation, and the headset 200 also stops acquiring the audio stream for conversation translation.

[0249] After step S111, that is, after the mobile phone 100 starts collecting the audio stream from the mobile phone side, the mobile phone 100 can execute steps S114 to S116 or steps S117 to S119.

[0250] S114. When the mobile phone detects that the user is not speaking, it performs sentence segmentation detection.

[0251] When S115 and mobile phone 100 detect a sentence break, they stop collecting the audio stream from the mobile phone side and obtain and display the translation result of speech 2.

[0252] In some embodiments, when the mobile phone 100 receives audio, if no user voice is detected in the audio stream on the mobile phone side, the mobile phone 100 determines that the user has not spoken, and the mobile phone 100 begins sentence segmentation detection. In some embodiments, the mobile phone 100 begins sentence segmentation detection by: starting a timer when the user is detected not speaking; and determining that a sentence segmentation has been detected when the timer for the period of user silence reaches one duration (e.g., 2 seconds). In this way, the terminal device 100 can accurately segment the user's voice on the mobile phone side and then translate it sentence by sentence, which can effectively improve the accuracy of translation.

[0253] In this embodiment, when a user's sentence segmentation is detected, the mobile phone 100 translates the collected voice 2 of the user on the mobile phone side and obtains the translation result of voice 2. In one implementation, the mobile phone 100 can obtain the text of language 2 corresponding to the user's voice 2 on the mobile phone side, and translate the text of language 2 into text of language 1 (i.e., the translation result of voice 2). For the specific implementation of the mobile phone 100 obtaining the translation result of voice 2, please refer to the relevant description of the mobile phone 100 obtaining the translation result of voice 1 in step S109, which will not be repeated here.

[0254] For example, referring to step (13) of Figure 9C, the mobile phone 100 obtains the translation result of voice 2, including: the translation APP of the mobile phone 100 can obtain the translation result of the user's voice 2 on the mobile phone side through the translation server 300 in the cloud.

[0255] In some embodiments, the mobile phone 100 can display the translation result of voice 2 on the screen via text; for example, referring to FIG7E, in the dialogue translation interface 13, the translation APP can display the English text 305B translated by the user's voice on the mobile phone side. The mobile phone 100 can also broadcast the translation result via its speaker; for example, referring to FIG7E, the voice playback control 305C is used to trigger the voice broadcast of the English text 305B. The mobile phone 100 can also send the translation result of voice 2 / the corresponding voice to the headset 200, instructing the headset 200 to broadcast the above translation result.

[0256] For example, referring to steps (14) to (18) of Figure 9C, the mobile phone 100 can also send the voice corresponding to the translation result to the headset 200, instructing the headset 200 to broadcast the above translation result, including: (14) The translation APP sends the voice 3 corresponding to the translation result of voice 2 to the audio frame; (15) The audio frame sends the voice 3 corresponding to the above translation result to the Bluetooth module; (16) The Bluetooth module of the mobile phone 100 encodes the voice 3 using the encoding module and sends the encoded voice 3 to the Bluetooth module of the headset 200; (17) The Bluetooth module of the headset 200 sends the encoded voice 3 to the encoding / decoding module; (18) The encoding / decoding module decodes the encoded voice 3, obtains the voice 3, and sends it to the speaker of the headset 200, and the speaker of the headset 200 plays the voice 3, that is, broadcasts the translation result of voice 2.

[0257] S116. In response to the detection of a sentence break, the mobile phone 100 indicates a mobile phone idle event to the headset 200 and executes S120.

[0258] In some implementations, as shown in Figure 9B, step S116 specifically includes steps S116A and S116B. S116A: In response to detecting a sentence break, the translation app of mobile phone 100 indicates a phone idle event to the audio manager. S116B: The audio manager of mobile phone 100 indicates a phone idle event to the headset 200.

[0259] In some embodiments, in response to detecting a sentence break, the mobile phone 100 also sends a message to the headset 200 to indicate a mobile phone idle event, i.e., the mobile phone 100 relinquishes the right to speak in the dialogue translation. Based on the mobile phone idle event, the headset 200 re-enters the dialogue translation mode.

[0260] S117, Mobile phone 100 receives user commands to stop mobile phone radio.

[0261] S118. In response to the above user operation, mobile phone 100 stops collecting audio streams from the mobile phone side and obtains and displays the translation result of voice 2.

[0262] In some embodiments, the translation app on the mobile phone 100 displays a microphone control. When the mobile phone 100 is capturing audio streams from the mobile phone side, the microphone control can also be used to trigger the mobile phone 100 to stop capturing audio. For example, referring to FIG7D, in response to a touch operation on the microphone control 301, the mobile phone 100 stops capturing audio streams from the mobile phone side. Not limited to the microphone control, this embodiment of the application can also trigger the mobile phone 100 to stop capturing audio in other ways, and no specific limitation is made here.

[0263] S119. In response to the above user operation, the mobile phone 100 indicates a mobile phone idle event to the headset 200 and executes S120.

[0264] In some implementations, as shown in Figure 9B, step S119 specifically includes steps S119A and S119B. S119A: In response to the aforementioned user operation, the translation app on mobile phone 100 indicates a phone idle event to the audio manager. S119B: The audio manager on mobile phone 100 indicates a phone idle event to the headset 200.

[0265] In some embodiments, in response to a user operation to stop the phone from recording, the phone 100 also sends a message to the headset 200 to indicate a phone idle event, meaning that the phone 100 is currently relinquishing the right to speak in the dialogue translation. Based on the phone idle event, the headset 200 re-enters the dialogue translation mode.

[0266] S120, in response to a phone idle event, the headset 200 enters conversation translation mode.

[0267] It is understandable that after the headset 200 enters the dialogue translation mode, it can execute S105 to S108 again.

[0268] In some embodiments, referring to Figure 8, the "phone radio" state of mobile phone 100 cannot preempt the "headset radio" state. After receiving the user's operation to start phone radio in step S110, mobile phone 100 does not need to execute S112, S113, S116, S119, and S120. That is, mobile phone 100 does not need to indicate the phone preemption event or the phone idle event, and when mobile phone 100 is receiving radio, headset 200 does not need to exit the dialogue translation mode. After step S110, headset 200 can continue to execute S105 to S108, and mobile phone 100 can normally receive the voice 1 or the end marker of the sentence segment from headset 200; if mobile phone 100 receives the end marker sent by headset 200, it translates the voice 1 of the user on the headset side before the end marker.

[0269] For example, Figure 10 shows a schematic diagram of dialogue translation scenario two.

[0270] As shown in Figure 10, the translation app on mobile phone 100 enters dialogue translation mode, and mobile phone 100 instructs the two earpieces of earphone 200 to enter dialogue translation mode respectively. The users performing the dialogue translation include User 1 and User 2. User 1 wears the right earpiece of earphone 200 and receives audio through the right earpiece, while User 2 wears the left earpiece of earphone 200 and receives audio through the left earpiece. For example, the right earpiece receives audio, capturing the speech of the right earpiece user (i.e., User 1) in language 1, and sends this speech to mobile phone 100; mobile phone 100 can obtain and display the translation result of the speech in language 2. The left earpiece receives audio, capturing the speech of the left earpiece user (i.e., User 2) in language 2, and sends this speech to mobile phone 100; mobile phone 100 obtains and displays the translation result of the speech in language 1.

[0271] The following describes the third example of the interface involved in the dialogue translation scenario.

[0272] Example 3: The translation app does not need to display the microphone controls on the phone side or the headphone side.

[0273] For example, Figures 11A to 11E show an interactive interface for dialogue translation in Interface Embodiment 3.

[0274] Figure 11A shows another dialogue translation interface 14. Compared to the dialogue translation interface 12, the dialogue translation interface 14 does not require displaying the audio controls corresponding to the two users of the dialogue translation. The right earpiece automatically controls the right earpiece to start and stop recording and translate by detecting the right earpiece user's speaking state and sentence breaks; the left earpiece automatically controls the left earpiece to start and stop recording and translate by detecting the left earpiece user's speaking state and sentence breaks.

[0275] The dialogue translation interface 14 includes a bilingual settings bar 401, which may include: a device identifier 401A for the right earpiece, a language 401B for the right earpiece user (e.g., Chinese), a drop-down control 401C for the right earpiece language, a device identifier 401D for the left earpiece, a language 401E for the left earpiece user (e.g., English), a drop-down control 401F for the left earpiece language, and a language switching control 401G. The drop-down control 401C is used to switch the language of the right earpiece user, and the drop-down control 401F is used to switch the language of the left earpiece user. The language switching control 401G is used to swap the languages ​​401B and 401E of the right earpiece user. For the specific implementation of the same interface elements in dialogue translation interface 14 and dialogue translation interface 12, please refer to the relevant description of dialogue translation interface 12; it will not be repeated here.

[0276] In some embodiments, in the dialogue translation mode, when the left earpiece detects the user speaking, it automatically starts to collect the audio stream of the left earpiece and sends the voice of the user in the left earpiece to the mobile phone 100; as shown in FIG11B, when the voice of the user in the left earpiece is received, the mobile phone 100 automatically displays a prompt message 402 on the left side of the dialog box display area of ​​the dialogue translation interface 14 to indicate that it is recording the voice of the user in the left earpiece.

[0277] In some embodiments, in dialogue translation mode, when the left earpiece detects that the user has stopped speaking, it performs sentence segmentation detection; when a sentence segmentation is detected, for example, when the duration of the user's silence reaches 2 seconds, the left earpiece sends an end marker to the mobile phone 100; as shown in FIG11C, in response to the above-mentioned end marker, the mobile phone 100 translates the speech of the user in the left earpiece and displays the translation result in dialog box 403. Dialog box 403 may include the English text 403A corresponding to the speech of the user in the left earpiece, the translated Chinese text 403B, and the voice playback control 403C. In one implementation, in response to a user operation (e.g., a click operation) on the voice playback control 403C, the mobile phone 100 broadcasts the translated Chinese text 403B through the speaker of the mobile phone 100, and / or instructs the left earpiece to broadcast the Chinese text 403B. In one implementation, when the mobile phone 100 displays dialog box 403, it automatically broadcasts the Chinese text 403B through the speaker of the mobile phone 100, and / or automatically instructs the left earpiece to broadcast the Chinese text 403B.

[0278] In some embodiments, during dialogue translation mode, the left earpiece stops capturing the audio stream when it detects that the user has stopped speaking or that a sentence has been interrupted.

[0279] In some embodiments, in response to the end marker of the left earpiece, if the mobile phone 100 detects that the translation app / dialogue translation interface 13 is running in the foreground, it obtains the translation result of the user's voice in the left earpiece and displays the translation result in the dialog box 403; otherwise, it is not necessary to obtain the above translation result. It can be understood that if the translation app / dialogue translation interface 13 is not running in the foreground, it indicates that the user is not interested in the voice of the dialogue translation, so there is no need to translate the voice of the dialogue translation to avoid unnecessary waste of resources.

[0280] In some embodiments, in the dialogue translation mode, when the right earpiece detects the user speaking, it automatically starts to collect the audio stream of the right earpiece and sends the voice of the right earpiece user to the mobile phone 100; as shown in Figure 11D, when the voice of the right earpiece user is received, the mobile phone 100 automatically displays a prompt message 404 on the right side of the dialog box display area of ​​the dialogue translation interface 14 to indicate that it is recording the voice of the right earpiece user.

[0281] In some embodiments, in dialogue translation mode, when the right earpiece detects that the user has stopped speaking, it performs sentence segmentation detection; when a sentence segmentation is detected, for example, when the duration of the user's silence reaches 2 seconds, the right earpiece sends an end marker to the mobile phone 100; as shown in FIG11E, in response to the above-mentioned end marker, the mobile phone 100 translates the speech of the user in the right earpiece and displays the translation result in dialog box 405. Dialog box 405 may include Chinese text 405A corresponding to the speech of the user in the right earpiece, translated English text 405B, and voice playback control 405C. In one implementation, in response to a user operation (e.g., a click operation) on the voice playback control 405C, the mobile phone 100 broadcasts the translated English text 405B through its speaker, and / or instructs the right earpiece to broadcast the translated English text 405B. In one implementation, when the mobile phone 100 displays dialog box 405, it automatically broadcasts the English text 405B through its speaker, and / or automatically instructs the right earpiece to broadcast the translated English text 405B.

[0282] In some embodiments, during dialogue translation mode, the right earpiece stops capturing the right earpiece audio stream when it detects that the user has stopped speaking or that a sentence has been interrupted.

[0283] In some embodiments, similarly, in response to the end marker of the right earpiece, if the mobile phone 100 detects that the translation APP / translation interface 13 of the translation APP is running in the foreground, it obtains the translation result of the right earpiece user's voice and displays the translation result in the dialog box 405; otherwise, it is not necessary to obtain the above translation result.

[0284] In some embodiments, one of the left and right earpieces has priority in dialogue translation. Taking the right earpiece as an example, the right earpiece with priority in dialogue translation can preempt the left earpiece's speaking time. In one implementation, when the right earpiece detects the user speaking, it indicates a right earpiece preemption event to the left earpiece; based on the right earpiece preemption event, it exits the dialogue translation mode; when the right earpiece detects the user pausing, it indicates a right earpiece idle event to the left earpiece; based on the right earpiece idle event, the left earpiece re-enters the dialogue translation mode.

[0285] Referring to Figures 11A to 11E, and so on, real-time bilingual dialogue translation between two users can be achieved by automatically starting and stopping the two earpieces to receive and translate audio.

[0286] The specific implementation of Interface Embodiment 3 can be found in the descriptions of Interface Embodiment 1 shown in Figures 7A to 7E, the state machine of Interface Embodiment 1 shown in Figure 8, and the dialogue translation method shown in Figures 9A to 9C, and will not be repeated here. For example, taking the right earphone as an example, the specific implementation of the right earphone body detecting user speech and phrasing can be found in the aforementioned implementation of earphone 200 detecting user speech and phrasing; the interaction process between the right earphone body and the mobile phone 100 can be found in the aforementioned interaction process between earphone 200 and mobile phone 100.

[0287] In the dialogue translation method described in Interface Embodiment 3, the two users in a face-to-face dialogue translation use the left and right earpieces of mobile phone 100 respectively for audio reception. Compared to both users using mobile phone 100 for audio reception, the operation is more convenient. For users wearing the left and right earpieces, no manual operation is required; the audio reception and translation for both users are automatically completed, achieving more convenient dialogue translation and effectively improving the user experience. Furthermore, Interface Embodiment 1 only uses audio reception controls to stop audio reception and translation without performing sentence segmentation detection. In Interface Embodiment 3, however, through the end marker after sentence segmentation detection for each earpiece, the terminal device 100 can accurately segment the speech of the user wearing each earpiece and then translate sentence by sentence, effectively improving the accuracy of the translation.

[0288] Based on the foregoing embodiments, as shown in FIG12, this application provides a dialogue translation method applied to a dialogue translation system. The dialogue translation system includes a terminal device and an earpiece, and the terminal device and the earpiece establish a communication connection. The method includes steps S201 to S203.

[0289] S201. When the earphone detects that the user is speaking on the earphone side, it collects the first voice.

[0290] S202, The headset sends the first voice message to the terminal device.

[0291] S203. When the earphone detects that the user has paused the sentence on the earphone side, it sends a first identifier to the terminal device. Based on the first identifier, the terminal device translates the first speech and obtains the translation result of the first speech.

[0292] The earphone can be the aforementioned earphone 200; the terminal device can be the aforementioned terminal device 100 or mobile phone 100. The first voice can be the aforementioned voice 1; the first identifier can be the aforementioned end identifier.

[0293] In one implementation, the step of translating the first speech based on the first identifier by the terminal device includes: translating the first speech by the first application running in the foreground on the terminal device, based on the first identifier; or translating the first speech by the first application running in the foreground on the terminal device's first interface, where the first interface is used for dialogue translation. The first application can be the aforementioned translation app, and the first interface can be the aforementioned dialogue translation interface, such as the aforementioned dialogue translation interface 13 or dialogue translation interface 14.

[0294] In one implementation, when acquiring the first voice, the method further includes: when the earpiece detects that the user on the earpiece side has stopped speaking, performing sentence segmentation detection; when the earpiece detects that the user on the earpiece side has stopped speaking for a duration reaching a first duration, determining that a sentence segmentation has been detected. The first duration can be the aforementioned duration of 1, for example, 2 seconds. This application embodiment does not specifically limit the first duration. Optionally, the user or terminal device can set the first duration according to actual needs.

[0295] In one implementation, when acquiring the first speech, the method further includes: stopping the acquisition of the first speech when the earphone detects that the user on the earphone side has stopped speaking; or, stopping the acquisition of the first speech when the earphone detects that the user on the earphone side has paused in speech.

[0296] In one implementation, before acquiring the first voice signal when the earpiece detects the user speaking, the method further includes: a terminal device receiving a first user operation; in response to the first user operation, the terminal device displaying a first interface and sending a first message to the earpiece; the first interface being used for dialogue translation; based on the first message, the earpiece entering a dialogue translation mode; acquiring the first voice signal when the earpiece detects the user speaking on its side, including: in the dialogue translation mode, acquiring the first voice signal when the earpiece detects the user speaking on its side; and sending a first identifier to the terminal device when the earpiece detects the user pausing at the earpiece, including: in the dialogue translation mode, sending a first identifier to the terminal device when the earpiece detects the user pausing at the earpiece. The first user operation can be the aforementioned user operation to enter the dialogue translation interface; for example, the first message can be the message sent from the mobile phone 100 to the earpiece 200 in step S103.

[0297] In one implementation, after the headset enters the dialogue translation mode, the method further includes: the terminal device receiving a second user operation; in response to the second user operation, the terminal device acquiring second voice and sending a second message to the headset; and based on the second message, the headset exiting the dialogue translation mode. The second user operation can be the aforementioned user operation to start the phone's microphone; for example, the second message can be the message sent from the phone 100 to the headset 200 in step S112, and the second message can be used to indicate a phone preemption event.

[0298] In one implementation, after the headset exits the dialogue translation mode, the method further includes: the terminal device receiving a third user operation; the terminal device responding to the third user operation by stopping the acquisition of second voice and sending a third message to the headset; and the headset entering the dialogue translation mode based on the third message. The third user operation can be the aforementioned user operation to stop the phone from receiving audio. For example, the third message can be the message sent from the phone 100 to the headset 200 in step S119, and the third message can be used to indicate a phone idle event.

[0299] In one implementation, after the headset exits the dialogue translation mode, the method further includes: when the terminal device detects that the user on the terminal side has paused a sentence, it sends a third message to the headset; based on the third message, the headset enters the dialogue translation mode. For example, the third message can be the message sent by the mobile phone 100 to the headset 200 in the aforementioned step S116, and the third message can be used to indicate a mobile phone idle event.

[0300] In one implementation, after acquiring the second voice, the method further includes: when the terminal device detects that the terminal-side user has paused phrasing, it stops acquiring the second voice. The second voice can be the aforementioned voice 2 from the mobile phone-side user.

[0301] The various embodiments of this application can be combined arbitrarily to achieve different technical effects.

[0302] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0303] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.

[0304] In summary, the above description is merely an embodiment of the technical solution of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made based on the disclosure of this application should be included within the scope of protection of this application.

Claims

1. A dialogue translation method, applied to a dialogue translation system, characterized in that, The dialogue translation system includes a terminal device and an earpiece, wherein the terminal device and the earpiece establish a communication connection, and the method includes: When the earphone detects that the user is speaking on the earphone side, it collects the first voice message; The earphones send the first voice message to the terminal device; When the earphone detects a sentence pause by the user on the earphone side, it sends a first identifier to the terminal device. Based on the first identifier, the terminal device translates the first speech and obtains the translation result of the first speech.

2. The method according to claim 1, characterized in that, The process of translating the first speech by the terminal device based on the first identifier includes: Based on the first identifier, when the first application is running in the foreground of the terminal device, the first application translates the first speech. Alternatively, based on the first identifier, when the first interface of the first application on the terminal device is running in the foreground, the first application translates the first speech, and the first interface is used for dialogue translation.

3. The method according to claim 1 or 2, characterized in that, When acquiring the first speech, the method further includes: When the earphone detects that the user on the earphone side has stopped speaking, it performs sentence segmentation detection; When the earphone detects that the duration of the user's silence on the earphone side has reached a first duration, it determines that the user on the earphone side has stopped speaking.

4. The method according to any one of claims 1 to 3, characterized in that, When acquiring the first speech, the method further includes: When the earphone detects that the user on the earphone side has stopped speaking, it stops collecting the first voice. Alternatively, the earphone may stop collecting the first speech when it detects that the user on the earphone side has paused their speech.

5. The method according to any one of claims 1 to 4, characterized in that, Before collecting the first voice signal when the earpiece detects the user speaking, the method further includes: The terminal device receives the first user's operation; In response to the first user operation, the terminal device displays a first interface and sends a first message to the headset; the first interface is used for dialogue translation. Based on the first message, the headset enters dialogue translation mode; When the earphone detects that the user is speaking on the earphone side, it collects a first voice recording, including: In the dialogue translation mode, the earphone detects that the user on the earphone side is speaking and collects the first voice; When the earphone detects a sentence segmentation by the user on the earphone side, it sends a first identifier to the terminal device, including: In the dialogue translation mode, when the earpiece detects that the user has paused the sentence on the earpiece side, it sends the first identifier to the terminal device.

6. The method according to claim 5, characterized in that, After the headset enters the dialogue translation mode, the method further includes: The terminal device receives a second user operation; In response to a second user operation, the terminal device acquires a second voice message and sends a second message to the headset. Based on the second message, the headset exits the dialogue translation mode.

7. The method according to claim 6, characterized in that, After the headset exits the dialogue translation mode, the method further includes: The terminal device receives operations from a third user. In response to the third user's operation, the terminal device stops collecting the second voice and sends a third message to the earpiece. Based on the third message, the headset enters dialogue translation mode.

8. The method according to claim 6, characterized in that, After the headset exits the dialogue translation mode, the method further includes: When the terminal device detects that the user on the terminal side has paused a sentence, it sends a third message to the earpiece; Based on the third message, the headset enters dialogue translation mode.

9. The method according to claim 8, characterized in that, After acquiring the second voice, the method further includes: When the terminal device detects that the user on the terminal side has paused a sentence, it stops collecting the second voice.

10. A dialogue translation method applied to headphones, characterized in that, The method includes: When the earphone detects that the user is speaking on the earphone side, it collects the first voice message; The earphones send the first voice message to the terminal device; When the earphone detects that the user has paused their speech on the earphone side, it sends a first identifier to the terminal device. The first identifier is used to instruct the terminal device to translate the first speech.

11. The method according to claim 10, characterized in that, When acquiring the first speech, the method further includes: When the earphone detects that the user on the earphone side has stopped speaking, it performs sentence segmentation detection; When the earphone detects that the duration of the user's silence on the earphone side has reached a first duration, it determines that the user has stopped speaking.

12. The method according to claim 10 or 11, characterized in that, When acquiring the first speech, the method further includes: When the earphone detects that the user on the earphone side has stopped speaking, it stops collecting the first voice. Alternatively, the earphone may stop collecting the first speech when it detects that the user on the earphone side has paused their speech.

13. The method according to any one of claims 10 to 12, characterized in that, Before collecting the first voice recording when the earphone detects that the user is speaking on the earphone side, the method further includes: The earpiece receives a first message sent by the terminal device; Based on the first message, the headset enters dialogue translation mode; When the earphone detects that the user is speaking on the earphone side, it collects the first voice, including: In the dialogue translation mode, the earphone detects that the user on the earphone side is speaking and collects the first voice; When the earphone detects a user's pause in speech, it sends a first identifier to the terminal device, including: In the dialogue translation mode, when the earpiece detects that the user has paused the sentence on the earpiece side, it sends the first identifier to the terminal device.

14. The method according to claim 13, characterized in that, After the headset enters the dialogue translation mode, the method further includes: The earphone receives a second message sent by the terminal device; the second message is sent by the terminal device when it detects a user operation that triggers the terminal device to start recording. Based on the second message, the headset exits the dialogue translation mode.

15. The method according to claim 14, characterized in that, After the headset exits the dialogue translation mode, the method further includes: The earphone receives a third message sent by the terminal device; the third message is sent when the terminal device detects a user operation that triggers the terminal device to stop receiving audio, or the third message is sent when the terminal device detects a user punctuation error on the terminal side. Based on the third message, the headset enters dialogue translation mode.

16. A dialogue translation method, applied to a terminal device, characterized in that, The method includes: The terminal device receives the first voice message sent by the earpiece; The terminal device receives a first identifier sent by the earpiece, the first identifier indicating the sentence segmentation by the user on the earpiece side; Based on the first identifier, the terminal device translates the first speech and obtains the translation result of the first speech.

17. The method according to claim 16, characterized in that, The process of translating the first speech by the terminal device based on the first identifier includes: Based on the first identifier, when the first application is running in the foreground of the terminal device, the first application translates the first speech. Alternatively, based on the first identifier, when the first interface of the first application on the terminal device is running in the foreground, the first application translates the first speech, and the first interface is used for dialogue translation.

18. The method according to claim 16 or 17, characterized in that, Before the terminal device receives the first voice message sent by the earpiece, the method further includes: The terminal device receives the first user's operation; In response to the first user operation, the terminal device displays a first interface and sends a first message to the headset; the first interface is used for dialogue translation, and the first message is used to instruct the headset to enter dialogue translation mode.

19. The method according to claim 18, characterized in that, After sending the first message to the earphone, the method further includes: The terminal device receives a second user operation; In response to a second user action, the terminal device acquires a second voice message and sends a second message to the headset, the second message being used to instruct the headset to exit the dialogue translation mode.

20. The method according to claim 19, characterized in that, After sending the second message to the earpiece, the method further includes: The terminal device receives operations from a third user. In response to a third user action, the terminal device stops acquiring the second voice and sends a third message to the headset, the third message being used to instruct the headset to enter dialogue translation mode.

21. The method according to claim 19, characterized in that, After acquiring the second voice, the method further includes: When the terminal device detects that the user has paused the sentence on the terminal side, it sends a third message to the headset, which is used to instruct the headset to enter the dialogue translation mode.

22. The method according to claim 19, characterized in that, After acquiring the second voice, the method further includes: When the terminal device detects that the user on the terminal side has paused a sentence, it stops collecting the second voice.

23. An electronic device, characterized in that, include: A processor and a memory, the memory being coupled to the processor, the memory being used to store computer program code, the computer program code including computer instructions, which, when the processor reads from the memory, cause the electronic device to perform the dialogue translation method as described in any one of claims 10 to 15 or 16 to 22.

24. An earphone, characterized in that, include: A processor and a memory, the memory being coupled to the processor, the memory being used to store computer program code, the computer program code including computer instructions, which, when the processor reads from the memory, cause the server to perform the dialogue translation method as described in claims 10 to 15 or claims 16 to 22.

25. A computer-readable storage medium, characterized in that, Includes computer instructions that, when executed on a server, cause the server to perform the dialogue translation method as described in claims 10 to 15 or claims 16 to 22.

26. A computer program product, characterized in that, When the computer program product is run on a computer, it causes the computer to perform the dialogue translation method as described in claims 10 to 15 or claims 16 to 22.

Citation Information

Patent Citations

  • Earphone control method, translation method, earphone and cloud server

    CN111783481A

  • Real-time dialogue translation method and system, earphone equipment and mobile terminal

    CN114245261A

  • Speech translation processing apparatus

    US20240370669A1