Voice signal processing method and device, electronic equipment and readable storage medium

By establishing two independent communication links on the same headset device for data transmission, the problem of manually setting the translation language in existing technologies is solved, enabling real-time multilingual translation and automatic voice signal conversion, thus improving communication efficiency.

CN121037488APending Publication Date: 2025-11-28VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511201832.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

In existing technologies, when electronic devices perform voice translation, users need to wait for the translation results and manually set the language, which makes the translation process cumbersome and communication inefficient.

Method used

By establishing two independent communication links on the same headset device, data can be transmitted to the first headset device and the second headset device respectively, enabling real-time translation between multiple languages ​​and automatically converting the language of the voice signal.

Benefits of technology

It simplifies the translation process, improves communication efficiency, and enables automatic speech signal conversion without the need for manual language settings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121037488A_ABST
    Figure CN121037488A_ABST
Patent Text Reader

Abstract

The invention discloses a voice signal processing method and device, electronic equipment and a readable storage medium, and belongs to the technical field of electronic equipment. The method comprises the following steps: receiving first input; in response to the first input, establishing a first communication link with the first earphone device, and establishing a second communication link with the second earphone device, the first earphone device and the second earphone device being the same pair of earphone devices; under the condition that a first voice signal sent by the first earphone device through the first communication link is received, the first voice signal is converted into a second voice signal, the second voice signal is sent to the first earphone device through the first communication link, and the language of the first voice signal is different from that of the second voice signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of electronic equipment technology, and specifically relates to a speech signal processing method, apparatus, electronic device and readable storage medium. Background Technology

[0002] With the continuous development of communication technology, electronic devices are becoming increasingly versatile. For example, they can be used to translate speech.

[0003] Specifically, when two users who speak different languages ​​are having a conversation, they can take turns speaking into the electronic device. The electronic device can collect the speaker's voice signal and translate the collected voice signal through a translation application in the electronic device, and then output the translated voice signal.

[0004] However, the above method is cumbersome and inefficient because a user must wait for the electronic device to output the translated voice signal of the other user before speaking; and each time a translation is performed, the user must manually set the translation language, such as translating Chinese to English or English to Chinese. Summary of the Invention

[0005] The purpose of this application is to provide a speech signal processing method, apparatus, electronic device, and readable storage medium that can simplify the translation process and improve communication efficiency.

[0006] In a first aspect, embodiments of this application provide a voice signal processing method, the method comprising: receiving a first input; in response to the first input, establishing a first communication link with a first earphone device and establishing a second communication link with a second earphone device, wherein the first earphone device and the second earphone device are the same earphone device; upon receiving a first voice signal transmitted by the first earphone device through the first communication link, converting the first voice signal into a second voice signal, and transmitting the second voice signal to the first earphone device through the first communication link, wherein the language of the first voice signal is different from the language of the second voice signal.

[0007] Secondly, embodiments of this application provide a voice signal processing apparatus, which includes: a receiving module, an establishing module, and a processing module; the receiving module is used to receive a first input; the establishing module is used to establish a first communication link with a first earphone device and a second communication link with a second earphone device in response to the first input received by the receiving module, wherein the first earphone device and the second earphone device are the same earphone device; the processing module is used to convert the first voice signal into a second voice signal when receiving a first voice signal sent by the first earphone device through the first communication link, and send the second voice signal to the first earphone device through the first communication link, wherein the language of the first voice signal is different from the language of the second voice signal.

[0008] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0009] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0010] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0011] In a sixth aspect, embodiments of this application provide a computer program / program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.

[0012] In this embodiment, a first input can be received; and in response to the first input, a first communication link is established with a first earphone device, and a second communication link is established with a second earphone device, wherein the first earphone device and the second earphone device are the same earphone device; upon receiving a first voice signal sent by the first earphone device through the first communication link, the first voice signal is converted into a second voice signal, and the second voice signal is sent back to the first earphone device through the first communication link, wherein the language of the first voice signal is different from that of the second voice signal. Through this scheme, since data transmission can be performed with the first earphone device and the second earphone device through two independent communication links, bidirectional audio streams can be processed independently, enabling real-time multilingual translation; moreover, the voice signal received from the earphone device can be automatically converted into a voice signal in another language, eliminating the need to manually set the translation language each time translation is performed. This simplifies the translation process and improves communication efficiency. Attached Figure Description

[0013] Figure 1 This is one of the flowcharts of the speech signal processing method provided in the embodiments of this application;

[0014] Figure 2 This is the second flowchart of the speech signal processing method provided in the embodiments of this application;

[0015] Figure 3 This is the third flowchart of the speech signal processing method provided in the embodiments of this application;

[0016] Figure 4 This is the fourth flowchart of the speech signal processing method provided in the embodiments of this application;

[0017] Figure 5 This is a schematic diagram of the headphone language setting interface in the voice signal processing method provided in the embodiments of this application;

[0018] Figure 6 This is the fifth flowchart of the speech signal processing method provided in the embodiments of this application;

[0019] Figure 7 This is the sixth flowchart of the speech signal processing method provided in the embodiments of this application;

[0020] Figure 8 This is the seventh flowchart of the speech signal processing method provided in the embodiments of this application;

[0021] Figure 9 This is the eighth flowchart of the speech signal processing method provided in the embodiments of this application;

[0022] Figure 10 This is a schematic diagram of the speech signal processing device provided in the embodiments of this application;

[0023] Figure 11 This is a schematic diagram of the electronic device provided in the embodiments of this application;

[0024] Figure 12 This is a hardware schematic diagram of the electronic device provided in the embodiments of this application. Detailed Implementation

[0025] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0026] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0027] The term "instruction" in this application can be either a direct instruction (or explicit instruction) or an indirect instruction (or implicit instruction). A direct instruction can be understood as one in which the sender explicitly informs the receiver of specific information, the operation to be performed, or the requested result, etc., in the instruction sent. An indirect instruction can be understood as one in which the receiver determines the corresponding information based on the instruction sent by the sender, or makes a judgment and determines the operation to be performed or the requested result, etc., based on the judgment result.

[0028] The terms "at least one," "at least one of," etc., used in the specification and claims of this application refer to any one, any two, or a combination of two or more of the included items. For example, at least one of a, b, and c can mean: "a," "b," "c," "a and b," "a and c," "b and c," and "a, b, and c," where a, b, and c can be single or multiple. Similarly, "at least two" refers to two or more items, and its meaning is similar to that of "at least one."

[0029] The following explains some concepts and terms involved in the speech signal processing method, apparatus, electronic device and readable storage medium provided in the embodiments of this application.

[0030] True Wireless Stereo (TWS) earbuds: These are earbuds with no physical cables connecting them to the device or between the left and right earbud units, relying entirely on wireless technology to transmit audio signals. This design completely eliminates the constraints of traditional wired headphones, providing a more flexible wearing experience. The left and right earbuds can work independently or collaborate to achieve synchronized stereo transmission. The core principle involves pairing the device via Bluetooth, and using a master / slave earbud signal relay to achieve channel separation, ensuring high sound quality and low latency.

[0031] Audio playback algorithms primarily encompass two core areas: playback gain algorithms and audio streaming playback technology. These are used to standardize audio loudness and optimize real-time audio transmission, respectively. The playback gain algorithm is an audio standardization technique that calculates the difference between the perceived loudness of the audio and the target loudness through psychoacoustic analysis and writes the gain value into the file's metadata. Players supporting this standard can automatically adjust the volume to unify the perceived loudness of different audio tracks, while avoiding volume fluctuations and clipping caused by recording differences. Audio streaming playback technology supports both static data and dynamic streaming transmission. Static mode is suitable for scenarios with small data volumes, while streaming mode is suitable for high sampling rates and real-time data transmission, such as online music or conferencing scenarios. This technology reduces latency and ensures sound quality by optimizing data transmission strategies.

[0032] The speech signal processing method, apparatus, electronic device, and readable storage medium provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.

[0033] The speech signal processing method provided in this application can be applied to scenarios involving real-time multilingual translation.

[0034] One possible application scenario is a meeting where participants speak different languages; another is a user traveling abroad and communicating with locals. Of course, this embodiment only illustrates these two scenarios, and the speech signal processing method provided in this embodiment can also be applied to any multilingual real-time translation scenario; this embodiment does not limit its application.

[0035] In the voice signal processing method provided in this application embodiment, a first input can be received; and in response to the first input, a first communication link is established with a first earphone device and a second communication link is established with a second earphone device, wherein the first earphone device and the second earphone device are the same earphone device; and upon receiving a first voice signal sent by the first earphone device through the first communication link, the first voice signal is converted into a second voice signal, and the second voice signal is sent to the first earphone device through the first communication link, wherein the language of the first voice signal is different from the language of the second voice signal.

[0036] Thus, since data can be transmitted to the first and second earphone devices through two independent communication links, bidirectional audio streams can be processed independently, enabling real-time multilingual translation. Furthermore, the audio signal received from the earphone device can be automatically converted into audio signals in other languages, eliminating the need to manually set the translation language each time. This simplifies the translation process and improves communication efficiency.

[0037] It should be noted that the speech signal processing method provided in this application can be executed by a speech signal processing device, an electronic device, or a functional module within an electronic device. Some embodiments of this application use an electronic device executing the speech signal processing method as an example to illustrate the speech signal processing method provided in this application.

[0038] Figure 1 A flowchart of a speech signal processing method provided in an embodiment of this application is shown. Figure 1 As shown, the speech signal processing method provided in this application embodiment may include the following steps 101 to 103.

[0039] Step 101: The electronic device receives the first input.

[0040] Optionally, in this embodiment of the application, the first input is used to establish an independent communication link with the headphone device.

[0041] Optionally, in this embodiment, the first input can be any possible form of input, such as touch input, voice input, or physical button input.

[0042] For example, taking the first input as a touch input, the first input may include, but is not limited to: single-click input, hard-press input, swipe input or other feasible inputs made by the user on the headphone connection settings interface with a finger or stylus. The specific input can be determined according to actual usage needs, and this application embodiment does not limit it.

[0043] For example, taking the first input as voice input, the first input may include, but is not limited to, any possible voice input such as the voice "establish an independent communication link with the headphone device" or the voice "establish an independent connection with the headphone device".

[0044] For example, taking the first input as a physical button input, the first input may include, but is not limited to, any of the following: pressing the lock screen button, pressing the volume "+" button, or pressing both the lock screen button and the volume "+" button simultaneously.

[0045] Optionally, in this embodiment, the first input can be used to launch a translation application in the electronic device. When the user launches the translation application through the first input, the electronic device can begin to establish an independent communication link with the headphone device.

[0046] Step 102: In response to the first input, the electronic device establishes a first communication link with the first earphone device and a second communication link with the second earphone device.

[0047] The first and second earphone devices are the same pair of earphone devices.

[0048] Optionally, in this embodiment of the application, the first communication link and the second communication link are different, that is, the first communication link and the second communication link are independent communication links.

[0049] Optionally, in this embodiment of the application, the first communication link and the second communication link can both be Bluetooth communication links, that is, the first communication link and the second communication link are independent Bluetooth communication links.

[0050] Optionally, in this embodiment of the application, the first earphone device can be the left earphone device in a TWS earphone device, and the second earphone device can be the right earphone device in the TWS earphone device; or, the first earphone device can be the right earphone device in a TWS earphone device, and the second earphone device can be the left earphone device in the TWS earphone device.

[0051] Optionally, in this embodiment of the application, before receiving the first input, the electronic device may have already established a connection with the first and second earphone devices in a conventional stereo mode, or the electronic device may not have established a connection with the first and second earphone devices.

[0052] Optionally, in this embodiment of the application, after receiving the first input, the electronic device can establish a first communication link with the first earphone device and a second communication link with the second earphone device through a dual-channel independent link architecture.

[0053] It's worth noting that the dual-channel independent link architecture can synchronously transmit multiple audio streams via asynchronous channels. Through system and firmware upgrades to older phones and headphones, the phone can simultaneously connect to both the left and right earbuds, establishing independent Bluetooth communication links with each, replacing the traditional master-slave forwarding mode. The left and right earbuds each have built-in independent microphone arrays, separately capturing the wearer's voice and transmitting it to the phone for processing via independent links, avoiding signal contention.

[0054] Exemplarily, assume that the first input is used to start the translation application in the electronic device. Then, when the user starts the translation application through the first input, the electronic device can automatically disconnect the traditional stereo mode with the first headset device and the second headset device, and switch to the dual independent communication mode through the dual-channel independent link architecture, and establish independent Bluetooth communication links with the first headset device and the second headset device respectively.

[0055] Optionally, in the embodiments of the present application, after the electronic device establishes the first communication link with the first headset device and the second communication link with the second headset device, data can be transmitted with the first headset device through the first communication link, and data can be transmitted with the second headset device through the second communication link.

[0056] Step 103: In the case of receiving the first voice signal sent by the first headset device through the first communication link, the electronic device converts the first voice signal into a second voice signal, and sends the second voice signal to the first headset device through the first communication link.

[0057] Wherein, the language of the first voice signal is different from the language of the second voice signal.

[0058] Optionally, in the embodiments of the present application, the first voice signal can be the voice signal of the wearer of the first headset device collected by the built-in microphone array of the first headset device.

[0059] Optionally, in the embodiments of the present application, when the electronic device converts the first voice signal into a second voice signal, it can be understood that: the electronic device translates the first voice signal into a second voice signal in another language.

[0060] Optionally, in the embodiments of the present application, the language of the first voice signal can be Chinese (or referred to as Chinese), and the language of the second voice signal can be English (or referred to as English); or, the language of the first voice signal can be French (or referred to as French), and the language of the second voice signal can be Japanese (or referred to as Japanese), etc.

[0061] It can be understood that the first voice signal and the second voice signal are voice signals with different languages but the same semantics.

[0062] For example, the first voice signal is the Chinese "你好" and the second voice signal is the English "hello".

[0063] Optionally, in the embodiments of the present application, the first communication link can be used for the first headset device to send the original voice signal to the electronic device, and for the electronic device to send the translation result of the original voice signal to the first headset device.

[0064] Optionally, in this embodiment of the application, the first voice signal may be: a voice signal obtained by the first earphone device compressing the original voice signal collected by the first earphone device through an audio frame compression algorithm, thereby reducing the amount of data transmitted to the electronic device.

[0065] Optionally, in the embodiments of this application, the above-mentioned audio frame compression algorithm can be the Opus audio frame compression algorithm.

[0066] It should be noted that the core goal of the Opus codec audio frame compression algorithm is to provide a low-latency, high-quality, and flexible single coding scheme suitable for various audio application scenarios such as real-time network voice transmission, music streaming, and video conferencing. It innovatively integrates two coding techniques: Speech-optimized Linear Predictive Coding (SILK) and Constrained Energy Lapped Transform (CELT). SILK, a speech-optimized linear predictive coding technique, excels at handling narrowband to wideband speech signals, while CELT, based on an improved discrete cosine transform music encoder, supports full-band high-fidelity music.

[0067] Optionally, in this embodiment of the application, the electronic device can detect Bluetooth channel interference through dynamic frequency band switching and automatically hop to a low-interference frequency band to ensure transmission stability.

[0068] Optionally, in this embodiment of the application, after receiving the second voice signal sent by the electronic device through the first communication link, the first earphone device can immediately decompress it and play it through the built-in speaker module, while simultaneously enabling dual-microphone environmental noise reduction to suppress external noise interference.

[0069] Optionally, in the embodiments of this application, combined with Figure 1 ,like Figure 2 As shown, step 103 can be implemented through steps 103a and 103b as described below.

[0070] Step 103a: Upon receiving a first voice signal sent by the first earphone device through the first communication link, the electronic device converts the first voice signal into a first intermediate voice signal.

[0071] Optionally, in this embodiment of the application, the first intermediate speech signal can be a speech signal obtained by translating the first speech signal.

[0072] Step 103b: The electronic device renders the first intermediate voice signal using an audio playback algorithm to obtain the second voice signal, and sends the second voice signal to the first earphone device through the first communication link.

[0073] Optionally, in this embodiment of the application, the electronic device can enhance the sound quality by rendering the first intermediate speech signal through an audio playback algorithm, that is, the sound quality of the second speech signal is enhanced compared to the sound quality of the first intermediate speech signal.

[0074] The method for rendering the first intermediate speech signal by an electronic device using an audio playback algorithm can be found in the specific descriptions in related technologies, and will not be repeated here to avoid repetition.

[0075] In this embodiment, since the electronic device can first convert the first speech signal into a first intermediate speech signal, then render the first intermediate speech signal through an audio playback algorithm, and send the rendered second speech signal to the first earphone device through the first communication link, the enhanced audio quality translation speech signal can be sent to the first earphone device, thereby improving the audio quality of the translation result playback.

[0076] Optionally, in the embodiments of this application, combined with Figure 1 ,like Figure 3 As shown, step 103 above can be implemented through step 103c below.

[0077] Step 103c: Upon receiving a first voice signal sent by the first earphone device through the first communication link, and if the similarity between the voiceprint features of the first voice signal and the first voiceprint features is greater than or equal to a first threshold, the electronic device converts the first voice signal into a second voice signal and sends the second voice signal to the first earphone device through the first communication link.

[0078] Among them, the first voiceprint feature is the voiceprint feature of the user of the first headphone device.

[0079] Optionally, in this embodiment of the application, the electronic device may first activate the voiceprint binding module to collect the first voiceprint feature of the user of the first earphone device, and then bind and store the first voiceprint feature with the input language of the first earphone device for subsequent automatic language routing.

[0080] Optionally, in this embodiment, the first threshold can be the system default or can be arbitrarily set by the user according to actual usage needs; this embodiment does not limit this.

[0081] It is understandable that when the similarity between the voiceprint features of the first voice signal and the first voiceprint feature is greater than or equal to the first threshold, the first voice signal can be considered to have been emitted by the user of the first earphone device. In this case, the electronic device can convert the first voice signal into a second voice signal and send the second voice signal to the first earphone device through the first communication link. When the similarity between the voiceprint features of the first voice signal and the first voiceprint feature is less than the first threshold, the first voice signal can be considered not to have been emitted by the user of the first earphone device. In this case, the electronic device may not perform any voice signal processing.

[0082] In this embodiment, the electronic device will only convert the first voice signal into a second voice signal and send the second voice signal to the first earphone device through the first communication link if the similarity between the voiceprint feature of the first voice signal and the first voiceprint feature is greater than or equal to the first threshold. Therefore, the translation action can be performed on the premise that the first voice signal is sent by the user of the first earphone device, thereby improving the accuracy and security of voice signal processing.

[0083] In the speech signal processing method provided in this application embodiment, since data transmission can be performed with the first and second earphone devices through two independent communication links, bidirectional audio streams can be processed independently, enabling real-time multilingual translation. Furthermore, the speech signal received from the earphone device can be automatically converted into speech signals in other languages, eliminating the need to manually set the translation language each time a translation is performed. This simplifies the translation process and improves communication efficiency.

[0084] Optionally, in the embodiments of this application, combined with Figure 1 ,like Figure 4 As shown, after step 102 and before step 103, the speech signal processing method provided in this application embodiment may further include the following steps 104 and 105.

[0085] Step 104: The electronic device receives the second input.

[0086] Optionally, in this embodiment of the application, the second input is used to determine the language of the voice signal output by the headphone device.

[0087] Optionally, in this embodiment of the application, the second input can be any possible form of input, such as touch input, voice input, or physical button input.

[0088] For example, taking the above-mentioned second input as a touch input, the second input may include, but is not limited to: single-click input, hard-press input, swipe input or other feasible inputs made by the user on the headphone language setting interface with a finger or stylus. The specific input can be determined according to actual usage needs, and this application embodiment does not limit it.

[0089] For example, taking the second input as voice input, the second input may include, but is not limited to, any possible voice input such as the voice "the first headphone device outputs Chinese" or the voice "the second headphone device outputs English".

[0090] For example, taking the above-mentioned second input as a physical button input, the second input may include, but is not limited to, any of the following: input by pressing the lock screen button, input by pressing the volume "+" button, and input by pressing the lock screen button and the volume "+" button simultaneously.

[0091] Step 105: The electronic device responds to the second input and determines the first language of the voice signal converted from the voice signal sent to the first earphone device, and the second language of the voice signal converted from the voice signal sent to the second earphone device.

[0092] The language of the second speech signal is the first language.

[0093] Optionally, in the embodiments of this application, the first language and the second language may be the same or different.

[0094] Optionally, in the embodiments of this application, the first language can be Chinese and the second language can be English; or, the first language can be Japanese and the second language can be Italian, etc.

[0095] Optionally, in this embodiment of the application, the first language of the voice signal converted from the voice signal sent by the first earphone device can be understood as: the language of the voice signal output by the first earphone device.

[0096] Optionally, in this embodiment of the application, the second language of the voice signal converted from the voice signal sent by the second earphone device can be understood as: the language of the voice signal output by the second earphone device.

[0097] For example, the first earphone device is the left earphone device of a TWS earphone device, and the second earphone device is the right earphone device of the TWS earphone device. Then, as... Figure 5 As shown, users can set the first language (Chinese) of the converted voice signal sent to the left ear device and the second language (English) of the converted voice signal sent to the right ear device via the second input on the headphone language setting interface 50. This way, upon receiving a voice signal from one of the headphone devices, the voice signal can be automatically converted to the corresponding set language.

[0098] In this embodiment, since the electronic device can determine the first language of the voice signal converted from the voice signal sent to the first earphone device and the second language of the voice signal converted from the voice signal sent to the second earphone device through the second input, the target language output by the earphone device can be set so that after receiving the voice signal sent by a certain earphone device, the voice signal can be processed directly according to the determined language, thereby improving the accuracy of translation and simplifying the translation process.

[0099] Optionally, in the embodiments of this application, combined with Figure 4 ,like Figure 6 As shown, step 103 above can be implemented through step 103d below.

[0100] Step 103d: When the first voice signal sent by the first earphone device through the first communication link is received and the language of the first voice signal is the second language, the electronic device converts the first voice signal into a second voice signal and sends the second voice signal to the first earphone device through the first communication link.

[0101] Optionally, in this embodiment, after the electronic device sets the first language of the converted voice signal from the voice signal sent to the first earphone device and the second language of the converted voice signal from the voice signal sent to the second earphone device via the second input, the electronic device will only convert the first voice signal into the second voice signal and send the second voice signal to the first earphone device via the first communication link if it receives the first voice signal sent by the first earphone device through the first communication link and the language of the first voice signal is the second language. This improves the accuracy of voice signal processing and avoids unnecessary voice signal processing.

[0102] Optionally, in this embodiment of the application, the electronic device may also automatically distinguish the input source based on the first voiceprint features and contextual semantic analysis after receiving the first voice signal, and then trigger the corresponding translation engine to translate the first voice signal into a second voice signal in a second language.

[0103] Optionally, in this embodiment of the application, if a language recognition conflict is detected when the electronic device is converting voice signals, such as when a bilingual mixed sentence is detected, the semantic verification module is activated to correct the routing result in combination with the dialogue context.

[0104] Optionally, in the embodiments of this application, the electronic device may use different models to process the received speech signals for different first and second languages.

[0105] For example, for common language pairs, such as Chinese as the first language and English as the second language, the received speech signal can be processed by the edge model, and the overall latency can be less than 100 milliseconds; for less common language pairs, such as Arabic as the first language and Italian as the second language, the received speech signal can be processed by the cloud model, and the overall latency can be less than 150 milliseconds.

[0106] In this embodiment, since the electronic device can first determine the language of the voice signal received from the first earphone device, and then convert and send the voice signal, unnecessary and useless translation can be reduced, the accuracy of translation can be improved, and the translation efficiency can be increased.

[0107] Optionally, in the embodiments of this application, combined with Figure 4 ,like Figure 7 As shown, after step 105 and before step 103, the speech signal processing method provided in this application embodiment may further include step 106.

[0108] Step 106: If the language data package stored in the electronic device does not include the language data package for the first language, the electronic device downloads the language data package for the first language.

[0109] Optionally, in this embodiment of the application, the electronic device first detects the stored language data packets, and then, if the stored language data packets do not include the language data packets for the first language, downloads the language data packets for the first language.

[0110] Optionally, in this embodiment of the application, the language data packet can be a language data packet cached by the system; for example, the language data packet may include Chinese data packets and English data packets, etc.

[0111] Optionally, in this embodiment of the application, the electronic device can download a language data package of the first language through a translation application.

[0112] Optionally, in this embodiment of the application, the electronic device may download a lightweight edge translation model corresponding to the first language, such as the TensorFlow Lite engine, while downloading the language data package of the first language.

[0113] For example, assuming the first language is Spanish, after the electronic device determines the first language of the voice signal converted from the voice signal sent to the first headset device and the second language of the voice signal converted from the voice signal sent to the second headset device through the second input, it can detect the stored language data packets. If the stored language data packets do not include the Spanish language data packets, the electronic device can download the Spanish language data packets and the corresponding lightweight edge translation model.

[0114] In the embodiments of the present application, since the electronic device can download the language data packet of the first language when the stored language data packets do not include the language data packet of the first language, there is no need for the user to manually download it, thus simplifying the process of downloading the language data packet and ensuring that subsequent voice signal processing operations can be executed normally.

[0115] Optionally, in the embodiments of the present application, in combination with Figure 1 , as Figure 8 shown, after the above step 103, the voice signal processing method provided by the embodiments of the present application may further include the following step 107.

[0116] It should be noted that Figure 8 only takes the execution of step 107 after step 103 for illustration. In actual implementation, step 107 may also be executed before step 103 or simultaneously with step 103, which is not limited in the embodiments of the present application.

[0117] Step 107: When the electronic device receives a third voice signal sent by the second earphone device through the second communication link, the electronic device converts the third voice signal into a fourth voice signal and sends the fourth voice signal to the second earphone device through the second communication link.

[0118] Among them, the language of the third voice signal is different from the language of the fourth voice signal.

[0119] Optionally, in the embodiments of the present application, the third voice signal may be the voice signal of the wearer of the second earphone device collected by the built-in microphone array of the second earphone device.

[0120] Optionally, in the embodiments of the present application, when the electronic device converts the third voice signal into a fourth voice signal, it can be understood that: the electronic device translates the third voice signal into a fourth voice signal in another language.

[0121] Optionally, in the embodiments of the present application, the language of the third voice signal may be English (or called English), and the language of the fourth voice signal may be Chinese (or called Chinese); or, the language of the third voice signal may be Japanese (or called Japanese), and the language of the fourth voice signal may be French (or called French), etc.

[0122] It can be understood that the third voice signal and the fourth voice signal are voice signals with different languages but the same semantics.

[0123] For example, the third voice signal is the English "hello", and the fourth voice signal is the Chinese "你好".

[0124] Optionally, in the embodiments of this application, the language of the first speech signal and the language of the fourth speech signal may be the same, and the language of the second speech signal and the language of the third speech signal may be the same.

[0125] Optionally, in this embodiment of the application, the second communication link can be used for the second earphone device to send the original voice signal to the electronic device, and for the electronic device to send the translation result of the original voice signal to the second earphone device.

[0126] Optionally, in this embodiment, the third voice signal can be: a voice signal obtained by the second earphone device compressing the original voice signal collected by the encoding and decoding audio frame compression algorithm, thereby reducing the amount of data transmitted to the electronic device.

[0127] Optionally, in the embodiments of this application, the above-mentioned audio frame compression algorithm can be the Opus audio frame compression algorithm.

[0128] Optionally, in this embodiment of the application, the electronic device can detect Bluetooth channel interference through dynamic frequency band switching and automatically hop to a low-interference frequency band to ensure transmission stability.

[0129] Optionally, in this embodiment of the application, after receiving the fourth voice signal sent by the electronic device through the second communication link, the second earphone device can immediately decompress it and play it through the built-in speaker module, while simultaneously enabling dual-microphone environmental noise reduction to suppress external noise interference.

[0130] Optionally, in the embodiments of this application, step 107 can be implemented by steps 107a and 107b as described below.

[0131] Step 107a: Upon receiving a third voice signal transmitted by the second earphone device through the second communication link, the electronic device converts the third voice signal into a second intermediate voice signal.

[0132] Optionally, in this embodiment, the second intermediate speech signal can be a speech signal obtained by translating the third speech signal.

[0133] Step 107b: The electronic device renders the second intermediate voice signal using an audio playback algorithm to obtain the fourth voice signal, and sends the fourth voice signal to the second earphone device through the second communication link.

[0134] Optionally, in this embodiment of the application, the electronic device can enhance the sound quality by rendering the second intermediate speech signal through an audio playback algorithm, that is, the sound quality of the fourth speech signal is enhanced compared to the sound quality of the second intermediate speech signal.

[0135] The method for rendering the second intermediate speech signal by an electronic device using an audio playback algorithm can be found in the specific descriptions in related technologies, and will not be repeated here to avoid repetition.

[0136] In this embodiment, since the electronic device can first convert the third speech signal into a second intermediate speech signal, then render the second intermediate speech signal through an audio playback algorithm, and send the rendered fourth speech signal to the second earphone device through the second communication link, the enhanced audio quality translation speech signal can be sent to the second earphone device, thereby improving the audio quality of the translation result playback.

[0137] Optionally, in the embodiments of this application, step 107 can be implemented by step 107c as described below.

[0138] Step 107c: When the electronic device receives a third voice signal sent by the second earphone device through the second communication link and the language of the third voice signal is the first language, the electronic device converts the third voice signal into a fourth voice signal and sends the fourth voice signal to the second earphone device through the second communication link.

[0139] Optionally, in this embodiment of the application, the language of the fourth speech signal can be the second language.

[0140] Optionally, in this embodiment, after the electronic device determines the first language of the converted speech signal from the speech signal sent to the first earphone device and the second language of the converted speech signal from the speech signal sent to the second earphone device via the second input, the electronic device will only convert the third speech signal into a fourth speech signal and send the fourth speech signal to the second earphone device via the second communication link if it receives a third speech signal sent by the second earphone device through the second communication link and the language of the third speech signal is the first language. This can improve the accuracy of translation and avoid useless translation operations.

[0141] Optionally, in this embodiment of the application, the electronic device may first collect the voiceprint features of the wearer of the second earphone device, and then, after receiving the third voice signal, automatically distinguish the input source based on the voiceprint features and contextual semantic analysis, and then trigger the corresponding translation engine to translate the third voice signal into a fourth voice signal in the first language.

[0142] Optionally, in this embodiment of the application, if a language recognition conflict is detected when the electronic device is converting voice signals, such as when a bilingual mixed sentence is detected, the semantic verification module is activated to correct the routing result in combination with the dialogue context.

[0143] Optionally, in the embodiments of this application, the electronic device may use different models to translate the received speech signals for different first and second languages.

[0144] For example, for common language pairs, such as Chinese as the first language and English as the second language, the received speech signal can be translated using an edge-side model, with an overall latency of less than 100 milliseconds; for less common language pairs, such as Arabic as the first language and Italian as the second language, the received speech signal can be translated using a cloud-based model, with an overall latency of less than 150 milliseconds.

[0145] In this embodiment, since the electronic device can first determine the language of the voice signal received from the second earphone device, and then convert and send the voice signal, unnecessary and useless translation can be reduced, the accuracy of translation can be improved, and the translation efficiency can be increased.

[0146] For example, suppose participant A, speaking Chinese, and participant B, speaking English, are discussing a meeting. Participant A can wear a first headset device and a second headset device. First, through a first input, an electronic device establishes a first communication link with the first headset device and a second communication link with the second headset device. Then, participant A begins to speak, and the first headset device transmits a first speech signal in Chinese to the electronic device via the first communication link. Upon receiving the first speech signal, the electronic device converts it into a second speech signal in English and transmits it back to the first headset device via the first communication link. Simultaneously, participant B can also speak, and the second headset device transmits a third speech signal in English to the electronic device via the second communication link. Upon receiving the third speech signal, the electronic device converts it into a fourth speech signal in Chinese and transmits it back to the second headset device via the second communication link.

[0147] For example, suppose user C is traveling abroad and communicating with a local person D. User C can wear a first headset device, and local person D can wear a second headset device. First, through a first input, the electronic device establishes a first communication link with the first headset device and a second communication link with the second headset device. Then, user C begins to speak, and the first headset device sends a first speech signal in Chinese to the electronic device via the first communication link. Upon receiving the first speech signal, the electronic device converts it into a second speech signal in Japanese and sends the second speech signal back to the first headset device via the first communication link. Simultaneously, local person D can also speak, and the second headset device sends a third speech signal in Japanese to the electronic device via the second communication link. Upon receiving the third speech signal, the electronic device converts it into a fourth speech signal in Chinese and sends the fourth speech signal back to the second headset device via the second communication link.

[0148] This allows two people to speak simultaneously, with the system processing two audio streams in parallel to achieve "seamless dialogue." The dual independent links and parallel threads break through the limitations of traditional TWS master-slave architectures, enabling simultaneous "speaking-translation-listening" with an end-to-end latency of <200ms. Compared to traditional technologies that require alternating speaking, this improves communication efficiency by 50%.

[0149] Optionally, in this embodiment of the application, if the language data package stored in the electronic device does not include a language data package for the second language, the electronic device may download a language data package for the second language.

[0150] Optionally, in this embodiment of the application, the electronic device can download a language data package of a second language through a translation application.

[0151] Optionally, in this embodiment of the application, the electronic device may download a lightweight edge translation model corresponding to the second language, such as the TensorFlow Lite engine, while downloading the language data package of the second language.

[0152] In this embodiment, since the stored language data package does not include the language data package for the second language, the electronic device can download the language data package for the second language. Therefore, the user does not need to download it manually, which simplifies the process of downloading the language data package and ensures that subsequent translation operations can be performed normally.

[0153] Optionally, in the embodiments of this application, combined with Figure 1 ,like Figure 9 As shown, after step 103 above, the speech signal processing method provided in this application embodiment may further include steps 108 and 109 as described below.

[0154] Step 108: The electronic device receives a third input.

[0155] Optionally, in this embodiment of the application, the third input is used to disconnect the first communication link and the second communication link.

[0156] Optionally, in this embodiment, the third input can be any possible form of input, such as touch input, voice input, or physical button input.

[0157] For example, taking the above-mentioned third input as touch input, the third input may include, but is not limited to: single-click input, hard-press input, swipe input or other feasible input by the user through a finger or stylus, etc., which can be determined according to actual usage needs, and this application embodiment does not limit it.

[0158] For example, taking the above-mentioned third input as voice input, the third input may include, but is not limited to, any possible voice input such as the voice "disconnect the first communication link and the second communication link" or the voice "disconnect the dual-channel independent link architecture".

[0159] For example, taking the above-mentioned third input as a physical button input, the third input may include, but is not limited to, any of the following: input by pressing the lock screen button, input by pressing the volume "+" button, and input by pressing the lock screen button and the volume "+" button simultaneously.

[0160] Optionally, in this embodiment, the third input can be used to close the voice signal processing application in the electronic device. When the user closes the voice signal processing application via the third input, the electronic device can begin to disconnect the first and second communication links.

[0161] Step 109: In response to the third input, the electronic device disconnects the first communication link and the second communication link.

[0162] Optionally, in this embodiment of the application, the electronic device may, in response to a third input, disconnect the first communication link and the second communication link, and immediately switch back to stereo mode to establish a connection with the first headphone device and the second headphone device, thereby releasing translation resources.

[0163] Optionally, in this embodiment of the application, the electronic device may, in response to a third input, disconnect the first communication link and the second communication link, and no longer establish a connection with the first earphone device and the second earphone device.

[0164] In this embodiment, since the electronic device can disconnect the first and second communication links through a third input, it can flexibly switch the connection mode with the headphone device and free up translation resources.

[0165] Optionally, the speech signal processing method provided in this application embodiment may further include the following step 110.

[0166] It should be noted that step 110 can be executed at any time during the above process when the triggering condition is met.

[0167] Step 110: If the duration during which the target headphone device is in a silent state is greater than or equal to the second threshold, the electronic device controls the target headphone device to enter a sleep state.

[0168] The target earphone device is at least one of the first earphone device and the second earphone device.

[0169] Optionally, in this embodiment, the second threshold can be the system default or can be arbitrarily set by the user according to actual usage needs.

[0170] For example, the second threshold can be any duration set by the user, such as 10 seconds, 8 seconds, or 15 seconds.

[0171] In this embodiment, when the duration of the target headphone device being in a silent state is greater than or equal to a second threshold, the electronic device controls the target headphone device to enter a sleep state. Therefore, the headphone device can be controlled to enter a sleep state when there is no need for voice signal processing, thereby reducing the power consumption of the headphone device.

[0172] The speech signal processing method provided in the embodiments of this application will be described exemplarily below with reference to the accompanying drawings.

[0173] For example, assuming the first earphone device is the left earphone device of a TWS earphone device, and the second earphone device is the right earphone device of the TWS earphone device, then the voice signal processing method provided in this application embodiment may include the following steps a to e:

[0174] Step a: Initialization and Link Establishment, i.e., the User Configuration Phase

[0175] (1) Device pairing and mode switching

[0176] When a user launches a translation application on their phone, the TWS earbuds automatically disconnect from the traditional stereo mode and switch to a dual independent communication mode via a dual-channel independent link architecture. This means that the left and right earbuds each establish an independent Bluetooth link with the phone: left ear → link 1, right ear → link 2.

[0177] like Figure 5 As shown, the mobile app displays the headphone language settings interface, where users can set the target language for the left and right earphones respectively, for example, the left earphone output language is Chinese and the right earphone output language is English.

[0178] (2) Translation engine loading

[0179] The system checks the local cache status of language data packets. If the target language, such as Spanish, is not downloaded, it automatically downloads the voice data packets and a lightweight edge translation model, such as the TensorFlow Lite engine, through a translation application.

[0180] (3) Start the voiceprint binding module, collect the voiceprint features of the wearer in both ears, and bind them to the preset language, i.e., left ear voiceprint → Chinese, right ear voiceprint → English, for subsequent automatic language routing.

[0181] Step b: Voice acquisition and preprocessing, i.e., binaural parallel input.

[0182] (1) Directional sound pickup and noise reduction

[0183] Left earphone device: The microphone array activates beamforming to focus and capture the wearer A's Chinese voice, which is then transmitted to the mobile phone in real time via link 1. At the same time, the left earphone speaker module is turned off to avoid feedback.

[0184] Right earphone device: Simultaneously collects wearer B's English speech and transmits it to the mobile phone via link 2. The two signals do not compete with each other, that is, the secondary earphone is directly connected to the mobile phone, which is not a traditional master-slave forwarding.

[0185] It adopts a Bluetooth 5.0+ dual-path Synchronous Connection-Oriented (SCO) link to independently transmit the original voice and translation results, and low-latency, anti-interference transmission ensures bidirectional data stream synchronization.

[0186] (2) Audio compression and anti-interference

[0187] The headphone device compresses the original audio frames using the Opus codec audio frame compression algorithm, reducing the amount of data transmission with a compression rate of over 50%.

[0188] The mobile device detects Bluetooth channel interference through dynamic frequency band switching and automatically hops to a low-interference frequency band to ensure transmission stability.

[0189] Step c: Core translation processing, i.e., mobile smart routing.

[0190] (1) Language identification and task distribution

[0191] The language routing engine, based on voiceprint binding and contextual semantic analysis, automatically distinguishes the input source.

[0192] Left ear speech → Recognized as Chinese → Triggers Chinese-English translation engine → Outputs English result;

[0193] Right ear speech → recognized as English → triggers English-Chinese translation engine → outputs Chinese result.

[0194] If there is a language conflict, such as in a bilingual mixed sentence, the semantic verification module is activated to correct the routing results based on the dialogue context.

[0195] (2) Low-latency translation execution

[0196] An edge-cloud collaborative architecture is adopted: common language pairs are processed by the edge model with a latency of <100ms; less common language pairs call the cloud API with a latency of <150ms.

[0197] The dual translation threads run in parallel, supporting bidirectional simultaneous translation, allowing wearer A and wearer B to speak at the same time.

[0198] Step d: Playback of translation results

[0199] (1) Audio playback rendering: The translation results are optimized and rendered using an audio playback algorithm to enhance sound quality.

[0200] (2) Anti-interference playback: After receiving the translation data, the headphone device immediately decompresses it and plays it through the speaker module. At the same time, dual-microphone environmental noise reduction (ENR>30dB) is enabled to suppress external noise interference.

[0201] Step e: Session management and termination, i.e., dynamic control

[0202] Anomaly Handling: If the mute timeout of one earphone device is detected, for example, if it exceeds 10 seconds, it will automatically enter sleep power saving mode.

[0203] Exit mechanism: When the user gives the voice command "Exit Translation" or manually closes the translation application, the system switches back to stereo mode and releases translation resources.

[0204] In the voice signal processing method provided in this application embodiment, the dual-channel independent link architecture can synchronously transmit multiple audio streams through asynchronous channels. By upgrading the system and firmware of older mobile phones and earphones, the mobile phone can simultaneously connect to the left and right earphones and establish independent Bluetooth communication links with each earphone, replacing the traditional master-slave forwarding mode. The left and right earphones have built-in independent microphone arrays to collect the wearer's voice and transmit it to the mobile phone for processing through independent links, avoiding signal contention.

[0205] Furthermore, the mobile application provides a headset language settings interface, allowing users to set the target language for both the left and right earbuds separately. The system automatically identifies the input voice language and routes it to the corresponding translation engine, supporting bidirectional synchronous translation. Through voiceprint recognition and language detection algorithms, it distinguishes the source of voice from both ears, eliminating the need for manual switching. Employing a dual-channel SCO link with Bluetooth 5.0+, it independently transmits the original voice and translation results, ensuring synchronized bidirectional data streams. Audio frame compression algorithms such as Opus codec are introduced to reduce the amount of data transmitted, keeping end-to-end latency within 200ms.

[0206] This enables independent voice capture by two individuals, real-time translation, and directional playback, solving core issues such as independent communication links between the two ears, real-time multilingual translation, and low-latency interaction. Furthermore, it innovatively incorporates voiceprint routing and spatial audio directional playback, forming a complete technological loop that can be further expanded to multi-earphone networking in conference scenarios.

[0207] The voice signal processing method provided in this application allows two people to speak simultaneously, with the system processing two audio streams in parallel to achieve "seamless dialogue." The dual independent links and parallel threads break through the limitations of traditional TWS master-slave architectures, enabling synchronous "speaking-translating-listening" with an end-to-end latency of <200ms. Compared to related technologies that require alternating speaking, communication efficiency is improved by 50%. No dedicated translation headset is needed; firmware upgrades and a mobile application support existing TWS headsets for independent dual-person voice capture, real-time translation, and targeted playback, resulting in low hardware modification costs. Voiceprint binding and semantic verification solve language misidentification problems, eliminating the need for manual language pair switching. It is adaptable to multiple scenarios, including: business meetings, real-time translation between Chinese and English bilingual conferences; tourist guides, cross-language communication between tour guides and tourists; and barrier-free communication, allowing hearing-impaired individuals to engage in dialogue via text-to-speech.

[0208] The above-described method embodiments, or various possible implementations of the method embodiments, can be executed individually, or, provided there are no contradictions, they can be combined with each other. The specific implementation can be determined according to actual usage requirements, and this application embodiment does not impose any restrictions on this.

[0209] The speech signal processing method provided in this application can be executed by a speech signal processing device. This application uses a speech signal processing device executing the speech signal processing method as an example to illustrate the speech signal processing device provided in this application.

[0210] like Figure 10 As shown in the figure, this application provides a speech signal processing device 10, which may include a receiving module 11, an establishing module 12 and a processing module 13.

[0211] The receiving module 11 can be used to receive a first input. The establishing module 12 can be used to establish a first communication link with a first earphone device and a second communication link with a second earphone device in response to the first input received by the receiving module 11, wherein the first earphone device and the second earphone device are the same earphone device. The processing module 13 can be used to convert the first voice signal sent by the first earphone device through the first communication link into a second voice signal, and send the second voice signal to the first earphone device through the first communication link, wherein the language of the first voice signal is different from the language of the second voice signal.

[0212] In one possible implementation, the receiving module 11 can also be used to receive a second input. The processing module 13 can also be used, in response to the second input received by the receiving module 11, to determine the first language of the voice signal converted from the voice signal sent to the first earphone device, and the second language of the voice signal converted from the voice signal sent to the second earphone device; wherein the language of the second voice signal is the first language.

[0213] In one possible implementation, the processing module 13 can also be used to download the language data package of the first language if the language data package stored in the electronic device does not include the language data package of the first language.

[0214] In one possible implementation, the processing module 13 can be used to convert the first speech signal into a second speech signal when the language of the first speech signal is the second language.

[0215] In one possible implementation, the processing module 13 can be used to convert the first speech signal into a first intermediate speech signal; and to render the first intermediate speech signal using an audio playback algorithm to obtain a second speech signal.

[0216] In one possible implementation, the processing module 13 can be used to convert the first speech signal into a second speech signal when the similarity between the voiceprint feature of the first speech signal and the first voiceprint feature is greater than or equal to a first threshold; wherein the first voiceprint feature is the stored voiceprint feature of the user of the first headphone device.

[0217] In one possible implementation, the processing module 13 can also be used to convert the third voice signal into a fourth voice signal when it receives the third voice signal sent by the second earphone device through the second communication link, and send the fourth voice signal to the second earphone device through the second communication link. The language of the third voice signal is different from that of the fourth voice signal.

[0218] In one possible implementation, the receiving module 11 can also be used to receive a third input. The processing module 13 can also be used to disconnect the first communication link and the second communication link in response to the third input received by the receiving module 11.

[0219] In one possible implementation, the processing module 13 can also be used to control the target headphone device to enter a sleep state when the duration of the detected target headphone device being in a silent state is greater than or equal to a second threshold; wherein the target headphone device is at least one of the first headphone device and the second headphone device.

[0220] In the speech signal processing apparatus provided in this application embodiment, since the speech signal processing apparatus can transmit data with the first earphone device and the second earphone device through two independent communication links respectively, it can independently process bidirectional audio streams and realize real-time translation between multiple languages; moreover, it can automatically convert the speech signal received from the earphone device into speech signals of other languages, eliminating the need to manually set the translation language each time translation is performed. This simplifies the translation process and improves communication efficiency.

[0221] The voice signal processing device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the scope of the device.

[0222] The speech signal processing device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0223] The speech signal processing apparatus provided in this application embodiment can implement the various processes implemented in the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0224] like Figure 11 As shown, this application embodiment also provides an electronic device 100, including a processor 101 and a memory 102. The memory 102 stores a program or instructions that can run on the processor 101. When the program or instructions are executed by the processor 101, they implement the various steps of the above-described speech signal processing method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0225] It should be noted that the electronic devices in the embodiments of this application include mobile electronic devices and non-mobile electronic devices.

[0226] Figure 12 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0227] like Figure 12 As shown, the electronic device 1000 includes, but is not limited to, components such as: radio frequency unit 1001, network module 1002, audio output unit 1003, input unit 1004, sensor 1005, display unit 1006, user input unit 1007, interface unit 1008, memory 1009, and processor 1010.

[0228] Those skilled in the art will understand that the electronic device 1000 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1010 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 12 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0229] The user input unit 1007 can be used to receive a first input. The processor 1010 can be used to establish a first communication link with a first headset device and a second communication link with a second headset device in response to the first input received by the user input unit 1007, wherein the first headset device and the second headset device are the same headset device; and upon receiving a first voice signal sent by the first headset device through the first communication link, converting the first voice signal into a second voice signal, and sending the second voice signal to the first headset device through the first communication link, wherein the language of the first voice signal is different from the language of the second voice signal.

[0230] In one possible implementation, the user input unit 1007 can also be used to receive a second input. The processor 1010 can also be used, in response to the second input received by the user input unit 1007, to determine a first language of the voice signal converted from the voice signal sent to the first headset device, and a second language of the voice signal converted from the voice signal sent to the second headset device; wherein the language of the second voice signal is the first language.

[0231] In one possible implementation, the processor 1010 can also be used to download the language data package of the first language when the language data package stored in the electronic device does not include the language data package of the first language.

[0232] In one possible implementation, the processor 1010 can be used to convert the first speech signal into a second speech signal when the language of the first speech signal is the second language.

[0233] In one possible implementation, the processor 1010 can be used to convert the first speech signal into a first intermediate speech signal; and to render the first intermediate speech signal using an audio playback algorithm to obtain a second speech signal.

[0234] In one possible implementation, the processor 1010 can be used to convert the first speech signal into a second speech signal when the similarity between the voiceprint features of the first speech signal and the first voiceprint features is greater than or equal to a first threshold; wherein the first voiceprint features are the stored voiceprint features of the user of the first headphone device.

[0235] In one possible implementation, the processor 1010 can also be used to convert the third voice signal into a fourth voice signal upon receiving a third voice signal sent by the second headset device through the second communication link, and send the fourth voice signal to the second headset device through the second communication link, wherein the language of the third voice signal is different from that of the fourth voice signal.

[0236] In one possible implementation, the user input unit 1007 can also be used to receive a third input. The processor 1010 can also be used to disconnect the first communication link and the second communication link in response to the third input received by the user input unit 1007.

[0237] In one possible implementation, the processor 1010 can also be used to control the target headphone device to enter a sleep state when the duration of the detected target headphone device being in a silent state is greater than or equal to a second threshold; wherein the target headphone device is at least one of a first headphone device and a second headphone device.

[0238] In the electronic device provided in this application embodiment, since the electronic device can transmit data with the first earphone device and the second earphone device through two independent communication links respectively, it can independently process bidirectional audio streams and realize real-time translation between multiple languages; moreover, it can automatically convert the voice signal received from the earphone device into voice signals in other languages, eliminating the need to manually set the translation language each time translation is performed. This simplifies the translation process and improves communication efficiency.

[0239] It should be understood that, in this embodiment, the input unit 1004 may include a graphics processing unit (GPU) 10041 and a microphone 10042. The GPU 10041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1006 may include a display panel 10061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1007 includes a touch panel 10071 and at least one of other input devices 10072. The touch panel 10071 is also called a touch screen. The touch panel 10071 may include a touch detection device and a touch controller. Other input devices 10072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0240] The memory 1009 can be used to store software programs and various data. The memory 1009 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1009 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1009 in this embodiment includes, but is not limited to, these and any other suitable types of memory.

[0241] The processor 1010 may include one or more processing units; optionally, the processor 1010 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into the processor 1010.

[0242] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described speech signal processing method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0243] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0244] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described speech signal processing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0245] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0246] This application provides a computer program / program product stored in a storage medium. The program / program product is executed by at least one processor to implement the various processes of the above-described speech signal processing method embodiments and achieve the same technical effects. To avoid repetition, further details are omitted here.

[0247] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0248] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0249] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A speech signal processing method, characterized in that, The method includes: Receive the first input; In response to the first input, a first communication link is established with a first earphone device and a second communication link is established with a second earphone device, wherein the first earphone device and the second earphone device are the same earphone device; Upon receiving a first voice signal sent by the first earphone device through the first communication link, the first voice signal is converted into a second voice signal, and the second voice signal is sent back to the first earphone device through the first communication link. The language of the first voice signal is different from that of the second voice signal.

2. The method according to claim 1, characterized in that, The method further includes: Receive the second input; In response to the second input, a first language of the voice signal converted from the voice signal sent to the first headphone device and a second language of the voice signal converted from the voice signal sent to the second headphone device are determined. The language of the second speech signal is the same as the first language.

3. The method according to claim 2, characterized in that, The method further includes: If the language data package stored in the electronic device does not include the language data package for the first language, download the language data package for the first language.

4. The method according to claim 2, characterized in that, The step of converting the first speech signal into a second speech signal includes: If the language of the first speech signal is the second language, the first speech signal is converted into the second speech signal.

5. The method according to claim 1, characterized in that, The step of converting the first speech signal into a second speech signal includes: Convert the first speech signal into a first intermediate speech signal; The first intermediate speech signal is rendered using an audio playback algorithm to obtain the second speech signal.

6. The method according to claim 1, characterized in that, The step of converting the first speech signal into a second speech signal includes: If the similarity between the voiceprint features of the first speech signal and the first voiceprint features is greater than or equal to a first threshold, the first speech signal is converted into the second speech signal. The first voiceprint feature is the stored voiceprint feature of the user of the first headphone device.

7. The method according to claim 1, characterized in that, The method further includes: Upon receiving a third voice signal sent by the second earphone device through the second communication link, the third voice signal is converted into a fourth voice signal and sent to the second earphone device through the second communication link. The language of the third voice signal is different from that of the fourth voice signal.

8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: Receive third input; In response to the third input, the first communication link and the second communication link are disconnected.

9. The method according to any one of claims 1 to 7, characterized in that, The method further includes: If the duration during which the target headphone device is in a silent state is greater than or equal to a second threshold, the target headphone device is controlled to enter a sleep state. The target earphone device is at least one of the first earphone device and the second earphone device.

10. A speech signal processing device, characterized in that, The device includes: a receiving module, an establishing module, and a processing module; The receiving module is used to receive the first input; The establishment module is used to establish a first communication link with the first earphone device and establish a second communication link with the second earphone device in response to the first input received by the receiving module, wherein the first earphone device and the second earphone device are the same earphone device; The processing module is configured to, upon receiving a first voice signal sent by the first earphone device through the first communication link, convert the first voice signal into a second voice signal, and send the second voice signal back to the first earphone device through the first communication link, wherein the language of the first voice signal is different from the language of the second voice signal.

11. An electronic device, characterized in that, It includes a processor and a memory, the memory storing programs or instructions that can run on the processor, the programs or instructions being executed by the processor to implement the steps of the speech signal processing method as described in any one of claims 1-9.

12. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the speech signal processing method as described in any one of claims 1-9.