Bluetooth earphone-based speech translation system and method
The Bluetooth earphone-based speech translation system uses signal processing and gain modules to differentiate between speech sources, ensuring accurate translation and reducing inefficiencies in communication.
Patent Information
- Application Number
- JP2024556570
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-08-30
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2043-08-30
AI Technical Summary
Existing Bluetooth earphone-based translation systems struggle with inconsistent speech recognition and translation due to the inability to differentiate between speech signals from different individuals, leading to inefficient communication and unnecessary translation operations.
A speech translation system using Bluetooth earphones that employs a voice signal processing center with Fourier transform, signal cross-correlation, and gain modules to identify the source of speech signals, setting appropriate gain factors for accurate translation.
The system accurately identifies the source of speech signals, reducing inconsistent recognition and translation errors, thereby enhancing communication efficiency and minimizing unnecessary translation operations.
Smart Images

Figure 2025530945000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to the field of speech translation, and in particular to a speech translation system and method based on Bluetooth earphones. [Background technology]
[0002] With economic globalization, commercial and daily communication between countries has become increasingly frequent. International academic conferences are often held between countries, and language differences mean that communication typically requires the use of translation devices. To ensure confidentiality and convenience, translation devices often use a pair of Bluetooth earphones, with each party wearing a Bluetooth earphone. Each earphone picks up audio signals, transmits the audio signals to a translation device or cloud translation engine for translation, and then transmits the translated audio signals to the other earphone for playback. This approach often facilitates efficient communication and exchanges between people who speak different languages. However, Bluetooth earphones have certain limitations. People typically communicate face-to-face, and the distance between them is relatively short. For example, suppose a pair of Bluetooth earphones are earphone A and earphone B, and the two people communicating are Person A and Person B. Person A is wearing earphone A and Person B is wearing earphone B. When Person A speaks, because the two people are close to each other, both Earphone A and Earphone B can pick up Person A's voice signal, but the translation machine or cloud translation engine cannot determine whether the voice signal is coming from Person A or Person B. In this case, the translation machine or cloud translation engine will identify and translate any audio signals from Earphone A and Earphone B. That is, the translation machine or cloud translation engine will identify and translate the voice signal from Person A picked up by Earphone A, and will also identify and translate the voice signal from Person A picked up by Earphone B. The same can happen when only Person B is speaking. This is likely to cause inaccurate recognition and translation by the translation machine or translation engine, which will affect normal communication and interaction and increase the number of unnecessary translation operations, resulting in a waste of translation resources. Summary of the Invention [Problem to be solved by the invention]
[0003] The purpose of the present invention is to overcome the drawbacks of the prior art by providing a speech translation system based on Bluetooth earphones, which recognizes which of the two parties in the communication the speech signals collected by the two Bluetooth earphones originate from, and transmits only the speech signal from the person wearing the earphone to a translation machine or cloud translation engine for speech recognition and translation, thereby solving the problem of inconsistent recognition and translation, improving communication efficiency, and significantly reducing unnecessary speech recognition and translation. [Means for solving the problem]
[0004] In order to achieve the above object, the present invention provides a speech translation system using Bluetooth earphones, The system comprises: a first translation Bluetooth earphone and a second translation Bluetooth earphone worn by each of the users who are to interact with each other; a voice signal processing center including a Fourier transform module, a signal cross-correlation processing module, a determination module, and a gain module, wherein the first voice signal and the second voice signal collected by the first translation Bluetooth earphone and the second translation Bluetooth earphone, respectively, are sent to the Fourier transform module for time-frequency signal processing, and then the signal cross-correlation processing module performs signal cross-correlation processing on the first voice signal and the second voice signal; based on the magnitude of the signal cross-correlation value of the first voice signal and the second voice signal, the determination module determines whether the first voice signal and the second voice signal originate from the same sound source; if the first voice signal and the second voice signal originate from the same sound source, the gain module determines the position of the sound source based on the time delay relationship between the first voice signal and the second voice signal; and the voice signal processing center sets a gain factor G1 of the first voice signal and a gain factor G2 of the second voice signal according to the acquired information on the position of the sound source; a translation module that recognizes and translates the first and second audio signals after they have been processed by the gain module, and transmits the translated audio signals to the first translation Bluetooth earphone or the second translation Bluetooth earphone.
[0005] According to one embodiment of the present invention, when the numerical range of the signal cross-correlation value between the first audio signal and the second audio signal is (0.7, 1), the first audio signal and the second audio signal are the same sound source signal.
[0006] According to one embodiment of the present invention, when the first audio signal and the second audio signal are both audio signals from a user wearing the first translation Bluetooth earphone, the first gain factor G1=1 and the second gain factor G2=0, and when the first audio signal and the second audio signal are both audio signals from a user wearing the second translation Bluetooth earphone, the first gain factor G1=0 and the second gain factor G2=1.
[0007] According to one embodiment of the present invention, the audio signal processing center further includes a signal amplitude detection module for detecting the magnitude of the signal amplitude of the first audio signal and the second audio signal. When it is determined that the first audio signal and the second audio signal originate from the same sound source, if the signal amplitude of the first audio signal is greater than the signal amplitude of the second audio signal, both the first audio signal and the second audio signal are audio signals from a user wearing the first translation Bluetooth earphone, and if the signal amplitude of the second audio signal is greater than the signal amplitude of the first audio signal, both the first audio signal and the second audio signal are audio signals from a user wearing the second translation Bluetooth earphone.
[0008] Another object of the present invention is to provide a speech translation method based on Bluetooth earphones, the method comprising the steps of: performing Fourier transform on the first speech signal and the second speech signal collected by the first translation Bluetooth earphone and the second translation Bluetooth earphone, respectively, to process time-frequency signals; performing signal cross-correlation processing on the first audio signal and the second audio signal after time-frequency signal processing to obtain a signal cross-correlation value between the first audio signal and the second audio signal; and determining whether the first audio signal and the second audio signal originate from the same sound source based on the magnitude of the signal cross-correlation value; a step of determining a position of a sound source from a time delay relationship between the first sound signal and the second sound signal when it is determined that the first sound signal and the second sound signal originate from the same sound source; and a step of setting a gain factor G1 of the first sound signal and a gain factor G2 of the second sound signal based on information about the position of the sound source. and a step of increasing the gain of the first audio signal and the second audio signal, outputting the signal to a translation machine or a cloud translation engine for recognition and translation, and transmitting the translated audio signal in association with the first translation Bluetooth earphone or the second translation Bluetooth earphone.
[0009] According to an embodiment of the present invention, the step of determining whether the first audio signal and the second audio signal originate from the same audio source comprises: The cross-correlation function between the first audio signal and the second audio signal is
number
number
number
number
number
number
[0010] According to one embodiment of the present invention, the position of a sound source can be determined from the time delay relationship between the first audio signal and the second audio signal. Equation (4) is
number
number
[0011] The present invention provides a speech translation system and method using Bluetooth earphones, which identifies whether a first speech signal and a second speech signal originate from the same sound source. After determining which specific sound source the first speech signal and the second speech signal originate from, the system sets gain factors corresponding to the first speech signal and the second speech signal, respectively. That is, if the first speech signal and the second speech signal are both from a user wearing a first translation Bluetooth earphone, the first gain factor G1=1 and the second gain factor G2=0. If the first speech signal and the second speech signal are both from a user wearing a second translation Bluetooth earphone, the first gain factor G1=0 and the second gain factor G2=1. The gained first speech signal and the second speech signal are sent to a translation machine or a cloud translation engine for translation, thereby solving the problem of inconsistent recognition and translation, improving communication efficiency, and significantly reducing unnecessary speech recognition and translation. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 2 is a diagram showing an application scene model according to an embodiment of the present invention. [Figure 2] 1 is a principle block diagram of a speech translation system based on a Bluetooth earphone of the present invention; [Figure 3] 2 is a flowchart of a voice translation method based on a Bluetooth earphone of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0013] The invention of the present application will now be described in more detail with reference to specific embodiments.
[0014] FIG. 1 is an example of an application scenario model of the present invention. A pair of Bluetooth earphones typically includes two, one on the left and one on the right, worn by two people communicating with each other. For example, a pair of Bluetooth earphones may include a first translation Bluetooth earphone (the left Bluetooth earphone) and a second translation Bluetooth earphone (the right Bluetooth earphone). The two earphones are worn with the first translation Bluetooth earphone located at position A and the second translation Bluetooth earphone located at position B. The two people communicating with each other are Person A, who wears the first translation Bluetooth earphone, and Person B, who wears the second translation Bluetooth earphone. The present invention will be described in detail below using this scenario model. FIG. 2 is a block diagram illustrating the principle of a speech translation system for Bluetooth earphones according to the present invention. The system includes a first translation Bluetooth earphone and a second translation Bluetooth earphone worn by each of the users communicating with each other. In the embodiment of the present invention, the first translation Bluetooth earphone and the second translation Bluetooth earphone may be two left and two right Bluetooth earphones of a pair, or may be two unpaired wireless Bluetooth earphones.
[0015] The speech translation system of the present invention also includes a speech signal processing center, including a memory module, processor, power supply module, and other components required to implement the translation system's functions. The speech signal processing center primarily includes a Fourier transform module, a signal cross-correlation processing module, a determination module, and a gain module as its signal processing core. The signal collected by the first translation Bluetooth earphone is a first speech signal, and the signal collected by the second translation Bluetooth earphone is a second speech signal. The first and second speech signals are sent to the Fourier transform module for time-frequency signal processing, after which the signal cross-correlation processing module performs signal cross-correlation processing on the first and second speech signals. A signal cross-correlation value is obtained, and a determination unit determines whether the first and second speech signals originate from the same sound source based on the magnitude of this signal cross-correlation value. If it is determined that the first and second speech signals originate from the same sound source, the location of the sound source is determined based on the time delay relationship between the first and second speech signals. If it is determined that the first and second audio signals do not originate from the same sound source, the audio signal processing center retains only the function of collecting the original signals without processing the two audio signals. The gain module sets a gain factor G1 for the first audio signal and a gain factor G2 for the second audio signal based on the acquired position information of the sound source.
[0016] In an embodiment of the present invention, the speech translation system further includes a translation module, wherein the first and second speech signals are processed by a gain module and then recognized and translated by the translation module. The translated speech signals are sent to the first or second translation Bluetooth earphone, i.e., the first translation Bluetooth earphone identifies and translates the first speech signal collected by the first translation Bluetooth earphone and sends it to the second translation Bluetooth earphone for playback. The second translation Bluetooth earphone identifies and translates the second speech signal collected by the second translation Bluetooth earphone and sends it to the first translation Bluetooth earphone for playback. This allows for accurate speech translation and avoids translation errors and confusion.
[0017] In one embodiment of the present invention, whether the first audio signal and the second audio signal are the same sound source signal is determined based on the magnitude of the signal cross-correlation value of the first and second audio signals.If the value of the signal cross-correlation value of the first and second audio signals is (0.7, 1), it can be determined that the first audio signal and the second audio signal are the same sound source signal; if not, the first audio signal and the second audio signal are not the same sound source signal.
[0018] In one embodiment of the present invention, when the first and second audio signals are from the same audio source, the following processing of the signals is required: For example, if the first and second audio signals are both from a user wearing a first translation Bluetooth earphone, the first gain factor G1 = 1 and the second gain factor G2 = 0; if the first and second audio signals are both from a user wearing a second translation Bluetooth earphone, the first gain factor G1 = 0 and the second gain factor G2 = 1. The purpose of setting the gain factors in this way is to facilitate accurate recognition and translation to obtain a target audio signal.
[0019] According to one embodiment of the present invention, the audio signal processing center further includes a signal amplitude detection module for detecting whether the first audio signal and the second audio signal are of equal amplitude. When it is determined that the first audio signal and the second audio signal originate from the same sound source, if the signal amplitude of the first audio signal is greater than the signal amplitude of the second audio signal, both the first audio signal and the second audio signal are from a user wearing a first translation Bluetooth earphone. Conversely, both the first audio signal and the second audio signal are from a user wearing a second translation Bluetooth earphone.
[0020] Another object of the present invention is to provide a voice translation method based on Bluetooth earphones. The method comprises: First, Fourier transform the first audio signal and the second audio signal collected by the first translation Bluetooth earphone and the second translation Bluetooth earphone, respectively, to process the time-frequency signal; Then, after the time-frequency signal processing, a signal cross-correlation process is performed on the first audio signal and the second audio signal to obtain a signal cross-correlation value between the first audio signal and the second audio signal, and a step of determining whether the first audio signal and the second audio signal originate from the same sound source based on the magnitude of the signal cross-correlation value; determining a position of the sound source based on a time delay relationship between the first sound signal and the second sound signal when it is determined that the first sound signal and the second sound signal originate from the same sound source; setting a gain factor G1 of the first audio signal and a gain factor G2 of the second audio signal based on the position information of the sound source; Finally, the method includes a step of increasing the gain of the first audio signal and the second audio signal, outputting them to a translation machine or a cloud translation engine for recognition and translation, and transmitting the translated audio signal in association with the first translation Bluetooth earphone or the second translation Bluetooth earphone.
[0021] According to one embodiment of the present invention, it is determined whether the first audio signal and the second audio signal originate from the same audio source as follows. The cross-correlation function between the first audio signal and the second audio signal is
number
number
number
number
number
number
number
[0022] According to one embodiment of the present invention, when a first audio signal and a second audio signal are both from the same sound source, the position of the sound source can be determined from the time delay relationship between the first audio signal and the second audio signal. Equation (4) is
number
number
number
number
[0023] As described above, the present invention provides a speech translation system and method based on Bluetooth earphones. The system identifies whether a first speech signal and a second speech signal originate from the same sound source. After determining the specific sound source from which the first speech signal and the second speech signal originate, the system sets gain factors corresponding to the first speech signal and the second speech signal, respectively. That is, if the first speech signal and the second speech signal are both speech signals from a user wearing a first translation Bluetooth earphone, the first gain factor G1=1 and the second gain factor G2=0. If the first speech signal and the second speech signal are both speech signals from a user wearing a second translation Bluetooth earphone, the first gain factor G1=0 and the second gain factor G2=1. The gained first speech signal and the second speech signal are then sent to a translation machine or a cloud translation engine for translation, thereby solving the problem of inaccurate recognition and translation, improving communication efficiency, and significantly reducing unnecessary speech recognition and translation.
[0024] The above describes in detail preferred embodiments of the present invention, but the present invention is not limited to the above embodiments, and various modifications can be made within the scope of the knowledge possessed by a person skilled in the art without departing from the spirit of the present invention.
Claims
1. A Bluetooth earphone-based speech translation system, comprising: a first translation Bluetooth earphone and a second translation Bluetooth earphone worn by each of the interacting users; a voice signal processing center including a Fourier transform module, a signal cross-correlation processing module, a determination module, and a gain module, wherein the first voice signal and the second voice signal collected by the first translation Bluetooth earphone and the second translation Bluetooth earphone, respectively, are sent to the Fourier transform module for time-frequency signal processing, and then the signal cross-correlation processing module performs signal cross-correlation processing on the first voice signal and the second voice signal; the determination module determines whether the first voice signal and the second voice signal originate from the same sound source based on the magnitude of the signal cross-correlation value of the first voice signal and the second voice signal; if the first voice signal and the second voice signal originate from the same sound source, the gain module determines the position of the sound source based on the time delay relationship between the first voice signal and the second voice signal; and the gain module sets a gain factor G1 of the first voice signal and a gain factor G2 of the second voice signal based on the acquired information on the position of the sound source; a translation module that recognizes and translates the first and second audio signals after the first and second audio signals are processed by the gain module, and transmits the translated audio signals to the first or second translation Bluetooth earphone; A speech translation system comprising:
2. 2. The speech translation system according to claim 1, wherein when the numerical range of the signal cross-correlation value between the first speech signal and the second speech signal is (0.7, 1), the first speech signal and the second speech signal are the same sound source signal.
3. When the first audio signal and the second audio signal are both audio signals from a user wearing the first translation Bluetooth earphone, the first gain factor G1=1 and the second gain factor G2=0; The speech translation system of claim 1, wherein when the first speech signal and the second speech signal are both speech signals from a user wearing the second translation Bluetooth earphone, the first gain factor G1 = 0 and the second gain factor G2 = 1.
4. the audio signal processing center further includes a signal amplitude detection module that detects the magnitude of the signal amplitude of the first audio signal and the second audio signal; When it is determined that the first audio signal and the second audio signal originate from the same sound source, If the signal amplitude of the first audio signal is greater than the signal amplitude of the second audio signal, both the first audio signal and the second audio signal are audio signals from a user wearing the first translation Bluetooth earphone; The speech translation system of claim 1, wherein if the signal amplitude of the second speech signal is greater than the signal amplitude of the first speech signal, both the first speech signal and the second speech signal are speech signals from a user wearing the second translation Bluetooth earphone.
5. A Bluetooth earphone-based speech translation method for a Bluetooth earphone-based speech translation system according to any one of claims 1 to 4, comprising: The speech translation method based on the Bluetooth earphone includes the steps of: Fourier transforming the first speech signal and the second speech signal collected by the first translation Bluetooth earphone and the second translation Bluetooth earphone, respectively, to process time-frequency signals; performing signal cross-correlation processing on the first audio signal and the second audio signal after time-frequency signal processing to obtain a signal cross-correlation value between the first audio signal and the second audio signal, and determining whether the first audio signal and the second audio signal originate from the same sound source based on the magnitude of the signal cross-correlation value; determining a position of the sound source based on a time delay relationship between the first sound signal and the second sound signal when it is determined that the first sound signal and the second sound signal originate from the same sound source; setting a gain factor G1 of the first audio signal and a gain factor G2 of the second audio signal based on position information of a sound source; a step of increasing the gain of the first audio signal and the second audio signal, outputting the first audio signal and the second audio signal to a translation device or a cloud translation engine, performing recognition and translation, and transmitting the translated audio signal to the first translation Bluetooth earphone or the second translation Bluetooth earphone; A speech translation method based on a Bluetooth earphone comprising:
6. The step of determining whether the first audio signal and the second audio signal originate from the same audio source comprises: The cross-correlation function between the first audio signal and the second audio signal is [Equation 20] Fulfilling x 1 (t) is a signal propagation model of the first audio signal, and x 2 (t) is a signal propagation model of the second audio signal; [Equation 21] Fulfilling [Equation 22] Fulfilling Here, t represents time, s(●) represents the sound source model, and n 1 (●) and n 2 (●) represents the noise model, and x 1 (●), x 2 (●) represent the signal models received by the first translating Bluetooth earphone and the second translating Bluetooth earphone, respectively, and τ 1 , τ 2 represent the time when the sound source propagates to the first translating Bluetooth earphone and the second translating Bluetooth earphone, respectively; α and β represent the energy attenuation factors when the sound source propagates to the first translating Bluetooth earphone and the second translating Bluetooth earphone, respectively; and τ represents the signal propagation delay; Furthermore, assuming that the noise signal and the speech signal are not correlated and that the noise signals are not correlated with each other, the cross-correlation function between the first speech signal and the second speech signal is [Equation 23] Fulfilling In equation (4), [0000] Calculate the value of [Equation 25] 6. The speech translation method using a Bluetooth earphone according to claim 5, wherein when the value range of σ is (0.7, 1), the first speech signal and the second speech signal are speech signals from the same sound source.
7. determining a position of a sound source from a time delay relationship between the first audio signal and the second audio signal, Equation (4) is [Equation 26] Convert to As can be seen from the properties of the correlation function, [0000] and the time delay τ between the first audio signal and the second audio signal is τ=τ 1 -τ 2 7. The method of claim 6, wherein a time delay τ is calculated, and if the time delay τ is positive, it indicates that both the first speech signal and the second speech signal are speech signals from a user wearing the second translation Bluetooth earphone, while if the time delay τ is negative, it indicates that both the first speech signal and the second speech signal are speech signals from a user wearing the first translation Bluetooth earphone.
Citation Information
Patent Citations
Voice processor, voice processing method, and voice processing program
JP2007264473A
Translation function providing method for electronic apparatus and ear set apparatus
JP2019012507A
Audio processing apparatus, audio processing method, and audio processing program
JP2019101385A
Method and apparatus for wind noise attenuation
JP2023509593A