Intelligent audio product with call inter-translation function

By integrating wireless reception and transmission modules, translation modules and other components in intelligent audio products, and using time and frequency domain features to extract voice information and perform real-time translation, the problem that existing call translation systems cannot achieve instant simultaneous translation, and efficient and accurate call translation is achieved.

CN120048267AInactive Publication Date: 2025-05-27AIPOWER TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510083388.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing call translation system cannot realize instant simultaneous translation, making it difficult for both parties to have a smooth and convenient conversation.

Method used

An intelligent audio product is designed, including wireless reception and transmission module, translation module, radio reception module, sound playback module and communication module. The speech information parameters are extracted based on time and frequency domain characteristics through the translation module, and real-time translation is completed using an Internet third-party translation server.

Benefits of technology

While real-time translation is realized, the information and emotional color of the voice signal is retained, the accuracy and authenticity of the translation are improved, and the immediacy of call translation is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120048267A_ABST
    Figure CN120048267A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent audio product with a call inter-translation function, and relates to the technical field of intelligent audio products. The intelligent audio product with the call inter-translation function comprises a wireless receiving and transmitting module, a translation module, a sound receiving module, a sound playing module and a communication module, the wireless receiving and transmitting module is connected with the sound receiving module, the sound receiving module is connected with the translation module, and the translation module completes translation based on voice signals based on voice information transmitted by the sound receiving module. According to the invention, content translation is completed through cooperation of each characteristic parameter with the Internet third-party translation server, local translation or cloud translation, information contained in the signals is concentrated while real-time translation is realized, the calculation amount in the voice signal processing process is greatly reduced, and the voice processing efficiency is improved. The problem that the translation difference of a traditional endpoint detection algorithm is large under the condition of a low signal-to-noise ratio is overcome, and the emotion color of a caller is still achieved while the translation timeliness is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent audio products, and in particular to an intelligent audio product with a call translation function. Background Art

[0002] With the development of society, global communication in various fields is becoming more and more frequent. In the process of cross-border economic and trade exchanges and cooperation, political exchanges, cultural communication, etc., when encountering language barriers, translation is required to communicate. When using a call terminal for remote calls, due to the language barriers and non-face-to-face communication, it becomes very difficult to introduce translation. Therefore, an effective tool and method are urgently needed to solve this problem. How to use translation software to achieve reasonable and accurate translation when talking in different languages ​​is still a global problem. At present, some call translation systems have also appeared on the market. However, these call translation systems generally have the following defects: they cannot solve the immediacy of call translation, that is, they cannot translate simultaneously, so that the two parties of the call cannot smoothly and conveniently conduct a call conversation.

[0003] For example, a call terminal and method for instant original voice translation of a call disclosed in China Invention Publication No. CN107465816A can easily realize instant translation of a call, and the translated voice has the characteristics of the speaker's original voice, with a stronger sense of reality, and the receiver can easily and accurately grasp the speaker's true expression intention. In addition, this call terminal uses the third-party translation service provided by the Internet, and does not need to be equipped with a server when used. As long as one party of the caller is equipped with the call terminal and the other party is equipped with a conventional call terminal, the call translation function of both parties can be realized, which is very flexible and convenient to use.

[0004] However, the sound quality, real-time translation, and independent use of the above-mentioned call terminals are limited by existing translation equipment, resulting in low translation efficiency and quality of current translation headphones, poor translation performance, and inability to be applied to more scenarios and user needs. It is unable to provide users with better translation services as speech recognition technology develops, and the translated language feature parameters are not processed well, resulting in poor sound quality, noise, and a lack of realism and emotional color. Summary of the invention

[0005] In view of the deficiencies in the prior art, the present invention provides an intelligent audio product with a call translation function, which solves the problems raised in the above background technology.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: an intelligent audio product with call translation function, including a wireless receiving and transmitting module, a translation module, a sound receiving module, a sound playing module, and a communication module;

[0007] The wireless receiving and transmitting module is connected to the radio module, the radio module is connected to the translation module, the translated information processed by the translation module is transmitted by the wireless receiving and transmitting module into the communication module, the communication module transmits the translated information based on the wireless receiving and transmitting module at the other end, and the voice playback module receives the translated information;

[0008] The translation module completes the translation based on the voice signal based on the voice information transmitted by the radio module. The translation module completes the content translation by an Internet third-party translation server, local translation, or cloud translation;

[0009] The source of the voice information received by the translation module is received by the wireless receiving and transmitting module.

[0010] During a two-person call, the product records and transmits the voice of caller A according to the wireless receiving and transmitting module and the radio, extracts and translates the characteristic parameters of caller A's voice according to the translation module, and completes the signal transmission and feedback of the communication base station according to the communication module and the wireless receiving and transmitting module, and caller B listens to the translated voice.

[0011] A further improvement of the technical solution of the present invention is that the translation module further includes a voice information parameter extraction unit. The voice information parameter extraction unit extracts parameters such as short-time energy, short-time average zero-crossing rate, formant, and pitch period in the voice information based on time-domain characteristics, and the voice information parameter extraction unit extracts parameters such as LPC, LPCC, LSP, short-time spectrum, and MFCC in the voice information based on frequency-domain characteristics;

[0012] Here, a brief description is given. Since short-time energy, short-time average zero-crossing rate, formant, and pitch period are based on time-domain characteristics, it is convenient to extract them with time as the basic unit. However, the extraction of LPC, LPCC, etc. based on frequency-domain characteristics is more troublesome, so a further description is given here;

[0013] After preprocessing such as frame division and pre-emphasis of the input voice signal, perform a short-time Fourier transform to obtain its spectrum, convert the time-domain signal into a frequency-domain signal, calculate the energy spectrum, and perform band-pass filtering on the energy spectrum based on the frequency domain. The center f(m) of the band-pass filter H m (k) is divided according to the Mel frequency scale, so:

[0014]

[0015] where f 1 、f n are the lowest and highest frequencies applied by the filter frequency, F s is the sampling frequency, N is the number of points of the FFT transformation, and according to the sampling formula Therefore, a frequency response wave is obtained at the outputs of m filters. After introducing DCT and differential cepstrum parameters, the dynamic characteristic parameters of LPC, LPCC, LSP, short-time spectrum, and MFCC are finally obtained.

[0016] A further improvement of the technical solution of the present invention is that the translation module further includes an identification unit. The identification unit completes parameter identification based on time-domain features and frequency-domain features through an identification algorithm. The identification algorithm includes the DTW algorithm and the HMM model algorithm. The translation module completes content translation by docking with an Internet third-party translation server, local translation, or cloud translation based on the parameter identification result.

[0017] A further improvement of the technical solution of the present invention is that the sound collection module and the sound playback module perform filtering and amplification based on voice information to complete signal enhancement and noise elimination of the voice information. The signal of the voice information enters through the MIC and enters the processor after amplification and filtering.

[0018] A further improvement of the technical solution of the present invention is that the communication module, the wireless receiving and transmitting module are implemented by a TF card. The TF card is installed on the intelligent audio product through plugging, fixing, and limiting.

[0019] A further improvement of the technical solution of the present invention is that the translation module further includes a voice synthesis unit. The voice synthesis unit completes voice synthesis TTS based on SAPI. The voice synthesis unit realizes the synthesis of the translation results of various parameters transmitted by an Internet third-party translation server, local translation, or cloud translation based on the SAPI interface;

[0020] The detailed disclosure of the common programming design of the voice synthesis unit is as follows:

[0021] Setting the output frequency of the synthesized voice ——

[0022] HRESULT hrOutputStream = m_cpVoice->GetOutputStream(&cpStream);

[0023] m_cpVoice->SetVolume((USHORT)hpos);

[0024] m_cpVoice->SetRate(hpos);

[0025] Setting the output of pitch and sound speed ——

[0026] m_cpVoice->SetVolume((USHORT)hpos);

[0027] m_cpVoice->SetRate(hpos);

[0028] If there is a mixture of Chinese and English, the involved decision function——

[0029] if(IsChina)

[0030] {if(chr <= 122 && chr >= 65)

[0031] {int iLen = i - iCbeg;

[0032] string strValue = strSpeak.Substring(iCbeg, iLen),

[0033] SpeakChina(strValue);

[0034] iEbeg = i;

[0035] IsChina = false;

[0036] }

[0037] A further improvement of the technical solution of the present invention is that the wireless receiving and transmitting module has a first channel and a second channel. The wireless receiving and transmitting module is connected to the sound receiving module through the first channel, and the wireless receiving and transmitting module is connected to the repository through the second channel. The repository is used to store the original audio file, the speech conversion text, the translation text, the translation text machine speech audio file, the original audio characteristic curve, and the translation text original voice audio file.

[0038] A further improvement of the technical solution of the present invention is that the sound receiving module, the sound playing module, and the translation module are implemented by an FPC flexible circuit board and are installed inside the product bracket. A metal reinforcement plate is provided at the top of the product bracket, and the metal reinforcement plate is fixedly connected to the product bracket to enhance the transmission signal.

[0039] Compared with the prior art, the beneficial effect of the present invention is that the translation module extracts the parameters of short-time energy, short-time average zero-crossing rate, formant, and pitch period in the voice information based on the time-domain characteristics, and extracts the parameters of LPC, LPCC, LSP, short-time spectrum, and MFCC in the voice information based on the frequency-domain characteristics. Finally, based on each characteristic parameter, combined with an Internet third-party translation server, local translation or cloud translation, content translation is completed. While realizing real-time translation, the information contained in the signal is concentrated, greatly reducing the calculation amount in the voice signal processing process, overcoming the large translation difference of the traditional endpoint detection algorithm in the case of low signal-to-noise ratio, and ensuring the timeliness of translation while still having the emotional color of the caller. Description of the Drawings

[0040] Figure 1 Schematic diagram of the structure of an intelligent audio product with call translation function;

[0041] Figure 2 Schematic diagram of the translation process during a two-person call of an intelligent audio product with call translation function;

[0042] Figure 3 Schematic diagram of the translation module of an intelligent audio product with call translation function extracting characteristic parameters of voice information based on the time domain;

[0043] Figure 4 Circuit diagram of the sound receiving module and the sound playing module of an intelligent audio product with call translation function;

[0044] 1. Metal reinforcement plate; 2. FPC flexible circuit board; 3. TF card; 4. Bracket. Detailed implementation manners

[0045] The following will describe various exemplary embodiments, features, and aspects of the present application in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements with the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings do not have to be drawn to scale unless otherwise specified.

[0046] The special term "exemplary" here means "serving as an example, an embodiment, or illustrative". Any embodiment described as "exemplary" here does not have to be construed as superior to or better than other embodiments.

[0047] In addition, in order to better illustrate the present application, numerous specific details are given in the following detailed embodiments. Those skilled in the art should understand that the present application can also be implemented without some specific details. In some instances, methods, means, and elements well known to those skilled in the art are not described in detail in order to highlight the gist of the present application.

[0048] The present invention provides an intelligent audio product with call translation function, including a wireless receiving and transmitting module, a translation module, a sound receiving module, a sound playing module, and a communication module;

[0049] The wireless receiving and transmitting module is connected to the sound receiving module, the sound receiving module is connected to the translation module, the translated information processed by the translation module is transmitted by the wireless receiving and transmitting module into the communication module, the communication module transmits the translated information based on the wireless receiving and transmitting module at the other end, and the sound playing module receives the translated information;

[0050] The translation module completes the translation based on the voice signal based on the voice information transmitted by the sound receiving module, and the translation module completes the content translation by an Internet third-party translation server, local translation, or cloud translation;

[0051] The source of the voice information received by the translation module is received by the wireless receiving and transmitting module.

[0052] During a two-person call, the product records and transmits the voice of caller A according to the wireless receiving and transmitting module and the microphone, extracts and translates the characteristic parameters of caller A's voice according to the translation module, and completes the signal transmission and feedback to the communication base station according to the communication module and the wireless receiving and transmitting module. Caller B listens to the translated voice.

[0053] The translation module further includes a voice information parameter extraction unit. The voice information parameter extraction unit extracts the parameters of short-time energy, short-time average zero-crossing rate, formant, and pitch period in the voice information based on time-domain features, and extracts the parameters of LPC, LPCC, LSP, short-time spectrum, and MFCC in the voice information based on frequency-domain features.

[0054] Here, a brief description is given. Since short-time energy, short-time average zero-crossing rate, formant, and pitch period are based on time-domain features, it is convenient to extract them with time as the basic unit. However, the extraction of LPC, LPCC, etc. based on frequency-domain features is more troublesome, so further description is given here.

[0055] After preprocessing such as frame division and pre-emphasis of the input voice signal, perform a short-time Fourier transform to obtain its spectrum, convert the time-domain signal into a frequency-domain signal, calculate the energy spectrum, and perform band-pass filtering on the energy spectrum based on the frequency domain. The center f(m) of the band-pass filter H m (k) is divided according to the Mel frequency scale, so:

[0056]

[0057] where f 1 、f n are the lowest and highest frequencies applied by the filter frequency, F s is the sampling frequency, N is the number of points of the FFT transformation. According to the sampling formula Therefore, the frequency response wave is obtained at the output of m filters. After introducing DCT and differential cepstrum parameters, the dynamic characteristic parameters of LPC, LPCC, LSP, short-time spectrum, and MFCC are finally obtained. The specific steps are as shown in the appendix Figure 3 shown;

[0058] The translation module further includes an identification unit. The identification unit completes parameter identification based on time-domain features and frequency-domain features by an identification algorithm. The identification algorithm includes the DTW algorithm and the HMM model algorithm. The translation module completes content translation based on the parameter identification results by connecting to an Internet third-party translation server, local translation, or cloud translation.

[0059] The sound receiving module and the sound playing module perform filtering and amplification based on voice information to complete signal enhancement and noise cancellation of the voice information. The signal of the voice information enters through the MIC, is amplified and filtered, and then enters the processor.

[0060] Among them, the sound receiving module and the reaction module are based on the amplification circuit in the audio product as shown in the appendix Figure 4 shown;

[0061] The communication module, the wireless receiving and transmitting module are implemented by a TF card. The TF card is installed on the intelligent audio product through plugging, fixing and limiting.

[0062] The translation module further includes a voice synthesis unit. The voice synthesis unit completes voice synthesis TTS based on SAPI. The voice synthesis unit synthesizes the translation results of each parameter transmitted by the Internet third-party translation server, local translation or cloud translation based on the SAPI interface;

[0063] The commonly used programming design of the voice synthesis unit is disclosed in detail:

[0064] Synthesized voice output frequency setting——

[0065] HRESULT hrOutputStream = m_cpVoice->GetOutputStream(&cpStream);

[0066] m_cpVoice->SetVolume((USHORT)hpos);

[0067] m_cpVoice->SetRate(hpos);

[0068] Pitch and speed output setting——

[0069] m_cpVoice->SetVolume((USHORT)hpos);

[0070] m_cpVoice->SetRate(hpos);

[0071] If there is a mixture of Chinese and English, the involved judgment function——

[0072] if(IsChina)

[0073] {if(chr <= 122 & chr >= 65)

[0074] {int iLen = i - iCbeg;

[0075] string strValue = strSpeak.Substring(iCbeg, iLen),

[0076] SpeakChina(strValue);

[0077] iEbeg = i;

[0078] IsChina = false;

[0079] }

[0080] The wireless receiving and transmitting module has a first channel and a second channel. The wireless receiving and transmitting module is connected to the sound receiving module through the first channel, and the wireless receiving and transmitting module is connected to the repository through the second channel. The repository is used to store original audio files, speech conversion texts, translation texts, translation text machine speech audio files, original audio characteristic curves, and original voice speech audio files of translation texts.

[0081] The sound receiving module, the sound playing module, and the translation module are implemented by an FPC flexible circuit board and are installed inside the product bracket. A metal reinforcement plate is provided at the top of the product bracket and is fixedly connected to the product bracket to enhance the transmission signal.

[0082] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.

[0083] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An intelligent audio product with call translation function, characterized in that: It includes wireless receiving and transmitting module, translation module, sound receiving module, sound playing module and communication module; The wireless receiving and transmitting module is connected to the sound receiving module, the sound receiving module is connected to the translation module, the translation information processed by the translation module is transmitted to the communication module by the wireless receiving and transmitting module, the communication module transmits the translation information based on the wireless receiving and transmitting module at the other end, and the sound playing module receives the translation information; The translation module completes the translation based on the voice signal based on the voice information transmitted by the sound receiving module, and the translation module completes the content translation by the Internet third-party translation server, local translation or cloud translation; The voice information received by the translation module comes from a wireless receiving and transmitting module; During a two-person call, the product completes recording and transmission of caller A's voice based on the wireless receiving and transmitting module and the radio, completes feature parameter extraction and translation of caller A's voice based on the translation module, completes signal transmission and return transmission of the communication base station based on the communication module and the wireless receiving and transmitting module, and caller B completes listening to the translated voice.

2. The intelligent audio product with call translation function according to claim 1, characterized in that: The translation module also includes a voice information parameter extraction unit, which extracts parameters of short-time energy, short-time average zero-crossing rate, formant, and pitch period in the voice information based on time domain features, and extracts parameters of LPC, LPCC, LSP, short-time spectrum, and MFCC in the voice information based on frequency domain features; The translation module performs short-time Fourier transform on the input speech signal after performing frame segmentation and pre-emphasis and other pre-processing to obtain its spectrum, converts the time domain signal into the frequency domain signal, obtains the energy spectrum, and performs band-pass filtering on the energy spectrum based on the frequency domain. The band-pass filter H m The center f(m) of (k) is divided according to the Mel frequency scale, so: Where f1, f n are the lowest and highest frequencies to which the filter frequency applies, F s is the sampling frequency, N is the number of points of FFT change, according to the sampling formula Therefore, a frequency response wave is obtained at the output of m filters. After the introduction of DCT and differential cepstrum parameters, the dynamic characteristic parameters of LPC, LPCC, LSP, short-time spectrum and MFCC are finally obtained.

3. The intelligent audio product with call translation function according to claim 1, characterized in that: The translation module also includes an identification unit, which uses an identification algorithm to complete parameter identification based on time domain features and frequency domain features. The identification algorithm includes a DTW algorithm and an HMM model algorithm. The translation module connects to an Internet third-party translation server, local translation or cloud translation based on the parameter identification result to complete content translation.

4. The intelligent audio product with call translation function according to claim 1, characterized in that: The sound receiving module and the sound playing module perform filtering and amplification based on the voice information to complete the signal enhancement and noise elimination of the voice information. The signal of the voice information enters from the MIC and enters the processor after amplification and filtering.

5. The intelligent audio product with call translation function according to claim 1, characterized in that: The communication module, wireless receiving and transmitting module are implemented by a TF card, and the TF card is installed on the smart audio product by plugging, fixing and limiting.

6. The intelligent audio product with call translation function according to claim 1, characterized in that: The translation module also includes a speech synthesis unit, which completes speech synthesis TTS based on SAPI. The speech synthesis unit realizes the synthesis of translation results of various parameters transmitted by an Internet third-party translation server, local translation or cloud translation based on the SAPI interface.

7. The intelligent audio product with call translation function according to claim 1, characterized in that: The wireless receiving and transmitting module has a first channel and a second channel. The wireless receiving and transmitting module is connected to the sound receiving module through the first channel, and the wireless receiving and transmitting module is connected to the storage library through the second channel. The storage library is used to store original audio files, speech conversion texts, translated texts, translated text machine voice audio files, original audio characteristic curves, and translated text original voice audio files.

8. The intelligent audio product with call translation function according to claim 1, characterized in that: The sound receiving module, the sound playing module and the translation module are implemented by an FPC soft circuit board and installed inside the product bracket. A metal reinforcement plate is provided on the top of the product bracket, and the metal reinforcement plate is fixedly connected to the product bracket.

Citation Information

Patent Citations

  • Speaker recognition method based on depth learning

    CN104157290A

  • Call terminal and method for performing instant original sound voice translation during call

    CN107465816A

  • Headset interpretation system

    CN108572950A

  • Push type fingerprint detection circuit

    CN208622121U

  • Vehicle-mounted audio information transmitter

    CN209964134U