Cross-language voice communication method and system
Through deep learning model combining streaming speech recognition with CTC and Attention mechanisms and machine translation model based on Transformer architecture, the problems of complexity and low accuracy of existing speech translation systems are solved, and efficient and accurate cross-language voice communication is achieved.
Patent Information
- Application Number
- CN202510471546.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing voice translation system has problems such as redundant functional, complex operation, unfriendly user interface, low accuracy of speech recognition and translation, poor real-time performance, and complexity of multilingual processing, which limits the application of real-time voice translation technology.
The deep learning model is adopted, a streaming speech recognition model combined with CTC and Attention mechanisms and a machine translation model based on Transformer architecture is realized to realize real-time recognition and accurate translation of speech. The system includes a registration and login module, a background management module, a call visualization module, a voice recognition interface and a machine translation interface, and supports cross-language voice communication.
It realizes the rapid and accurate text conversion of voice and translated into text in other languages, improving the efficiency and accuracy of cross-language voice communication. The system is simple and easy to use, and is suitable for most cross-language communication occasions.
Smart Images

Figure CN120015013A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech translation, and in particular to a cross-language speech communication method and system. Background Art
[0002] As international communication continues to increase, the need for cross-language communication is becoming increasingly urgent. In order to eliminate communication barriers between different languages, real-time speech translation systems have emerged. Traditional speech translation systems are mainly dedicated to translating spoken expressions into text in the target language, and have problems such as redundant functions, complex operations, and unfriendly user interfaces. At the same time, problems such as low speech recognition and translation accuracy, poor real-time performance, and high complexity in multilingual processing restrict the application of real-time speech translation technology. Summary of the invention
[0003] The purpose of the present invention is to provide a cross-language voice communication method and system, specifically to provide a system that is concise and stable and can provide a smooth and accurate real-time voice translation system.
[0004] To achieve the above-mentioned purpose, the present invention adopts the following technical scheme: a cross-language voice communication method, comprising an initiator A and at least one receiver B, the language of the initiator A is Aa, and the language of the receiver B is Bb; at the beginning of the communication, the initiator A initiates a call to the receiver B, and after the receiver B accepts the call, the initiator A and the receiver B conduct a cross-language voice call. During the cross-language voice call, either the initiator A or the receiver B can see the voice spoken by the other party converted into text in their own language, wherein the voice spoken by either party is converted into text in the language to which the voice belongs by calling a voice recognition interface, and the other party then calls a machine translation interface to translate the text in the language to which the voice belongs into text in their own language.
[0005] Specifically, the called speech recognition interface adopts a streaming speech recognition model based on CTC plus Attention, which is a deep learning model. The speech recognition model includes an encoder, a CTC decoder and an Attention decoder. The speech audio is processed to obtain a frame-level fbank feature sequence. The frame-level fbank feature sequence is then input into the encoder after undergoing Embedding operation, adding position encoding, and downsampling operations. The encoder outputs a hidden state feature sequence and inputs it into the CTC decoder and the Attention decoder for decoding to obtain the output text.
[0006] Specifically, the called machine translation interface adopts a translation model based on the Transformer architecture, which is a deep learning model. The translation model includes a four-layer encoder and a six-layer decoder. The encoder consists of a multi-head attention layer, a residual connection and a normalization layer, a feedforward fully connected layer, a residual connection and a normalization layer connected in sequence, and the decoder consists of a masked multi-head attention layer, a residual connection and a normalization layer, a multi-head attention layer, a residual connection and a normalization layer, a feedforward fully connected layer and a residual connection, and a normalization layer connected in sequence; after the source language text is input into the encoder for feature processing, a set of vectors of fixed length are output, and the decoder translates the vectors output by the encoder and outputs the translated text in the target language.
[0007] Specifically, when an initiator A and a receiver B conduct a cross-language voice call, the voice call is conducted through a protocol that publicly supports voice communication.
[0008] Specifically, the protocol that publicly supports voice communication adopts one of the http protocol, WebRTC protocol, SIP protocol, RTP protocol, XMPP protocol, H.323 protocol and MGCP protocol.
[0009] In order to realize the above-mentioned voice communication, the present invention also provides a cross-language voice communication system, including a registration and login module, a background management module, a call visualization module, a voice recognition interface and a machine translation interface. The registration and login module is used to provide registration and login services for the initiator A, and the background management module is used to save and verify the relevant information of the initiator A. The background management module is also used to establish a communication connection after the initiator A initiates the voice communication; after the receiver B receives the voice communication, both the initiator A and the receiver B can call the voice recognition interface and the machine translation interface.
[0010] Specifically, the speech recognition interface adopts a streaming speech recognition model based on CTC plus Attention, which is a deep learning model.
[0011] Specifically, the streaming speech recognition model based on CTC plus Attention consists of an encoder, a CTC decoder and an Attention decoder. The speech audio is processed to obtain a frame-level fbank feature sequence. The frame-level fbank feature sequence is then input to the encoder after undergoing embedding operations, adding position encoding, and downsampling operations. The encoder outputs a hidden state feature sequence and inputs it to the CTC decoder and Attention decoder for decoding to obtain the output text. Specifically, the machine translation interface adopts a translation model based on the Transformer architecture, which is a deep learning model.
[0012] Specifically, the translation model includes a four-layer encoder and a six-layer decoder, the encoder consists of a multi-head attention layer, a residual connection and normalization layer, a feedforward fully connected layer, a residual connection and normalization layer connected in sequence, and the decoder consists of a masked multi-head attention layer, a residual connection and normalization layer, a multi-head attention layer, a residual connection and normalization layer, a feedforward fully connected layer and a residual connection, and a normalization layer connected in sequence; after the source language text is input into the encoder for feature processing, a set of vectors of fixed length are output, and the decoder translates the vectors output by the encoder and outputs the translated text in the target language.
[0013] The beneficial effects of the present invention are as follows: by integrating and improving advanced speech recognition models and translation models, it is possible to quickly and accurately convert speech into text and translate it into text in other languages, thereby achieving efficient and accurate cross-language voice communication; and the overall system is simple and can be applied to most cross-language communication occasions. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Attached Figure 1 A schematic diagram of the connection structure of a streaming speech recognition model based on CTC plus Attention, which is a deep learning model, adopted by the speech recognition interface in the embodiment; Attached Figure 2 A schematic diagram of a specific connection structure of an encoder in a speech recognition model in an embodiment; Attached Figure 3 This is a schematic diagram of the connection structure of a machine translation interface in an embodiment that adopts a translation model based on a Transformer architecture that is a deep learning model. DETAILED DESCRIPTION
[0015] Embodiment 1, a cross-language voice communication method, includes an initiator A and at least one receiver B, the language of initiator A is Aa, and the language of receiver B is Bb; at the beginning of the communication, initiator A initiates a call to receiver B, and after receiver B accepts the call, initiator A and receiver B conduct a cross-language voice call; during the cross-language voice call, either initiator A or receiver B can see the voice spoken by the other party converted into text in their own language, wherein the voice spoken by either party is converted into text in the language to which the voice belongs by calling a voice recognition interface, and the other party then calls a machine translation interface to translate the text in the language to which the voice belongs into text in their own language. Among them, receiver B can be one or more, and when there are multiple receivers B, the languages of multiple receivers B can be different, and each receiver can use a voice recognition interface and a machine translation interface. When a cross-language voice call is made between an initiator A and a receiver B, the voice call is made through a protocol that publicly supports voice communication; wherein the protocol that publicly supports voice communication adopts one of the http protocol, WebRTC protocol, SIP protocol, RTP protocol, XMPP protocol, H.323 protocol and MGCP protocol.
[0016] In order to realize the above-mentioned cross-language voice communication, this embodiment provides a feasible system solution: a cross-language voice communication system, including a registration and login module, a background management module, a call visualization module, a voice recognition interface and a machine translation interface, wherein the registration and login module is used to provide registration and login services for the initiator A, the background management module is used to save and verify the relevant information of the initiator A, and the background management module is also used to establish a communication connection after the initiator A initiates the voice communication; after the receiver B receives the voice communication, both the initiator A and the receiver B can call the voice recognition interface and the machine translation interface. In this system, as long as the initiator A registers and logs in to the system, the voice communication can be initiated through the system. When initiating the voice communication, a link can be created by the background management module, which is sent by the initiator A to the receiver B through a conventional communication method, and the receiver B can communicate with the initiator A by clicking on the link.
[0017] Among them, refer to Figure 1, the speech recognition interface adopts a streaming speech recognition model based on CTC plus Attention, which is a deep learning model. Among them, CTC (Connectionist Temporal Classification) is a training target and decoding method for speech recognition and sequence learning tasks. The goal is to solve the sequence labeling problem (for details, please refer to the paper "Connectionist Temporal Classification: Labelling Unsegmented Sequence Datawith Recurrent Neural Networks"); Attention is an attention mechanism, which is an attention mechanism used in deep learning models. The speech recognition model includes an encoder, a CTC decoder and an Attention decoder. The speech audio is processed through audio to obtain a frame-level fbank feature sequence. The frame-level fbank feature sequence is then input into the encoder after undergoing Embedding operations, adding position encoding, and downsampling operations. The encoder outputs a hidden state feature sequence and inputs it into the CTC decoder and Attention decoder for decoding to obtain the output text. Among them, refer to Figure 2, the encoder of the speech recognition model consists of a sequentially connected Convolution layer, Multi-Head Self Attention layer, Convolution layer, CBHG layer, and layernorm. The above structure can better take into account global and local features compared with the traditional FFN (Feedforward Neural Network) layer, and the CBHG layer can comprehensively consider the prediction probability of speech of different lengths. Among them, the hidden state feature sequence output by the encoder of the speech recognition model is input into the CTC decoder, and the CTC decoder outputs the candidate sequence. The candidate sequence is then input into the Attention decoder together with the hidden state feature sequence to obtain the final output text. Among them, CBHG (Convolutional Block and Highway Network with Gated Recurrent Unit) is a neural network structure, which is usually used for audio and speech processing tasks. It mainly consists of a convolutional block: composed of a one-dimensional convolution layer and a batch normalization layer, which is used to capture the local features of the input sequence; a highway network: a mechanism that allows information to be transmitted deeper in the network; a gated recurrent unit (Gated Recurrent Unit, GRU): GRU is a variant of a recurrent neural network (RNN) and consists of three parts. For details, please refer to the paper "TACOTRON: TOWARDS END-TO-END SPEECH SYNTHESIS".
[0018] Specifically, refer to Figure 3 The machine translation interface adopts a translation model based on the Transformer architecture, which is a deep learning model. The Transformer architecture is a deep learning model architecture for natural language processing (NLP) and other sequence-to-sequence tasks (see the paper "Attention is All You Need" for details). The translation model includes a four-layer encoder and a six-layer decoder. The encoder consists of a multi-head attention layer, a residual connection and normalization layer, a feedforward fully connected layer, a residual connection and normalization layer connected in sequence, and the decoder consists of a masked multi-head attention layer, a residual connection and normalization layer, a multi-head attention layer, a residual connection and normalization layer, a feedforward fully connected layer and a residual connection, and a normalization layer connected in sequence. After the source language text is input into the encoder for feature processing, a set of vectors with a fixed length are output. The decoder translates the vectors output by the encoder and outputs the translated text in the target language.
[0019] Of course, the above are only preferred embodiments of the present invention, and are not intended to limit the scope of use of the present invention. Therefore, any equivalent changes made to the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A cross-language voice communication method, characterized in that: An initiator A and at least one receiver B, the language of initiator A is Aa, and the language of receiver B is Bb; at the beginning of the communication, initiator A initiates a call to receiver B, and after receiver B accepts the call, initiator A and receiver B conduct a cross-language voice call. During the cross-language voice call, either initiator A or receiver B can see the voice spoken by the other party converted into text in their own language, wherein the voice spoken by either party is converted into text in the language to which the voice belongs by calling a voice recognition interface, and the other party then calls a machine translation interface to translate the text in the language to which the voice belongs into text in their own language; the called voice recognition interface adopts a streaming voice recognition model based on CTC plus Attention, which is a deep learning model, and the voice recognition model includes an encoder, a CTC decoder, and an Attention decoder. The voice audio is processed by audio to obtain a frame-level fbank feature sequence, and the frame-level fbank feature sequence is then sequentially input into the encoder after undergoing an Embedding operation, adding a position code, and downsampling operations. The encoder outputs a hidden state feature sequence and inputs it into the CTC decoder and the Attention decoder for decoding to obtain an output text.
2. A cross-language voice communication method according to claim 1, characterized in that: The called machine translation interface adopts a translation model based on the Transformer architecture that belongs to a deep learning model. The translation model includes a four-layer encoder and a six-layer decoder. The encoder consists of a multi-head attention layer, a residual connection and a normalization layer, a feedforward fully connected layer, a residual connection and a normalization layer connected in sequence, and the decoder consists of a masked multi-head attention layer, a residual connection and a normalization layer, a multi-head attention layer, a residual connection and a normalization layer, a feedforward fully connected layer and a residual connection, and a normalization layer connected in sequence; after the source language text is input into the encoder for feature processing, a set of vectors of fixed length are output, and the decoder translates the vectors output by the encoder and outputs the translated text in the target language.
3. A cross-language voice communication method according to claim 1, characterized in that: When the initiator A and the receiver B conduct a cross-language voice call, the voice call is conducted through a protocol that publicly supports voice communication.
4. The cross-language voice communication method according to claim 3, characterized in that: The protocol publicly supporting voice communication adopts one of the following protocols: http protocol, webrtc protocol, SIP protocol, RTP protocol, XMPP protocol, H.323 protocol and MGCP protocol.
5. A cross-language voice communication system, characterized by: It includes a registration and login module, a background management module, a call visualization module, a speech recognition interface and a machine translation interface. The registration and login module is used to provide registration and login services for the initiator A, the background management module is used to save and verify the relevant information of the initiator A, and the background management module is also used to establish a communication connection after the initiator A initiates the voice communication; after the receiver B receives the voice communication, both the initiator A and the receiver B can call the speech recognition interface and the machine translation interface; the speech recognition interface adopts a streaming speech recognition model based on CTC plus Attention, which is a deep learning model. The streaming speech recognition model based on CTC plus Attention consists of an encoder, a CTC decoder and an Attention decoder. The speech audio is processed by audio to obtain a frame-level fbank feature sequence. The frame-level fbank feature sequence is then input into the encoder after undergoing Embedding operation, adding position coding, and downsampling operations in sequence. The encoder outputs a hidden state feature sequence and inputs it into the CTC decoder and the Attention decoder for decoding to obtain the output text.
6. A cross-language voice communication system according to claim 5, characterized in that: The machine translation interface adopts a translation model based on the Transformer architecture, which is a deep learning model.
7. A cross-language voice communication system according to claim 6, characterized in that: The translation model includes a four-layer encoder and a six-layer decoder, wherein the encoder is composed of a multi-head attention layer, a residual connection and normalization layer, a feedforward fully connected layer, a residual connection and normalization layer connected in sequence, and the decoder is composed of a masked multi-head attention layer, a residual connection and normalization layer, a multi-head attention layer, a residual connection and normalization layer, a feedforward fully connected layer and a residual connection, and a normalization layer connected in sequence; after the source language text is input into the encoder for feature processing, a set of vectors with a fixed length is output, and the decoder translates the vectors output by the encoder and outputs the translated text in the target language.
Citation Information
Patent Citations
Simultaneous interpretation system based on speech recognition technology
CN106486125A
Multilingual real-time conversation platform
CN109218038A
Echo suppression method based on coding and decoding neural network, audio device and equipment
CN111353258A
Neural machine translation model fusing key information based on Transformer model
CN113033153A
Interaction object driving and phoneme processing method and device, equipment and storage medium
CN113314104A