Novel radio station based on voice recognition and translation

The new radio station based on speech recognition and translation solves the problem of digital radio being unable to achieve multilingual communication and large bandwidth usage, and realizes barrier-free long-distance real-time communication across languages.

CN120673759APending Publication Date: 2025-09-19CHINESE PEOPLES LIBERATION ARMY UNIT 93995
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510934522.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In the existing technology, digital radio data occupies a large bandwidth, making it impossible to communicate in multiple languages. In addition, the data compression rate during communication is limited, and the bandwidth is large.

Method used

A new radio based on speech recognition and translation is used. The voice processing module recognizes and translates the speaker's voice, generates corresponding language coding information, and forms a radio frequency signal through digital modulation and mixing amplification for transmission. The receiving end demodulates and translates the voice and outputs the voice information in the target language.

Benefits of technology

It occupies much less bandwidth than traditional radio stations at the same speaking speed, and can achieve barrier-free communication across languages, making it suitable for long-distance real-time communication between various languages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673759A_ABST
    Figure CN120673759A_ABST
Patent Text Reader

Abstract

The invention discloses a novel radio station based on voice recognition and translation, which is characterized in that voice recognition is carried out on the voice of a sender to generate corresponding language coding information, the information is modulated by a digital modulation module and then forms a radio frequency signal through a frequency mixing amplification module, and the radio frequency signal is transmitted to the air through an antenna; after receiving a radio frequency signal through an antenna, the receiving station generates an intermediate frequency digital modulation signal through a frequency mixing amplification module, then demodulates the intermediate frequency digital modulation signal into language coding information of a speaker of the transmitting station through a digital demodulation module, and forms language digital voice information corresponding to a listener of the receiving station after the language coding information is processed by a voice translation module in a voice processing module. The digital-to-analog signal conversion is realized through the D / A conversion module, the analog signal is sent to the earphone through the power amplification module to form a voice message which can be understood by a receiver, the occupied bandwidth is far smaller than that of a traditional radio station, and meanwhile, barrier-free cross-language communication can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of wireless voice communication technology, and in particular to a new radio station based on voice recognition and translation. Background Art

[0002] A radio station, or simply a "radio," is a wireless terminal device used for point-to-multipoint communication. Compared to wired communications, it offers greater mobility, cost-effectiveness, and lower operating costs. Therefore, it is widely used in both military and civilian fields, such as broadcasting stations and military airborne radios. Radio stations are categorized by their technology into analog and digital. Analog stations typically use fixed-frequency communication, which is simple to implement but susceptible to interference. Digital stations, however, are increasingly popular due to their widespread use of frequency-hopping and spread spectrum technologies, which offer improved interference resistance compared to fixed-frequency communication.

[0003] Radio transmission requires a certain amount of electromagnetic signal bandwidth, which is positively correlated with the amount of voice signal data being transmitted. However, bandwidth is a limited resource, and people always prefer to minimize the bandwidth used for communication. Therefore, radio transmissions often use voice compression to significantly reduce the data volume before wireless transmission. The receiving station decompresses the compressed signal and then restores the voice, thus enabling voice communication while utilizing less bandwidth.

[0004] Although digital radio can achieve voice communication through compression and decompression, the following problems still exist:

[0005] 1) Limited data compression rate and large bandwidth occupation

[0006] Although the MELPe model can reduce voice rates to a minimum of 600 bits per second (bps), it has not been used in practice due to the complexity of the algorithm. The currently used Global System for Mobile Communications (GSM) voice coding rate is 13 kbps, and after channel coding, the rate is 22.8 kbit / s. Satellite communication voice rates can reach a minimum of 1.2 kbps. Although this compression significantly reduces the original voice data rate, in actual use, the data rate is still above 1000 bps.

[0007] 2) It is impossible to communicate in multiple languages

[0008] When using radio for voice communication, at least one of the two communicating parties needs to be familiar with and master the other party's language in order to conduct normal voice communication. However, with the increase in international cooperation, timely communication between different languages ​​is often required. For example, international joint military exercises and cross-regional flights require communicators to master multiple languages. When they cannot understand each other's language, they need to contact an intermediary for translation. This greatly hinders communication between people who speak different languages ​​and reduces the efficiency of radio use. Summary of the Invention

[0009] The main purpose of this application is to provide a new type of radio based on speech recognition and translation, aiming to propose a radio that occupies much less bandwidth than traditional radios and can achieve barrier-free communication across languages, and is suitable for long-distance real-time communication between various languages.

[0010] To achieve the above-mentioned purpose, the present application provides a new radio station based on speech recognition and translation, comprising: a transmitting module, a receiving module and a language control module; the transmitting module comprises an analog signal receiving module, a first speech processing module and a signal transmitting module which are communicatively connected in sequence; the receiving module comprises a digital signal receiving module, a second speech processing module and a signal playing module which are communicatively connected in sequence; the language control module is communicatively connected to the first speech processing module and the second speech processing module respectively, and the language control module is used to control the language conversion mode of the first speech processing module and the second speech processing module; the transmitting module is used to convert the speaker's speech into an amplified analog electrical signal through the analog signal receiving module and then input it into the first speech processing module to generate a code carrying the speaker's corresponding language The signal transmitting module converts the digital coded signal into a radio frequency signal and transmits it through an antenna; the receiving module is used to receive the radio frequency signal through the digital signal receiving module and demodulate the digital coded information carrying the speaker's language coding information, and use the second voice processing module to translate the digital coded information into digital voice information of the receiving station's target language; the signal playing module converts the digital voice information into an amplified analog signal for output; the language control module is used to obtain the speaker's language coding information and the receiving station's target language information, and control the first voice processing module to identify the speaker's original language in the digital coded signal according to the speaker's language coding information, and control the second voice processing module to translate the speaker's original language into the target language according to the target language information.

[0011] Optionally, the transmitting module includes a microphone and a low-noise amplification module; the first voice processing module includes an A / D conversion module and a voice recognition module; the signal transmitting module includes a digital modulation module and a first frequency mixing and amplification module; the transmitting module is used to convert the speaker's voice into an analog electrical signal through the microphone, and the analog electrical signal is amplified by the low-noise amplification module and input into the A / D conversion module to be converted into a digital signal; the voice recognition module recognizes the digital signal and generates digital coding information carrying coding information corresponding to the speaker's language; the digital modulation module modulates the digital coding information into a baseband signal, and the baseband signal is converted into a radio frequency signal by the first frequency mixing and amplification module and then transmitted through the antenna.

[0012] Optionally, the digital signal receiving module includes a second frequency mixing and amplifying module and a digital demodulation module; the second voice processing module includes a voice translation module and a D / A conversion module; the signal playing module includes a power amplification module and headphones; wherein the receiving module receives the radio frequency signal through the antenna and inputs it into the second frequency mixing and amplifying module to convert the radio frequency signal into an intermediate frequency digital modulation signal; the digital demodulation module demodulates the intermediate frequency digital modulation signal to obtain digital coding information carrying the language coding information of the speaker at the transmitting station; the voice translation module translates the digital coding information into digital voice information of the target language of the receiving station; the D / A conversion module converts the digital voice information into an analog signal, which is amplified by the power amplification module and output through the headphones.

[0013] Optionally, the language control module has two working modes, including a same-language mode and a cross-language mode, wherein the cross-language mode includes two or more mainstream languages.

[0014] Optionally, the encoding information includes but is not limited to: using GB2312 encoding for Chinese, using ASCII encoding for English, or performing custom encoding on language words to reduce the transmission data rate.

[0015] Optionally, the coded information transmission rate does not exceed 64 bit / s, and the voice compression data rate is lower than 1000 bps.

[0016] Optionally, the digital modulation module and the digital demodulation module adopt spread spectrum or frequency hopping technology to enhance the communication anti-interference capability; the first mixing and amplifying module and the second mixing and amplifying module are used to realize frequency conversion and power amplification of baseband signals and radio frequency signals.

[0017] The embodiment of the present application proposes a new type of radio based on speech recognition and translation, wherein the transmitting module includes an analog signal receiving module, a first speech processing module and a signal transmitting module which are communicatively connected in sequence; the receiving module includes a digital signal receiving module, a second speech processing module and a signal playing module which are communicatively connected in sequence; the language control module is communicatively connected to the first speech processing module and the second speech processing module respectively, and the language control module is used to control the language conversion mode of the first speech processing module and the second speech processing module; the transmitting module is used to convert the speaker's speech into an amplified analog electrical signal through the analog signal receiving module and then input the signal into the first speech processing module to generate a digital coded signal carrying the coding information of the speaker's corresponding language, and the signal transmitting module converts the digital coded signal into a radio frequency signal and transmits it through an antenna; the receiving module is used to receive the radio frequency signal through the digital signal receiving module and then decode the signal The digital coding information carrying the speaker's language coding information is called out, and the second voice processing module is used to translate the digital coding information into digital voice information of the target language of the receiving station. The signal playback module converts the digital voice information into an amplified analog signal for output; the language control module is used to obtain the speaker's language coding information and the target language information of the receiving station, and control the first voice processing module to identify the speaker's original language in the digital coding signal according to the speaker's language coding information, and control the second voice processing module to translate the speaker's original language into the target language according to the target language information. Compared with traditional radio stations, at the same speaking speed, this application proposes a method that occupies much less bandwidth than traditional radio stations and can achieve barrier-free communication across languages. It is suitable for long-distance real-time communication between various different languages, such as international joint exercise command and cross-international regional flight voice communication. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 A block diagram of the principles of a novel radio station provided in accordance with an embodiment of the present application, which is based on speech recognition and translation;

[0019] Figure 2 This is a flowchart of the Chinese-English communication processing of a new radio station provided in an embodiment of the present application based on speech recognition and translation.

[0020] The realization of the objectives, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0021] It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application.

[0022] In order to overcome the shortcomings of existing digital radio stations that occupy a large bandwidth and cannot realize communication between multiple languages, the present application provides a new radio station based on speech recognition and translation. By performing speech recognition on the sender's voice, the corresponding language coding information is generated, which is modulated by a digital modulation module and then formed into a radio frequency signal by a mixing and amplification module and transmitted into the air through an antenna; after the receiving station receives the radio frequency signal through the antenna, it generates an intermediate frequency digital modulation signal after passing through the mixing and amplification module, and then demodulated into the language coding information of the speaker at the transmitting station by a digital demodulation module. After being processed by the speech translation module in the speech processing module, digital voice information in the language corresponding to the listener at the receiving station is formed, and the digital-to-analog signal conversion is realized by the D / A conversion module. The analog signal is sent to the earphone through the power amplification module to form voice information that the receiver can understand.

[0023] This application discloses a novel radio station based on speech recognition and translation, which may include the following modules:

[0024] The transmitting module 1 includes a transmitting module 10, a first voice processing module 20, and a signal transmitting module 30, which are communicatively connected in sequence. The receiving module 2 includes a digital signal receiving module 70, a second voice processing module 60, and a signal playing module 50, which are communicatively connected in sequence. The language controlling module 3 is communicatively connected to the first voice processing module 20 and the second voice processing module 60, respectively, and is used to control the language conversion mode of the first voice processing module 20 and the second voice processing module 60. The transmitting module 1 is used to convert the speaker's voice into an amplified analog electrical signal through the analog signal receiving module 10, and then input it into the first voice processing module 20 to generate a digital coded signal carrying the coding information of the speaker's corresponding language. The signal transmitting module 30 converts the digital coded signal into a radio frequency signal and transmits it through an antenna. The receiving module 2 is used to receive the radio frequency signal through the digital signal receiving module 70 and demodulate the digital coded information carrying the speaker's language coding information. The coding information includes, but is not limited to, GB2312 encoding for Chinese, ASCII encoding for English, or custom encoding of language words to reduce the transmission data rate. The coded information transmission rate does not exceed 64 bits / s, and the voice compression data rate is less than 1000 bps. The second voice processing module 60 translates the digital coded information into digital voice information in the receiving station's target language. The signal playback module 50 converts the digital voice information into an amplified analog signal for output. The language control module 33 is used to obtain the speaker's language coding information and the receiving station's target language information. Based on the speaker's language coding information, it controls the first voice processing module 20 to identify the speaker's original language in the digital coded signal and, based on the target language information, controls the second voice processing module 60 to translate the speaker's original language into the target language. The language control module 3 has two operating modes: a same-language mode and a cross-language mode. The cross-language mode includes two or more mainstream languages.

[0025] Furthermore, the transmitting module includes a microphone and a low noise amplifier module; the first voice processing module 20 includes an A / D conversion module and a voice recognition module; the signal transmitting module 30 includes a digital modulation module and a first mixing amplifier module;

[0026] The transmitting module converts the speaker's speech from microphone 101 into an analog electrical signal. This analog signal is amplified by the low-noise amplifier module and then fed into the A / D conversion module for conversion into a digital signal. The speech recognition module identifies the digital signal and generates digital coded information containing the speaker's language. The digital modulation module modulates the digital coded information into a baseband signal, which is then converted into a radio frequency signal by the first mixer-amplifier module and transmitted via the antenna. The digital modulation module and digital demodulation module utilize spread spectrum or frequency hopping technology to enhance communication anti-interference capabilities. The first and second mixer-amplifier modules perform frequency conversion and power amplification between the baseband and radio frequency signals.

[0027] Furthermore, the digital signal receiving module 70 includes a second frequency mixing and amplifying module 206 and a digital demodulation module 205; the second voice processing module 60 includes a voice translation module 204 and a D / A conversion module 203; the signal playing module 50 includes a power amplification module 202 and an earphone 201; wherein, the receiving module receives the radio frequency signal through the antenna and inputs it into the second frequency mixing and amplifying module 206, which converts the radio frequency signal into an intermediate frequency digital modulation signal; the digital demodulation module 205 demodulates the intermediate frequency digital modulation signal to obtain digital coding information carrying the language coding information of the speaker at the transmitting station; the voice translation module 204 translates the digital coding information into digital voice information of the target language of the receiving station; the D / A conversion module 203 converts the digital voice information into an analog signal, which is amplified by the power amplification module 202 and output through the earphone 201.

[0028] This application provides a new radio station based on speech recognition and translation, including two parts: transmitting mode and receiving mode (such as Figure 1 In the transmitting mode, the microphone 101 converts the speaker's voice into an analog electrical signal. After amplification by the low-noise amplifier module 102, the signal is sent to the A / D converter module 103 in the first speech processing module 20 to convert the analog signal into a digital signal. The signal is then recognized by the speech recognition module 104 as the language encoding information corresponding to the speaker. The encoded information is modulated into the baseband signal required for transmission by the digital modulation module 105. The signal is then converted into a radio frequency signal by the first mixer-amplifier module 106 and sent to the antenna for transmission into the air. In the receiving mode, the antenna sends the received high-frequency small signal to the second mixer-amplifier module 206, which outputs an intermediate frequency digital modulated signal. The digital demodulator module 205 demodulates the signal into the language encoding information of the transmitting station speaker. After processing by the speech translation module 204 in the first speech processing module 60, the signal is converted into digital speech information in the language corresponding to the receiving station listener. The signal is then converted from digital to analog by the D / A converter module 203 and sent to the earphone 201 through the power amplifier module 202 to form speech information that the receiver can understand.

[0029] In one embodiment of the present application, the digital signal receiving module 70 includes a second mixing and amplifying module 206 and a digital demodulation module 205; the second voice processing module 60 includes a voice translation module 204 and a D / A conversion module 203; the signal playing module 50 includes a power amplification module 202 and an earphone 201; wherein, the receiving module receives the radio frequency signal through the antenna and inputs it into the second mixing and amplifying module 206, which converts the radio frequency signal into an intermediate frequency digital modulation signal; the digital demodulation module 205 demodulates the intermediate frequency digital modulation signal to obtain digital coding information carrying the language coding information of the speaker at the transmitting station; the voice translation module 204 translates the digital coding information into digital voice information in the target language of the receiving station; the D / A conversion module 203 converts the digital voice information into an analog signal, which is amplified by the power amplification module 202 and output through the earphone 201.

[0030] In one embodiment of the present application, the digital modulation module 105 and the digital demodulation module 205 use spread spectrum or frequency hopping technology to enhance communication anti-interference capability. The first mixer-amplifier module 106 and the second mixer-amplifier module 206 are used to implement frequency conversion and power amplification of baseband signals and radio frequency signals.

[0031] Furthermore, the speech processing module is based on an artificial intelligence model and achieves speech recognition and translation through dedicated training.

[0032] Furthermore, the language encoding information can be used to customize the language words to reduce the transmission data rate.

[0033] Furthermore, the language control module 3 can select the same world mainstream language mode. In this mode, the speech translation module 204 does not need to perform translation.

[0034] The following is an explanation of a new radio processing flow based on speech recognition and translation combined with Chinese and English communication. Figure 2 As shown in the Chinese-English communication processing flow chart, the user of radio station A speaks Chinese and the user of radio station B speaks English.

[0035] Before starting a call, radio station A places the language control module 3 in Chinese-English mode, the radio station A speech recognition module 104 in Chinese recognition mode, and the radio station A speech translation module 204 in English-to-Chinese translation mode; radio station B places the language control module 3 in English-Chinese mode, the radio station B speech recognition module 104 in English recognition mode, and the radio station B speech translation module 204 in Chinese-to-English translation mode.

[0036] When a call starts, radio A is in sending mode and radio B is in receiving mode.

[0037] Radio Station A uses microphone 101 to convert the speaker's Chinese voice into a small analog signal. After low-noise amplification, it is sent to Radio Station A's voice processing module for A / D conversion to a digital signal. Voice recognition in Radio Station A's voice processing module then generates the GB2312 encoding value corresponding to Radio Station A's Chinese voice. This encoding value is digitally modulated and sent to the mixer-amplifier module to form an RF signal, which is then wirelessly transmitted via the antenna. Since the normal speaking speed of Chinese is 3-4 characters per second, and the GB2312 encoding length of each character is 16 bits, recognizing Chinese as GB2312 encoding and transmitting it requires no more than 64 bits / s at normal speaking speeds to achieve voice information transmission. This significantly reduces the data rate compared to traditional digital radio stations.

[0038] Radio station B receives the high-frequency small signal sent by radio station A through the antenna and outputs the intermediate frequency modulated signal to the mixing amplifier. After digital demodulation, the GB2312 code corresponding to the Chinese characters is obtained. The voice translator in the voice processing module of radio station B translates the GB2312 code to generate the corresponding English audio stream. The audio stream is power amplified after D / A digital-to-analog conversion and sent to earphone 201, thereby expressing the Chinese voice information of the user of radio station A into English voice that can be understood by the user of radio station B.

[0039] The communication mode processing flow for Radio B sending mode and Radio A receiving mode is similar to that for Radio A sending mode and Radio B receiving mode.

[0040] The beneficial effects of the present application are as follows: The present application provides a new type of radio based on speech recognition and translation. Compared with traditional radios, at the same speaking speed, the bandwidth occupied by the present application is far less than that of traditional radios. At the same time, it can realize barrier-free communication between languages ​​and is suitable for long-distance real-time communication between various languages, such as international joint exercise command and cross-international regional flight voice communication.

[0041] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A new radio station based on speech recognition and translation, characterized in that: include: Transmitting module, receiving module and language control module; The transmitting module includes an analog signal receiving module, a first voice processing module and a signal transmitting module which are communicatively connected in sequence; The receiving module includes a digital signal receiving module, a second voice processing module and a signal playing module which are communicatively connected in sequence; The language control module is communicatively connected to the first voice processing module and the second voice processing module respectively, and the language control module is used to control the language conversion mode of the first voice processing module and the second voice processing module; The transmitting module is used to convert the speaker's voice into an amplified analog electrical signal through the analog signal receiving module and input it into the first voice processing module to generate a digital coded signal carrying the coding information of the speaker's corresponding language. The signal transmitting module converts the digital coded signal into a radio frequency signal and transmits it through the antenna; The receiving module is used to receive the radio frequency signal through the digital signal receiving module and demodulate the digital coded information carrying the speaker's language coding information, use the second voice processing module to translate the digital coded information into digital voice information in the target language of the receiving station, and the signal playing module converts the digital voice information into an analog signal after amplification and output; The language control module is used to obtain the speaker's language coding information and the target language information of the receiving station, and control the first voice processing module to identify the speaker's original language in the digital coding signal according to the speaker's language coding information, and control the second voice processing module to translate the speaker's original language into the target language according to the target language information.

2. The novel radio station based on speech recognition and translation as claimed in claim 1, characterized in that: The transmitting module includes a microphone and a low-noise amplifier module; the first voice processing module includes an A / D conversion module and a voice recognition module; the signal transmitting module includes a digital modulation module and a first mixing amplifier module; The transmitting module is used to convert the speaker's voice into an analog electrical signal through the microphone. The analog electrical signal is amplified by the low-noise amplifier module and then input into the A / D conversion module to convert it into a digital signal. The speech recognition module recognizes the digital signal and generates digital code information carrying the code information of the speaker's corresponding language. The digital modulation module modulates the digital coding information into a baseband signal, and the first mixer amplification module converts the baseband signal into a radio frequency signal before transmitting it through the antenna.

3. The novel radio station based on speech recognition and translation as claimed in claim 2, characterized in that: The digital signal receiving module includes a second mixing and amplifying module and a digital demodulation module; the second voice processing module includes a voice translation module and a D / A conversion module; the signal playing module includes a power amplifying module and headphones; Among them, the receiving module receives the radio frequency signal through the antenna and inputs it into the second mixing and amplifying module, which converts the radio frequency signal into an intermediate frequency digital modulation signal; the digital demodulation module demodulates the intermediate frequency digital modulation signal to obtain digital coding information carrying the language coding information of the speaker at the transmitting station; the speech translation module translates the digital coding information into digital speech information of the target language of the receiving station; the D / A conversion module converts the digital speech information into an analog signal, which is amplified by the power amplifier module and output through the headphones.

4. The novel radio station based on speech recognition and translation as claimed in claim 1, characterized in that: The language control module has two working modes, including a same-language mode and a cross-language mode, wherein the cross-language mode includes two or more mainstream languages.

5. The novel radio station based on speech recognition and translation as claimed in claim 4, characterized in that: The encoding information includes but is not limited to: using GB2312 encoding for Chinese, using ASCII encoding for English, or performing custom encoding on language words to reduce the transmission data rate.

6. The novel radio station based on speech recognition and translation as claimed in claim 5, characterized in that: The coded information transmission rate shall not exceed 64 bit / s, and the voice compression data rate shall be lower than 1000 bps.

7. The novel radio station based on speech recognition and translation as claimed in claim 2, characterized in that: The digital modulation module and the digital demodulation module adopt spread spectrum or frequency hopping technology to enhance the communication anti-interference capability; The first mixer-amplifier module and the second mixer-amplifier module are used to realize frequency conversion and power amplification of baseband signals and radio frequency signals.