Voice communication method based on pinyin and voiceprint coding and low-power-consumption voice terminal
By converting the voice signal into pinyin and voiceprint coded text information for transmission, the bottlenecks in bandwidth, power consumption and confidentiality of traditional voice communication systems are solved, and voice communication with high data compression ratio, low power consumption and high semantic restoration rate is realized.
Patent Information
- Application Number
- CN202510783218.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-08-15
AI Technical Summary
Traditional voice communication systems have bottlenecks in bandwidth, power consumption and confidentiality, making it difficult to achieve high compression ratio and low power consumption while maintaining high intelligibility and personality reduction.
The voice communication method based on pinyin and voiceprint encoding is adopted to convert the voice signal into standardized text information containing pinyin and four-sound tone markings, and is sent to the receiving end through wireless communication, and then restored to the voice signal from the receiving end, and personalized restore is performed by combining the voiceprint packet data.
It realizes a high data compression ratio (transmission information is only 1% to 5% of the original voice data), low power consumption (average power consumption is reduced by 30 to 50 times), and maintains high semantic restoration rate and personalized sound quality, suitable for multiple languages and environments.
Smart Images

Figure CN120496573A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of wireless communication and voice signal processing, and in particular relates to a voice communication method based on pinyin and voiceprint coding and a low-power voice terminal. Background Art
[0002] Traditional voice communication systems (such as walkie-talkies, VoIP, and cellular voice) generally rely on direct compression of audio waveforms for transmission, which presents bottlenecks in bandwidth, power consumption, and confidentiality. The maturity of existing TTS (text-to-speech) and ASR (speech recognition) technologies has made it possible to convert speech into an intermediate form of text and then reconstruct it.
[0003] As a language with a strict phonetic-phonetic system, Chinese can accurately express the physical characteristics of speech through the combination of pinyin and tones. Based on this characteristic, the present invention proposes a novel voice communication method: "speech → pinyin text + tones + voiceprint encoding + text information such as emotion codes → wireless transmission → speech restoration." This method achieves high compression ratios and low power consumption while maintaining high intelligibility and individuality. Summary of the Invention
[0004] In view of this, the present invention provides a voice communication method based on pinyin and voiceprint coding and a low-power voice terminal to solve the above problems.
[0005] To solve the above technical problems, the present invention provides a voice communication conversion method based on pinyin and voiceprint coding, comprising the following steps:
[0006] S1, converting the voice signal collected by the sending end into standardized text information including pinyin and four tone annotations;
[0007] S2. Sending the text message to the receiving end via wireless communication;
[0008] S3. The receiving end converts the received pinyin text information into a voice signal for playback.
[0009] As an optional method, the speech signal is converted into standardized text information including pinyin and four-tone tones through a speech recognition model, including:
[0010] By using speech semantic recognition software, after processing the ambiguity of multiple characters with the same pronunciation during the online speech-to-text conversion process, the data of one pronunciation and one symbol is obtained as a training set to automatically train and optimize the speech recognition model;
[0011] When it is necessary to enrich the sample dimensions, improve the generalization ability and accuracy of the model, and adapt to polyphonetic characters, polysemy, and voice input with different accents, this training will also be combined with the open speech corpus; the training process automatically constructs the training set through batch recording and pinyin annotation, automatically optimizes the training, and automatically iterates.
[0012] As an option, the standardized text information also includes the speaker's voiceprint data;
[0013] The voiceprint package is generated by collecting a large amount of voice data from the speaker, using professional software running online to extract voiceprint features, and then compressed and encoded through an algorithm. It is uniquely bound to the speaker's voice, and each voiceprint package corresponds to a digital code;
[0014] The voiceprint package and its code are pre-stored in the communication terminal. The voiceprint package data contains the compressed spectrum template, formant parameters, rhythmic features and pronunciation habit parameters; the pre-storage mechanism adopts Flash address index mapping.
[0015] As an optional method, the pinyin text restoration speech process uses a lightweight local TTS module, or performs direct waveform synthesis based on a pre-trained voiceprint package, and performs reverse comparison and optimization to generate results through an online speech recognition system.
[0016] As an optional method, the process of restoring speech from pinyin text is applicable to different languages, and multilingual adaptation is achieved by adjusting the phonetic phoneme spelling system, intonation annotation system and voiceprint expression model.
[0017] On the other hand, the present invention also provides a low-power voice terminal, which uses the above-mentioned voice communication conversion method based on pinyin and voiceprint coding for signal processing. As an optional method, the terminal at least includes a microphone, a voice recognition module, a voiceprint packet decoding module, a text-to-speech module, a speaker playback module, a wireless communication module, and preferably an LDSW communication module;
[0018] The speech recognition module, wireless communication module, voiceprint packet decoding module and text-to-speech module can be integrated into a SoC chip with a Cortex-M4 or other microcontroller with at least similar functions as the core, including the MCU of a smartphone, and support fully offline operation.
[0019] As an optional method, the wireless communication module adopts the low-power sleep wake-up method specified in the YD / T 6334-2025 standard, that is, after waking up from periodic sleep, it will enter a low-power standby state for a moment on a pre-set public monitoring channel to listen for the wake-up work command signal sent by the communication initiator; the wireless transceiver module that initiates voice communication will use the low-power sleep wake-up method specified in the YD / T 6334-2025 communication standard to wake up the wireless communication module of the signal receiver that is in an intermittent low-power working state, so that it can enter a state where it can receive digital information at any time.
[0020] As an optional method, the speech recognition module and TTS module can be deployed online or offline locally;
[0021] When deployed online, the communication channel is cellular network or WiFi, and it supports loading voiceprint data from the cloud to restore personalized voice after receiving the voiceprint package code;
[0022] When deployed offline locally, the speech recognition module, TTS module, and voiceprint package data must be deployed and stored locally; the voiceprint package code includes the mobile phone number, and the personal voiceprint packages of all mobile phone users will be stored in the cloud.
[0023] The beneficial effects of the present invention are:
[0024] The present invention has a high data compression ratio, and the transmitted information is only 1% to 5% of the original voice data; it combines the LDSW protocol for wake-up and transmission, and is suitable for ultra-low power devices; high semantic restoration rate: through the trained "one sound, one word" voice physical feature model, it ensures that both communicating parties have consistent understanding; voiceprint personalization: pre-stores the "voiceprint package" and only transmits its identification code, avoiding repeated transmission and improving the restored sound quality; real-time emotional expression: the speaker's emotions are extracted in real time and transmitted in the form of emotional codes; strong versatility: adapts to multiple languages through language pinyin system conversion and voiceprint training; the present invention changes the existing voice communication mechanism, releases spectrum resources, and reduces system power consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a flowchart of the method of the present invention;
[0026] Figure 2 This is a diagram of the terminal system framework and module structure of the present invention;
[0027] Figure 3 This is a diagram for automatically training a new speech-to-pinyin mathematical model based on an existing ASR model of the present invention;
[0028] Figure 4 The present invention uses a dedicated tool to generate a personal voiceprint package online;
[0029] Figure 5 This is a schematic diagram of the present invention converting a traditional half-duplex intercom into a full-duplex intercom;
[0030] Figure 6 This is a schematic diagram of the LDSW ultra-low power sleep and wake-up function of the present invention. DETAILED DESCRIPTION
[0031] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention is further described in detail below in conjunction with specific implementation methods.
[0032] See also Figures 1-6 This embodiment provides a voice communication conversion method based on pinyin and voiceprint coding, comprising the following steps:
[0033] S1. Convert the voice signal collected by the transmitter into standardized text information including pinyin and four-tone tones. Prior to this step, the sleep wake-up method in the YD / T 6334-2025 low-power communication standard or other low-power wake-up methods can be used to "wake up" the signal receiving communication terminal in an intermittent low-power working state, so that it enters the state of receiving digital information.
[0034] S2. Sending the text information to a receiving end via a low-power wireless communication method;
[0035] S3. The receiving end converts the received text information into a voice signal for playback.
[0036] The standardized text information mentioned above also includes information such as the speaker's voiceprint package. The voiceprint package is generated by extracting the voiceprint features of a large number of speakers' pre-collected audio materials using professional online software tools. After compression encoding (such as VQ vector quantization), it is uniquely bound to the speaker's voice, and each voiceprint package corresponds to a digital code. The voiceprint package and its code are pre-stored in the communication terminal or cloud: the sender transmits the code number, and the receiver calls the voiceprint package stored locally or in the cloud according to the code, combining the pinyin text information with TTS technology to restore the speech. The voiceprint package data here includes the compressed spectral template, formant parameters, prosodic features, and pronunciation habit parameters. The data volume is usually 0.128KB to several MB (depending on the feature dimension, model complexity, and compression method). The pre-storage mechanism uses Flash address index mapping to ensure fast and reliable access in resource-constrained chips such as Cortex-M4. For smartphones with more abundant storage resources, the voiceprint package can support higher-dimensional features (such as multiple emotion patterns or multi-speaker models), further improving the naturalness of speech synthesis.
[0037] It should be noted here that since the extraction of voiceprint packages requires a lot of computing resources, especially "generalized" voiceprint packages, the voiceprint packages we are talking about here, in addition to collecting a large amount of speech data from the speaker, using professional software running online to extract voiceprint features (such as MFCC, resonance peaks), and generating feature vectors or model parameters through algorithm training (such as GMM or deep learning models), can also include more content that is basically unchanged during the speaker's speaking process, such as prosodic features and pronunciation habit parameters. In addition, we train a local offline mathematical model (such as PocketSphinx or TensorFlow Lite Micro) embedded in a voice communication terminal for real-time speech-to-pinyin conversion. This involves using a large amount of actual speech data and existing online speech-to-text semantic recognition models (such as the iFlytek API and the Doubao API) to process the ambiguity of multiple characters with one sound and one symbol (one-to-one corresponding pinyin annotation) to train the model. To further enrich the sample dimensions, improve the generalization and accuracy of the model, and adapt to polyphonetic characters, multiple meanings, and speech input with different accents, the training is also combined with open speech corpora (such as AIShell, THCHS-30, and CommonVoice). The training process automatically constructs a training set through batch recording and pinyin annotation, automatically optimizes the training, and supports automatic iteration, eliminating the need for manual verification of each item.
[0038] For applications in resource-limited voice communication terminals such as intercoms, a lightweight local TTS module (such as eSpeak-NG) is used to restore speech from pinyin text, or direct waveform synthesis is performed based on pre-generated voiceprint packages, and the results are optimized through reverse comparison using an online speech recognition system. For applications on online smartphones, a higher-performance online or offline TTS module can be used.
[0039] It should also be emphasized here that although the method of this embodiment is described in the name of Pinyin, it is applicable to different languages and languages, and multilingual adaptation is achieved by using the International Phonetic Alphabet to adjust the phoneme spelling system, intonation annotation system and voiceprint expression model.
[0040] When the method of this embodiment is specifically implemented, it is necessary to at least include a microphone, a speech recognition module, a wireless communication module, preferably a wireless low-power communication module (LDSW communication module) that meets the YD / T 6334-2025 standard, a voiceprint packet decoding module, a text-to-speech module, and a speaker playback module. The speech recognition module, wireless communication module, voiceprint packet decoding module, and text-to-speech module can be integrated into a SoC chip with a Cortex-M4 or other single-chip microcomputer with at least similar functions as the core, including the MCU of a smartphone, and support complete offline operation. When deployed in the MCU of a smartphone, the speech recognition module and TTS module can be deployed online or offline locally. When deployed online, the communication channel is a cellular network or WiFi, and supports loading voiceprint data from the cloud after receiving the voiceprint package code to restore personalized voice. When deployed offline locally, the speech recognition module and TTS module, as well as the voiceprint package data, must be locally deployed and stored. With the advancement of technology, the voiceprint package code is the mobile phone number, and the personal voiceprint packages of all mobile phone users will be stored in the cloud.
[0041] The wireless transceiver and channel used by the smartphone depend on the connection with the smartphone and can be used by the smartphone, including wireless communication modules and channels that can be used offline, including at least mobile network wireless communication modules and channels used when the smartphone is working online. At this time, the smartphone is an independently working walkie-talkie when offline.
[0042] The LDSW communication module described here uses the low-power sleep wake-up method specified in the YD / T 6334-2025 standard. That is, after waking up from periodic sleep, it will enter a low-power standby state for a moment on a pre-set public monitoring channel to monitor the wake-up work command signal sent by the communication initiator; the wireless transceiver module that initiates voice communication will use the low-power sleep wake-up method specified in the YD / T 6334-2025 communication standard to "wake up" the wireless communication module of the signal receiver that is in an intermittent low-power working state, so that it can enter a state where it can receive digital information at any time.
[0043] When using the method of this embodiment, especially when using the ultra-low power sleep-wake-up technology in the YD / T 6334-2025 standard, the communication data volume of all wireless communication terminals will be reduced by more than 90% (calculated based on a comparison of the voice coding ratio and the average transmission frame length of the LDSW protocol), the average power consumption will be reduced by more than 30 to 50 times (calculated by comparing the power consumption of a typical walkie-talkie with the total power consumption of the terminal standby + wake-up + transmission of this embodiment), and the spectrum reuse efficiency will be increased by more than 10 times (derived by the ratio of the maximum number of simultaneous connections per unit bandwidth to the data occupancy time). The above indicators are obtained through theoretical modeling and experimental simulation tests, and the specific values will be detailed in the embodiments.
[0044] The speech recognition, emotion recognition, voiceprint extraction and text-to-speech modules in the systems or methods described above have cross-platform adaptability and support deployment in resource-constrained embedded SOC chips (such as Cortex-M4), general-purpose smart terminal main chips, and local servers or cloud platforms with edge computing capabilities. They can also switch or trim corresponding functional modules according to the operating environment, including the hardware resources available to the available communication terminals, to ensure flexible deployment and stable operation of the communication system under different power consumption, network conditions and performance requirements.
[0045] If necessary, in order to restore the voice features sent by the sender more realistically and subtly at the receiving end, we can also mark non-habitual accents when sending the voice and restore them at the receiving end.
[0046] Because the method of this embodiment transmits text information during wireless transmission, the data volume is greatly reduced. The content of the transmission becomes standardized, truly digital, and easily confidential. The system is also very simple and easy to deploy. Therefore, in remote areas where resources are scarce due to natural disasters or war, a simple hardware device, even a common smartphone, can be used to quickly establish an emergency rescue communication system, including connection to the Beidou satellite network or other low-orbit satellites. Therefore, this method is suitable for the following three typical applications:
[0047] Offline intercom scenario: Resource-constrained devices (such as voice intercoms and low-power wireless communicators) rely on pinyin, voiceprint package encoding, emotion codes, etc. to achieve effective communication with local TTS;
[0048] Online smartphone scenario: The app completes voice-to-pinyin conversion, voiceprint package encoding, and emotion coding. The information is transmitted via the Internet or cellular network, and the receiving end calls the local or cloud TTS module to synthesize high-fidelity audio.
[0049] Smartphone disconnected from the Internet: Voice to Pinyin + voiceprint package code + emotion code is completed by the App. The information can at least be transmitted through the wireless transceiver channel used when the smartphone is online, and can be used for voice-to-text communication with other smartphones, and then back to voice intercom communication. The advantage of small data volume during wireless transmission of voice-to-text can be used to establish a simple low-power emergency communication system.
[0050] Multiple computing platform adaptation methods: This embodiment system supports three main deployment methods to adapt to different computing resource conditions and communication requirements:
[0051] Embedded intercom terminal deployment mode: With a SoC chip with a Cortex-M4 core or at least equivalent core functions as the core, it integrates voice-to-pinyin conversion, voiceprint packet storage and decoding, simple emotion recognition and local TTS modules, and is adapted to low-power wireless communication modules for resource-constrained and offline operation scenarios;
[0052] Smartphone online operation mode: Utilizes the smartphone app combined with the main chip computing power to complete high-quality speech-to-pinyin conversion, voiceprint packet decoding, and speech synthesis; TTS and ASR can be deployed locally or in the cloud, suitable for Internet or mobile network communication scenarios;
[0053] Smartphone offline operation mode: The above-mentioned smartphone system can be switched to offline operation when the network is disconnected, and a local intercom network is built with the help of the local wireless communication module of the mobile phone (the wireless transceiver of the mobile phone itself for mobile network communication, as well as Bluetooth, WiFi, etc.). The key modules are all deployed locally, which is suitable for communication needs in disaster emergencies and special environments.
[0054] On the other hand, the system part of the voice terminal of this embodiment includes:
[0055] Microphone input module, used to collect the speaker's voice;
[0056] A SoC chip with a Cortex-M4 core or at least a core with equivalent functionality, running the speech-to-phonetic-to-tone-emotion module;
[0057] Local or external TTS module generates speech based on pinyin + voiceprint package + emotional expression;
[0058] Use LoRa wireless transceiver or other wireless communication modules for data transmission;
[0059] The receiving end also deploys voice reconstruction function and voiceprint packet matching module;
[0060] Below we use the application of LDSW ultra-low power artificial intelligence walkie-talkie (LDSW walkie-talkie for short) to illustrate:
[0061] All LDSW radios are normally in a low-power standby state, where they monitor signals on the public monitoring channel for a short period of time after waking up from a periodic sleep cycle. The sleep-wake cycle is set to 1 second. At this time, the average standby current of the LDSW radio is <20uA, which is 1 / 500 of the 10mA standby current of a typical radio. We can assume that the ratio of operating time to standby time for a typical radio is 1:9. Therefore, the radio spends most of its time in a low-power "standby state" of <20uA, ready to receive signals at any time.
[0062] When any LDSW intercom initiates intercom communication with another LDSW intercom, it continuously transmits a wake-up command signal to the other party (identified by ID number) on a public monitoring channel for a period longer than the sleep-wake cycle (1 second). This signal indicates that after 1 second, both parties will begin voice-to-text wireless communication on a new communication channel designated by the command signal, using a certain encryption method. Simultaneously, the speaker's voice information is converted using voice-to-pinyin software embedded in its own SoC chip, along with the speaker's voiceprint packet code. The converted text message is then transmitted to the recipient via the SoC transceiver using a long-distance, high-power (e.g., 1-5 watt) wireless transmission method. Because the transmission time (also known as frequency occupancy time) required to transmit the same voice information is at least 300 times longer than the corresponding text information, the corresponding power consumption is also more than 300 times higher. The power consumption of the SOC chip is also very limited when processing voice-to-pinyin conversion, or when receiving signals, and when using TTS to synthesize pinyin and voiceprint packets to restore the original voice. Therefore, the power consumption of the entire voice information transmission will be reduced by at least 200 times.
[0063] When the receiving end receives the relevant text information, it calls the voiceprint package corresponding to the voiceprint package code and uses the TTS software embedded in the SOC chip to restore the received text information to the original voice.
[0064] Since voice input and voice playback use two different hardware channels, voice input and voice playback can be carried out simultaneously. After the voices of both parties are converted into text messages, they can be transmitted simultaneously in a time-division manner, thus realizing low-cost and ultra-low power dual-function intercom communication.
[0065] Because the entire LDSW intercom communication uses very small amounts of wireless data and consumes very low power, we can easily use LDSW intercoms or even ordinary offline smartphones to quickly establish a low-cost, narrowband wireless communication network for emergency rescue that can relay "voice" information and connect to the BeiDou-3 satellite network and other ground-orbit satellites.
[0066] The above are merely preferred embodiments of the present invention. It should be noted that the above preferred embodiments should not be construed as limiting the present invention, and the scope of protection of the present invention should be determined by the scope defined in the claims. Persons skilled in the art will appreciate that improvements and modifications may be made without departing from the spirit and scope of the present invention, and such improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A voice communication conversion method based on pinyin and voiceprint coding, characterized in that: The following steps are involved: S1, converting the voice signal collected by the sending end into standardized text information including pinyin and four tone annotations; S2. Sending the text message to a receiving end via wireless communication; S3. The receiving end converts the received pinyin text information into a voice signal for playback.
2. The method for voice communication conversion based on pinyin and voiceprint coding according to claim 1, characterized in that: The speech recognition model converts speech signals into standardized text information including pinyin and four-tone annotations, including: By using speech semantic recognition software, after processing the ambiguity of multiple characters with the same pronunciation during the online speech-to-text conversion process, the data of one pronunciation and one symbol is obtained as a training set to automatically train and optimize the speech recognition model; When it is necessary to enrich the sample dimensions, improve the generalization ability and accuracy of the model, and adapt to polyphonetic characters, polysemy, and voice input with different accents, this training will also be combined with the open speech corpus; the training process automatically constructs the training set through batch recording and pinyin annotation, automatically optimizes the training, and automatically iterates.
3. The method for voice communication conversion based on pinyin and voiceprint coding according to claim 1, characterized in that: The standardized text information also includes the speaker's voiceprint data; The voiceprint package is generated by collecting a large amount of voice data from the speaker, using professional software running online to extract voiceprint features, and then compressed and encoded through an algorithm. It is uniquely bound to the speaker's voice, and each voiceprint package corresponds to a digital code; The voiceprint package and its code are pre-stored in the communication terminal, and the voiceprint package data includes a compressed spectrum template, formant parameters, rhythmic features and pronunciation habit parameters; The pre-storage mechanism uses Flash address index mapping.
4. The method for voice communication conversion based on pinyin and voiceprint coding according to claim 1, characterized in that: The process of restoring phonetic text to speech uses a lightweight local TTS module, or performs direct waveform synthesis based on a pre-trained voiceprint package, and performs reverse comparison and optimization to generate results through an online speech recognition system.
5. The method for voice communication conversion based on pinyin and voiceprint coding according to claim 4, characterized in that: The process of restoring speech from pinyin text is applicable to different languages, and multilingual adaptation is achieved by adjusting the phonetic phoneme spelling system, intonation annotation system and voiceprint expression model.
6. A low-power voice terminal, which uses the voice communication conversion method based on pinyin and voiceprint coding according to any one of claims 1 to 5 for signal processing, characterized in that: The terminal includes at least a microphone, a voice recognition module, a voiceprint packet decoding module, a text-to-speech module, a speaker playback module, a wireless communication module, preferably an LDSW communication module; The speech recognition module, wireless communication module, voiceprint packet decoding module and text-to-speech module can be integrated into a SoC chip with a Cortex-M4 or other single-chip microcomputer with at least similar functions as the core, including the MCU of a smartphone, and support fully offline operation.
7. A low-power voice terminal according to claim 6, characterized in that: The wireless communication module adopts the low-power sleep wake-up method specified in the YD / T 6334-2025 standard, that is, after waking up from periodic sleep, it will enter a low-power standby state for a moment on a pre-set public monitoring channel to monitor the wake-up work command signal sent by the communication initiator; the wireless transceiver module that initiates voice communication will use the low-power sleep wake-up method specified in the YD / T6334-2025 communication standard to wake up the wireless communication module of the signal receiver that is in an intermittent low-power working state, so that it can enter a state where it can receive digital information at any time.
8. The low-power voice terminal according to claim 7, characterized in that: The speech recognition module and TTS module can be deployed online or offline locally; When deployed online, the communication channel is cellular network or WiFi, and it supports loading voiceprint data from the cloud to restore personalized voice after receiving the voiceprint package code; When deployed offline locally, the speech recognition module, TTS module, and voiceprint package data must be deployed and stored locally; the voiceprint package code includes the mobile phone number, and the personal voiceprint packages of all mobile phone users will be stored in the cloud.
Citation Information
Patent Citations
Voice data compression method and device, electronic equipment and readable storage medium
CN117292697A
End-to-end encrypted voice communication method
CN118785149A
Cited By
Voice communication method based on LDSW voice compression technology and smart phone
CN120954420A