Communication methods, devices and systems based on voiceprint recognition and speech reconstruction

By employing local voiceprint recognition and voice reconstruction technologies, the problems of voiceprint leakage and bandwidth are solved, achieving a highly secure and low-bandwidth call experience.

CN122090850APending Publication Date: 2026-05-26深圳市南山区明涵软硬件技术服务工作室
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
深圳市南山区明涵软硬件技术服务工作室
Filing Date
2026-02-12
Publication Date
2026-05-26

Smart Images

  • Figure CN122090850A_ABST
    Figure CN122090850A_ABST
Patent Text Reader

Abstract

This invention discloses a communication method, apparatus, and system based on voiceprint recognition and speech reconstruction, relating to the field of communication security technology. The communication apparatus based on voiceprint recognition and speech reconstruction includes a voice acquisition module, a storage module, a voice conversion module, a data transmission module, a data reception module, a speech synthesis module, and a voice playback module, all connected to a central processing module. The voice acquisition module is connected to the storage module, the voice conversion module is connected to the speech synthesis module and the data reception module, and the speech synthesis module is connected to the voice playback module. The communication method of this invention includes: S1. Voiceprint extraction, S2. Speech-to-text conversion, S3. Data transmission, S4. Data reception, and S5. Text-to-speech conversion and playback. This invention solves the problems of high risk of voiceprint biometric leakage and difficulty in balancing natural call quality and security in existing technologies. It achieves encryption while preventing voiceprint feature transmission, maintaining a natural call experience, and significantly reducing bandwidth requirements due to the small data volume.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of communication security technology, specifically relating to voice call encryption technology, voice processing technology, and privacy protection technology, and particularly to a communication method, device, and system that performs voice-to-text conversion locally on both parties' terminals and combines it with voiceprint features to achieve secure communication. Background Technology

[0002] With the widespread use of mobile communication and Voice over Internet Protocol (VoIP), call privacy and security issues have become increasingly prominent, and the data traffic generated during calls is enormous. Existing technologies primarily employ voice call encryption schemes, including: End-to-end encryption of the voice waveform can be applied directly, such as using Signal Protocol or ZRTP in services like Signal and WhatsApp to encrypt the voice stream. This method can prevent man-in-the-middle eavesdropping, but if the key is leaked, an interceptor can still reconstruct the original voice waveform, thereby revealing the speaker's voiceprint and other biometric information.

[0003] Speech obfuscation or voice alteration, such as through spectrum inversion, frequency band segmentation and reconstruction, or the addition of masking noise to achieve speech blurring (related patents such as US7184952B2, etc.), can reduce intelligibility to some extent, but usually significantly reduce the naturalness of the call and the sound quality, and advanced reverse engineering algorithms may still be able to partially restore the original speech and voiceprint.

[0004] The transmitted compressed speech feature parameters (such as Mel spectrum, fundamental frequency, etc.) are then used to synthesize speech at the receiving end via a vocoder or neural network. This approach can reduce bandwidth usage, but the parameters themselves still contain reversible voiceprint information, posing a risk of reverse analysis to extract voiceprints, and the naturalness of the synthesized speech is limited.

[0005] Plain text communication solutions (such as encrypted instant messaging) completely avoid voice transmission, but they completely lose the naturalness and real-time emotional expression of voice calls, and cannot meet users' needs for voice communication.

[0006] The common shortcomings of existing technologies are: It is difficult to completely eliminate the risk of leakage of biometric information such as voiceprints in the transmission link; It is difficult to strike a good balance between security, bandwidth efficiency, and call naturalness; Some solutions rely on cloud processing, which poses risks of data leaving the domain and privacy compliance (such as violating GDPR and the data minimization principle in China's Personal Information Protection Law). The data traffic generated during a call is enormous, requiring a huge amount of bandwidth.

[0007] To this end, we propose a communication method, device, and system based on voiceprint recognition and speech reconstruction. Summary of the Invention

[0008] To address the shortcomings of existing technologies, the present invention aims to provide a communication method, device, and system based on voiceprint recognition and speech reconstruction. This invention addresses the problems of high risk of voiceprint biometric leakage and the difficulty in balancing natural call quality and security in existing technologies. It provides a communication method, device, and system based on voiceprint recognition and speech reconstruction that achieves high-strength encryption while completely preventing voiceprint feature transmission, maintaining a natural call experience, and significantly reducing text transmission data traffic and bandwidth requirements.

[0009] The objective of this invention can be achieved through the following technical solutions: A communication method based on voiceprint recognition and speech reconstruction, comprising a voiceprint extraction stage and a formal communication stage, includes the following steps: S1. Voiceprint Extraction: After communication is established, in the voiceprint extraction stage, the voice acquisition module collects the voice samples of the peer user, extracts the voiceprint feature vector of the peer user and stores it in the storage module; S2. Voice to Text: During the formal communication phase, the voice conversion module converts the user's voice stream into text data in real time. S3. Data Transmission: The data transmission module sends text data to the other end via the network; S4. Data Reception: Simultaneously, during the formal communication phase, the data receiving module receives text data from the other end via the network; S5. Text-to-speech and playback: The speech synthesis module calls the identity information of the peer user in the storage module, fuses the text data with the peer user's voiceprint feature vector in the storage module, synthesizes a speech stream with the peer user's timbre features, and plays it through the speech playback module.

[0010] Furthermore, a communication method based on voiceprint recognition and speech reconstruction also includes the following steps: S2-3. Text Data Processing: The central processing module processes the text data to generate secure data; Step S2-3 is between steps S2 and S3; S3′. Data Transmission: The data transmission module sends secure data to the peer via the network; Step S3′ replaces step S3; S4′. Data reception: Simultaneously, during the formal communication phase, the data receiving module receives secure data from the other end via the network; Step S4' replaces step S4; S4-5. Security Data Processing: The central processing module processes the security data to obtain text data; Steps S4-5 are located between steps S4 and S5.

[0011] Furthermore, a communication method based on voiceprint recognition and speech reconstruction includes a central processing module that processes secure data, including encryption and / or compression, decryption and / or decompression.

[0012] The present invention also provides a communication device based on voiceprint recognition and speech reconstruction, which includes a central processing module, a voice acquisition module, a storage module, a voice conversion module, a data transmission module, a data reception module, a voice synthesis module, and a voice playback module. The central processing module is communicatively connected to the voice acquisition module, the storage module, the voice conversion module, the data transmission module, the data reception module, the voice synthesis module, and the voice playback module, respectively. The voice acquisition module is communicatively connected to the storage module. The voice conversion module is communicatively connected to the voice synthesis module and the data reception module, respectively. The voice synthesis module is communicatively connected to the voice playback module.

[0013] The present invention also provides a communication system based on voiceprint recognition and speech reconstruction, comprising at least two communication devices as described above, wherein the communication devices are each other’s communication counterparts.

[0014] Furthermore, the voiceprint extraction stage adopts a two-way interactive protocol: the two systems alternately or simultaneously prompt the user to read the same preset or random text, record their own local audio and extract voiceprint feature vectors, and complete two-way authentication by comparing the text content sent and received.

[0015] Furthermore, both the speech acquisition module and the speech synthesis module are locally deployed lightweight neural network models, so the raw speech data does not need to be uploaded to the cloud.

[0016] Furthermore, the voiceprint feature vector is used as a conditional input to control the timbre, prosody, and intonation features of the speech synthesis model.

[0017] The present invention also provides a communication system based on voiceprint recognition and speech reconstruction. The communication system includes: a memory storing a computer program; and a processor communicatively connected to the memory. When the computer program is executed by the processor, it implements any of the methods described above.

[0018] Furthermore, the speech conversion module also includes a sentiment analysis module, which generates emotion, stress, and speech rate tags.

[0019] The present invention also provides a communication device readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the methods described above.

[0020] Compared with the prior art, the beneficial effects of the present invention are: Voiceprint features are always stored locally on the terminal and are never transmitted over the network, thus fundamentally eliminating the risk of voiceprint biometric leakage. The transmitted content consists only of encrypted text data, which is very small in size, thus significantly reducing bandwidth usage. Through local personalized voice synthesis, the timbre and rhythm characteristics of the other end user are preserved, resulting in a high degree of naturalness in the call; The entire process is localized, adhering to the principles of data minimization and privacy-preserving design, without relying on cloud servers. Attached Figure Description

[0021] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings.

[0022] Figure 1 This is a schematic diagram of the communication device structure of the present invention; Figure 2 This is a schematic diagram of the overall system architecture of the present invention; Figure 3 This is a flowchart of the method according to Embodiment 1 of the present invention; Figure 4 This is a flowchart of the method according to Embodiment 2 of the present invention. Detailed Implementation

[0023] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] This invention can be applied to terminals with voice input and output functions, such as smartphones, tablets, IoT devices, and network communication apps or devices. Example 1

[0025] Please see Figures 1-3 As shown, the technical solution provided by this invention is: a communication system based on voiceprint recognition and speech reconstruction, comprising two communication devices, such as user A device and user B device, wherein the communication devices are each other's communication counterparts. It may also include more than two communication devices, such as three or four, but at least two must be included.

[0026] Please see Figure 1As shown, the communication device includes a central processing module, a voice acquisition module, a storage module, a voice conversion module, a data transmission module, a data reception module, a voice synthesis module, and a voice playback module. The central processing module is communicatively connected to the voice acquisition module, storage module, voice conversion module, data transmission module, data reception module, voice synthesis module, and voice playback module. The voice acquisition module is communicatively connected to the storage module. The voice conversion module is communicatively connected to the voice synthesis module and the data reception module. The voice synthesis module is communicatively connected to the voice playback module. The voice acquisition module is used to acquire voice samples from the peer user and extract the peer user's voiceprint feature vector; the storage module is used to store the voice samples and voiceprint feature vectors acquired by the voice acquisition module; the voice conversion module is used to convert the local user's voice stream into text data; the data sending module is used to send the text data to the peer user via the network; the data receiving module is used to receive text data from the peer user via the network; the voice synthesis module is used to call the peer user's identity information in the storage module, fuse the text data with the peer user's voiceprint feature vector in the storage module, and synthesize a voice stream with the peer user's timbre characteristics; the voice playback module is used to play the voice; the central processing module is used to process data and communicate with the voice acquisition module, storage module, voice conversion module, data sending module, data receiving module, voice synthesis module, and voice playback module. The processing methods include encryption and / or compression, decryption and / or decompression.

[0027] The communication in this invention is divided into two stages: the voiceprint extraction stage and the formal communication stage.

[0028] Voiceprint extraction stage: After the call is established, the voiceprint extraction stage begins and lasts for 10-30 seconds, with extraction performed every 10 seconds. This can be repeated 1-3 times to improve feature stability. User A and User B engage in a normal, seamless call. User A's voice acquisition module extracts the voiceprint feature vector VB from User B and stores it in the storage module. Simultaneously, User B's voice acquisition module extracts the voiceprint feature vector VA from User A and stores it in the storage module.

[0029] Formal call phase: User A's device's speech conversion module (e.g., a lightweight model based on Streaming Transformer or Conformer) converts the speech stream into text data in real time; User B's device's speech conversion module converts the speech stream into text data in real time. The speech conversion module may also include a sentiment analysis module, which generates emotion, stress, and speech rate tags; these are transmitted along with the text to guide speech synthesis at the receiving end. User A's data transmission module sends text data over the network; User B's data transmission module sends text data over the network. Meanwhile, the data receiving module of user A device receives text data from user B device; the data receiving module of user B device receives text data from user A device. User A's speech synthesis module (e.g., a neural network model based on Tacotron2, FastSpeech2, or VITS) uses the VB stored in the local storage module as the speaker conditional code and combines it with emotion tags to synthesize speech waveforms; User B's speech synthesis module uses the VA stored in the local storage module as the speaker conditional code and combines it with emotion tags to synthesize speech waveforms. The voice playback modules of User A's device and User B's device respectively play their respective synthesized voices.

[0030] This invention also provides a communication method based on voiceprint recognition and speech reconstruction. This communication method includes a voiceprint extraction stage and a formal communication stage, comprising the following steps: S1. Voiceprint Extraction: After communication is established, in the voiceprint extraction stage, the voice acquisition module collects the voice samples of the peer user, extracts the voiceprint feature vector of the peer user and stores it in the storage module; S2. Voice to Text: During the formal communication phase, the voice conversion module converts the user's voice stream into text data in real time. S3. Data Transmission: The data transmission module sends text data to the other end via the network; S4. Data Reception: Simultaneously, during the formal communication phase, the data receiving module receives text data from the other end via the network; S5. Text-to-speech and playback: The speech synthesis module calls the identity information of the peer user in the storage module, fuses the text data with the peer user's voiceprint feature vector in the storage module, synthesizes a speech stream with the peer user's timbre features, and plays it through the speech playback module.

[0031] The voiceprint extraction stage adopts a two-way interactive protocol: the two systems alternately or simultaneously prompt the user to read the same preset or random text, record their own local audio and extract voiceprint feature vectors, and complete two-way authentication by comparing the text content sent and received.

[0032] Both the speech acquisition module and the speech synthesis module are locally deployed lightweight neural network models, and the raw speech data does not need to be uploaded to the cloud.

[0033] The voiceprint feature vector serves as a conditional input, used to control the timbre, prosody, and intonation features of the speech synthesis model.

[0034] The text data is accompanied by sub-language information tags (such as emotion, stress, and speech rate tags), which are generated by the sentiment analysis module and transmitted along with the text to guide speech synthesis at the receiving end.

[0035] The present invention also provides a communication system based on voiceprint recognition and speech reconstruction. The communication system includes: a memory storing a computer program; and a processor communicatively connected to the memory. When the computer program is executed by the processor, a communication method based on voiceprint recognition and speech reconstruction as described above is implemented.

[0036] The present invention also provides a communication device readable storage medium having a computer program stored thereon, characterized in that the program, when executed by a processor, implements a communication method based on voiceprint recognition and speech reconstruction as described above.

[0037] In the field of digital communications, the bandwidth required for storage and transmission of voice and text differs significantly. Below is a simple calculation example, assuming voice is sampled at 8kHz and text is sampled using 16-bit digital data: Voice: One minute of voice content, after 16-bit analog-to-digital conversion, requires the following storage space: 16 bits x 8 x 1024 x 60 = 7680 Kbits = 3840 Kbytes.

[0038] Text: Compared to voice, plain text occupies relatively less space because it only needs to store the text itself according to a certain format. A long text message or a long article may only require a few thousand bytes. In summary, in terms of communication modes, both voice and text have their own advantages and disadvantages. However, from a space storage perspective, without compression, voice occupies significantly more space than text.

[0039] Taking 200 characters as an example, this is the speaking speed of most people in one minute. One Chinese character occupies 16 bits, so 200 Chinese characters would require 16 * 200 = 3200, which is 3.2 Kbits = 1.6 Kbytes. At a speaking speed of 200 characters per minute, the bandwidth required for 200 Chinese characters is 1.6 K / 3840 = 0.42% of the bandwidth required for the same amount of speech.

[0040] In conclusion, transmitting text in Chinese characters will save a significant amount of bandwidth. Example 2

[0041] Unlike Example 1, please refer to Figure 4 Before the data is sent, it is encrypted and / or compressed; after the data is received, it is decrypted and / or decompressed, which includes the following steps: S1. Voiceprint Extraction: After communication is established, in the voiceprint extraction stage, the voice acquisition module collects the voice samples of the peer user, extracts the voiceprint feature vector of the peer user and stores it in the storage module; S2. Voice to Text: During the formal communication phase, the voice conversion module converts the user's voice stream into text data in real time. S2-3. Text Data Processing: The central processing module processes the text data to generate secure data; Step S2-3 is between steps S2 and S3; S3′. Data Transmission: The data transmission module sends secure data to the peer via the network; Step S3′ replaces step S3; S4′. Data reception: Simultaneously, during the formal communication phase, the data receiving module receives secure data from the other end via the network; Step S4' replaces step S4; S4-5. Security Data Processing: The central processing module processes the security data to obtain text data; Steps S4-5 are located between steps S4 and S5; S5. Text-to-speech and playback: The speech synthesis module calls the identity information of the peer user in the storage module, fuses the text data with the peer user's voiceprint feature vector in the storage module, synthesizes a speech stream with the peer user's timbre features, and plays it through the speech playback module.

[0042] The central processing module processes secure data, including through encryption and / or compression, decryption and / or decompression.

[0043] Preferably, the encryption process employs a session key derived from the voiceprint feature vectors of both parties during the voiceprint exchange phase, further enhancing security.

[0044] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0045] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0046] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A communication method based on voiceprint recognition and speech reconstruction, characterized in that, The communication method includes a voiceprint extraction stage and a formal communication stage, which includes the following steps: S1. Voiceprint Extraction: After communication is established, in the voiceprint extraction stage, the voice acquisition module collects the voice samples of the peer user, extracts the voiceprint feature vector of the peer user and stores it in the storage module; S2. Voice to Text: During the formal communication phase, the voice conversion module converts the user's voice stream into text data in real time. S3. Data Transmission: The data transmission module sends text data to the other end via the network; S4. Data Reception: Simultaneously, during the formal communication phase, the data receiving module receives text data from the other end via the network; S5. Text-to-speech and playback: The speech synthesis module calls the identity information of the peer user in the storage module, fuses the text data with the peer user's voiceprint feature vector in the storage module, synthesizes a speech stream with the peer user's timbre features, and plays it through the speech playback module.

2. The communication method based on voiceprint recognition and speech reconstruction according to claim 1, characterized in that, It also includes the following steps, S2-3. Text Data Processing: The central processing module processes the text data to generate secure data; Step S2-3 is between steps S2 and S3; S3′. Data Transmission: The data transmission module sends secure data to the peer via the network; Step S3′ replaces step S3; S4′. Data reception: Simultaneously, during the formal communication phase, the data receiving module receives secure data from the other end via the network; Step S4' replaces step S4; S4-5. Security Data Processing: The central processing module processes the security data to obtain text data; Steps S4-5 are located between steps S4 and S5.

3. The communication method based on voiceprint recognition and speech reconstruction according to claim 2, characterized in that, The central processing module processes secure data, including through encryption and / or compression, decryption and / or decompression.

4. A communication device based on voiceprint recognition and speech reconstruction as described in any one of claims 1-3, characterized in that, It includes a central processing module, a voice acquisition module, a storage module, a voice conversion module, a data transmission module, a data reception module, a voice synthesis module, and a voice playback module. The central processing module is communicatively connected to the voice acquisition module, storage module, voice conversion module, data transmission module, data reception module, voice synthesis module, and voice playback module. The voice acquisition module is communicatively connected to the storage module. The voice conversion module is communicatively connected to the voice synthesis module and the data reception module. The voice synthesis module is communicatively connected to the voice playback module.

5. A communication system based on voiceprint recognition and speech reconstruction, characterized in that, It includes at least two communication devices as described in claim 4, wherein the communication devices are each other's communication counterparts.

6. The communication method based on voiceprint recognition and speech reconstruction according to any one of claims 1-3, characterized in that, The voiceprint extraction stage adopts a two-way interactive protocol: the two systems alternately or simultaneously prompt the user to read the same preset or random text, record their own local audio and extract voiceprint feature vectors, and complete two-way authentication by comparing the text content sent and received.

7. The communication device based on voiceprint recognition and speech reconstruction according to claim 4, characterized in that, Both the speech acquisition module and the speech synthesis module are locally deployed lightweight neural network models, and the raw speech data does not need to be uploaded to the cloud.

8. The communication method based on voiceprint recognition and speech reconstruction according to any one of claims 1-3, characterized in that, The voiceprint feature vector serves as a conditional input, used to control the timbre, prosody, and intonation features of the speech synthesis model.

9. The communication device based on voiceprint recognition and speech reconstruction according to claim 4, characterized in that, The speech conversion module also includes an emotion analysis module, which generates emotion, stress, and speech rate tags.

10. A communication system based on voiceprint recognition and speech reconstruction, characterized in that, The communication system includes: Memory, which stores computer programs; A processor, communicatively connected to the memory, implements the method described in any one of claims 1-3 when the computer program is executed by the processor.

Citation Information

Patent Citations

  • Method and system for masking speech

    US7184952B2