Speech Communication Privacy Protection Method, System and Medium Based on Speech Confusion

By generating and parsing voice confusing noise, the problem that the receiver cannot decrypt voice data is solved, and voice data recovery and privacy protection are achieved under lossy channels to ensure that communication quality is not affected.

CN119449493BActive Publication Date: 2025-07-08HANGZHOU HIGH-TECH ZONE (BINJIANG) INSTITUTE OF BLOCKCHAIN & DATA SECURITY +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510032417.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-07-08
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

In the prior art, the recipient cannot effectively decrypt the voice data, resulting in the privacy leakage and identity theft in voice communications not being effectively resolved.

Method used

The speech data is obfuscated by generating speech obfuscating noise, and random numbers and identity information are used to find the target speech in the speech data set, generating speech obfuscating noise, and then deconfused through one-dimensional optimization parameters and improvement factors to restore the original speech data.

Benefits of technology

In the case of loss of channel, ensure that the receiver can successfully decrypt voice data, protect user privacy, and do not significantly affect communication quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119449493B_ABST
    Figure CN119449493B_ABST
Patent Text Reader

Abstract

The present application relates to a method, system and medium for protecting voice communication privacy based on voice confusion, which receives first voice confusion data sent by a first user to a second user; obtains a first random number corresponding to the first user and a first identity of the first user; searches for a target voice in a voice dataset according to the first random number and the first identity, and splices the target voice to generate voice confusion noise; deconfuses the first voice confusion data according to the voice confusion noise to restore the first voice data; even if data transmission is lossy, the voice can be restored based on deconfusion, so that the receiving party can successfully decrypt the voice data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of communication security technologies, and in particular, to a method, a system, and a medium for protecting voice communication privacy based on voice obfuscation. Background Art

[0002] Voice communication technologies, including traditional telephones, Voice over Internet Protocol (VoIP), and voice messages, have developed rapidly in recent years. As these technologies become increasingly important and appear more frequently in people's lives, they have also raised a large number of privacy issues. During the process of voice communication, the transmission of voice data on different links may be accessed and eavesdropped by unauthorized attackers. Since voice data itself contains rich user personal privacy information, the leakage of voice data may lead to users being subject to unauthorized surveillance, user privacy leakage, and even identity theft.

[0003] To reduce the above privacy leakage problems, some end-to-end encryption transmission schemes have been proposed in related technologies. For example, providers of communication services integrate end-to-end encryption methods at the network protocol layer of voice transmission. Another example is that users encrypt voice data with a key at the sending end and then transmit it. However, this encrypted transmission method is prone to data loss during transmission, making it difficult for the receiving end to decrypt.

[0004] Currently, for the problem that the receiving end in related technologies cannot decrypt voice data, no effective solution has been proposed. Summary of the Invention

[0005] Based on this, in view of the above technical problems, it is necessary to provide a method, a system, and a medium for protecting voice communication privacy based on voice obfuscation that can ensure the receiving end successfully decrypts voice data.

[0006] In a first aspect, the present application provides a method for protecting voice communication privacy based on voice obfuscation. The method includes:

[0007] Receiving first voice obfuscated data sent from a first user to a second user;

[0008] Obtaining a first random number corresponding to the first user and a first identity of the first user;

[0009] Searching for a target voice in a voice dataset according to the first random number and the first identity, and splicing the target voice to generate voice obfuscation noise;

[0010] Deobfuscating the first voice obfuscated data according to the voice obfuscation noise to restore the first voice data.

[0011] In one embodiment, before obtaining the first random number corresponding to the first user and the first identity of the first user, the method further includes:

[0012] Receiving a first certificate sent by the first user and returning a second certificate of the second user;

[0013] Receiving first encrypted data sent by the first user, decrypting the first encrypted data by using the private key of the second user to obtain the first random number and the first identity;

[0014] Generating a second random number, encrypting the second random number and the second identity of the second user by using the public key of the first user, and generating and returning second encrypted data.

[0015] In one embodiment, according to the voice obfuscation noise, de-obfuscating the first voice obfuscation data to restore the first voice data, including:

[0016] Segmenting and aligning the first voice obfuscation data and the voice obfuscation noise according to each signal sampling point to obtain multiple groups of corresponding first voice obfuscation data blocks and voice obfuscation noise blocks;

[0017] Setting a parameter set, where the parameter set includes multiple one-dimensional optimization parameters related to signal frequencies;

[0018] Based on the one-dimensional optimization parameters, setting an objective function, traversing the parameter set, obtaining the minimum value of the objective function, and using the one-dimensional optimization parameters corresponding to the minimum value as target parameters; wherein, the objective function is used to measure the degree of noise residue in the original signal after applying the one-dimensional optimization parameters;

[0019] Based on the target parameters, calculating an improvement factor corresponding to each frequency sampling point; wherein, the improvement factor is used to measure the optimized noise suppression effect;

[0020] Weighting the time-frequency spectrum of each of the first voice obfuscation data blocks with the improvement factor to obtain the time-frequency spectrum of the de-obfuscated first voice data.

[0021] In one embodiment, based on the one-dimensional optimization parameters, setting an objective function, traversing the parameter set, and obtaining the minimum value of the objective function, including:

[0022] For each frequency sampling point, performing a scaling process on the voice obfuscation noise block based on the one-dimensional optimization parameters to obtain a first value;

[0023] Take the absolute value of the first numerical value and compare it with the corresponding first speech confusion data block to obtain a ratio, and take 1 minus the ratio to obtain a second numerical value;

[0024] Use the second numerical value as a weight to weight each of the first speech confusion data blocks to obtain the value of the objective function.

[0025] In one embodiment, based on the target parameter, calculating an improvement factor corresponding to each frequency sampling point includes:

[0026] For each frequency sampling point, perform a scaling process on the speech confusion noise block based on the target parameter to obtain a first numerical value;

[0027] Take the absolute value of the first numerical value and compare it with the corresponding first speech confusion data block to obtain a ratio, and take 1 minus the ratio to obtain a second numerical value;

[0028] Compare the second numerical value with zero, and take the larger value of the two as the value of the improvement factor.

[0029] In one embodiment, after calculating the improvement factor corresponding to each frequency sampling point based on the target parameter, the method further includes:

[0030] Obtain the channel loss rate of the first link, and compare the improvement factor with the channel loss rate; the first link is a communication link from the first user to the second user;

[0031] If the improvement factor is less than or equal to the channel loss rate, set the improvement factor to zero.

[0032] In one embodiment, before deconfusing the first speech confusion data according to the speech confusion noise and restoring the first speech data, the method further includes:

[0033] Obtain the frequency response of the first link, and use the frequency response to compensate the speech confusion noise.

[0034] In one embodiment, before deconfusing the first speech confusion data according to the speech confusion noise and restoring the first speech data, the method further includes:

[0035] Receive the second speech data sent by the first user based on the first link, and obtain the prior data of the second speech data;

[0036] According to the second speech data and the prior data, perform channel state estimation on the first link to obtain channel state parameters; wherein, the channel state parameters include frequency response and / or channel loss rate.

[0037] In a second aspect, the present application provides a method for protecting voice communication privacy based on voice confusion. The method includes:

[0038] Receiving a voice input from a first user to obtain first voice data;

[0039] Determining a first identity in a voice dataset that matches the voiceprint information of the first user;

[0040] Determining a first random number for the current session, and generating voice confusion noise based on the first random number, the first identity, and the voice dataset;

[0041] Superimposing the voice confusion noise on the first voice data to obtain first voice-confused data.

[0042] In one embodiment, generating voice confusion noise based on the first random number, the first identity, and the voice dataset includes:

[0043] Generating an index sequence based on the first random number;

[0044] Searching for target voice corresponding to the first identity in the voice dataset according to the index sequence;

[0045] Concatenating the target voice to obtain the voice confusion noise.

[0046] In one embodiment, before determining the first random number for the current session, the method further includes:

[0047] Sending a first certificate of the first user to the second user;

[0048] Receiving a second certificate returned by the second user;

[0049] Generating the first random number, encrypting the first random number and the first identity using the public key of the second user, and sending the generated first encrypted data to the second user;

[0050] Receiving second encrypted data returned by the second user, and decrypting the second encrypted data using the private key of the first user; wherein, the second encrypted data is generated by encrypting a second random number and a second identity of the second user using the public key of the first user.

[0051] In one embodiment, after superimposing the voice confusion noise on the first voice data to obtain first voice-confused data, the method further includes:

[0052] Sending the first voice-confused data to the second user; or,

[0053] Send the first voice obfuscation data to a first communication terminal corresponding to the first user, and send it to the second user via the first communication terminal.

[0054] In a third aspect, the present application provides a voice processing system. The system includes: a processor, a speaker, and a microphone, and the speaker and the microphone are respectively connected to the processor; wherein,

[0055] The processor is configured to execute the voice communication privacy protection method based on voice obfuscation described in the first aspect or the second aspect above.

[0056] In a fourth aspect, the present application provides a voice communication system, including: a voice processing system and a communication terminal; wherein,

[0057] The communication terminal is configured to receive voice obfuscation data and send the voice obfuscation data to the voice processing system, and the voice processing system is configured to execute the voice communication privacy protection method based on voice obfuscation described in the first aspect above; or,

[0058] The voice processing system is configured to execute the voice communication privacy protection method based on voice obfuscation described in the second aspect above and send the generated voice obfuscation data to the communication terminal.

[0059] In a fifth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect or the second aspect above are implemented.

[0060] For the above voice communication privacy protection method, system, and medium based on voice obfuscation, by obfuscating the original voice data, the generated voice obfuscation data does not need to maintain a high degree of data integrity during transmission. Even if the data transmission is lossy, the voice can be restored based on de-obfuscation. Therefore, through the method of this embodiment, it can be ensured that the receiving party can successfully decrypt the voice data and can also work properly in the case of a lossy channel. Description of the Drawings

[0061] Figure 1 It is a structural block diagram of a voice processing system in an embodiment;

[0062] Figure 2 It is a flowchart of a voice communication privacy protection method for a receiving party in an embodiment;

[0063] Figure 3 It is a flowchart of data exchange between users in an embodiment;

[0064] Figure 4Flowchart of the voice deobfuscation method in an embodiment;

[0065] Figure 5 Flowchart of the voice communication privacy protection method based on voice obfuscation for the sender in an embodiment;

[0066] Figure 6 Schematic diagram of the application environment of the voice communication system in an embodiment;

[0067] Figure 7 For Figure 6 The operating principle diagram of the voice communication system in

[0068] Figure 8 Experimental result diagram for the user voice privacy protection ability in an embodiment;

[0069] Figure 9 Experimental result diagram for the voice communication quality in an embodiment. Detailed implementation manners

[0070] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0071] Unless otherwise defined, the technical terms or scientific terms involved in the present application shall have the general meanings understood by those with ordinary skills in the technical field to which the present application belongs. In the present application, words such as "a", "one", "a kind of", "the", "these" and the like do not represent a limitation in quantity, and they can be singular or plural. The terms "including", "comprising", "having" and any variations thereof involved in the present application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device including a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent in these processes, methods, products or devices. The terms "connected", "coupled" and the like involved in the present application do not limit to physical or mechanical connections, but may include electrical connections, whether directly or indirectly. The "multiple" involved in the present application means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. Usually, the character " / " indicates that the objects before and after are in an "or" relationship. The terms "first", "second", "third", etc. involved in the present application are only used to distinguish similar objects and do not represent a specific order for the objects.

[0072] In one embodiment, Figure 1 a structural block diagram of a voice processing system is provided, as Figure 1 shown. The voice processing system includes: a processor 101, a speaker 102, and a microphone 103. The speaker 102 and the microphone 103 are respectively connected to the processor 101. The processor 101 can be one or more ( Figure 1 only one is shown in the figure), and the processor 101 can include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA.

[0073] Optionally, the above voice processing system further includes a memory 104 for storing data and a transmission device 105 for communication functions. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above voice processing system. For example, the voice processing system may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 that shown in the figure.

[0074] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the voice communication privacy protection method based on voice confusion in this embodiment. The processor 101 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 can include high-speed random access memory and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 can further include a memory remotely set relative to the processor 101, and these remote memories can be connected to the terminal through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0075] The transmission device 105 is used to receive or send data via a network. The above network includes a wireless network provided by a communication provider. In one instance, the transmission device 105 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 105 can be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.

[0076] When the user corresponding to the voice processing system is the sender, the processor 101 executes Figure 2The steps shown. When the user corresponding to the voice processing system is the recipient, the processor 101 executes Figure 5 The steps shown. Of course, the voice processing system can be both a sender and a recipient. Therefore, the processor 101 simultaneously has the function of executing Figure 2 and Figure 5 The steps shown. Some of the following embodiments will introduce the voice communication privacy protection method based on voice obfuscation of the present application from the perspectives of the recipient and the sender respectively.

[0077] In one embodiment, Figure 2 A flowchart of a voice communication privacy protection method for the recipient based on voice obfuscation is provided. Taking the application of this method to Figure 1 The voice processing system as an example, the process includes the following steps:

[0078] Step S21, receive the first voice obfuscation data sent by the first user to the second user.

[0079] The first voice obfuscation data is generated by superimposing the voice obfuscation noise of the first user on the original voice data. In this embodiment, the first voice obfuscation data can be directly received by the voice processing system on the second user side, or can be first received by the communication terminal on the second user side, and then the communication terminal sends the first voice obfuscation data to the voice processing system. This embodiment is not limited.

[0080] Optionally, before step S21, the method further includes a system registration step: when each voice processing system instance leaves the factory, it registers its identity in the same trusted third-party certification authority to obtain the voice processing system certificate Cert dn And the certificate Cert ca Of this trusted third-party certification authority, and at the same time obtain a public voice data set D, such as the ai_datatang data set. After the user obtains the corresponding voice processing system, entering the user registration stage, a pair of key pairs (p k , s k ) can be generated using the RSA-1024 algorithm, and a user voice signal is recorded, with a length of about 20 to 60 seconds, to accurately extract the user's voiceprint information. The user's voiceprint information is extracted from this voice signal, and based on this voiceprint information, the speaker identity s with the closest voiceprint information in the voice data set D is matched using cosine similarity, which is the user identity.

[0081] Step S22, obtain the first random number corresponding to the first user and the first identity of the first user.

[0082] The first random number and the first identity can be obtained through data exchange between users. For example, the first user directly sends the first random number and the first identity to the second user; or, after the first user verifies the identity of the second user, the first random number and the first identity are sent to the second user.

[0083] Step S23: According to the first random number and the first identity, search for the target voice in the voice dataset, and splice the target voice to generate voice confusion noise.

[0084] First, an index sequence can be generated according to the first random number; the target voice corresponding to the first identity is searched in the voice dataset according to the index sequence; the target voice is spliced to obtain the voice confusion noise.

[0085] Step S24: According to the voice confusion noise, deconfuse the first voice confusion data to restore the first voice data.

[0086] In this step, the target parameter that can minimize the ratio of the voice confusion noise to the original signal can be found from multiple one-dimensional optimization parameters related to frequency, and then the improvement factor is calculated using the target parameter, and the first voice confusion data is processed based on the improvement factor to restore the clear first voice data.

[0087] From an end-to-end perspective, the channel is lossy and will damage the encrypted data. Traditional voice communication privacy protection methods have relatively high requirements for the integrity of encrypted data. If some bits are missing during the transmission of encrypted data, it will be difficult for the receiving party to decrypt. In the above steps S21 to S24, by confusing the original voice data, the generated voice confusion data does not need to maintain high data integrity during transmission. Even if the data transmission is lossy, the voice can be restored based on deconfusion. Therefore, through the method of this embodiment, it can be ensured that the receiving party can successfully decrypt the voice data and can work properly even in the case of a lossy channel.

[0088] In one embodiment, Figure 3 A flowchart of data exchange between users is provided, as Figure 3 shown. Before step S22 of obtaining the first random number corresponding to the first user and the first identity of the first user, the method further includes the following steps:

[0089] Step 1: Receive the first certificate sent by the first user and return the second certificate of the second user.

[0090] The two communication parties use a voice communication channel for data exchange. The data modulation method during the data exchange process is frequency modulation, and the carrier frequency is 800 + i×400 Hz; where i = 0, 1, …, 12, and the length of each data modulation signal is 0.075 seconds. Among the 13-bit data, the first bit is the reference bit and is always 0; the 10th to 13th bits are the check bits, and the check method is CRC-4.

[0091] During the data exchange process, the two communication parties first generate a random number and send it to the other party. The two parties compare the sizes of the random numbers to determine the master-slave relationship. Assume that the two communication parties are a (the second user) and b (the first user). The random number generated by the former is 0.5, and the random number generated by the latter is 0.6. Then, during this communication process, b is the host and a is the slave. After determining the master-slave relationship, the host first sends its certificate Cert to the slave. b After receiving the host's certificate, the slave will first send an acknowledgment signal ACK to the host and then send its certificate Cert a to the host. The acknowledgment signal ACK consists of two swept-frequency signals with a frequency range of 1 Hz to 8000 Hz and a length of 0.24 seconds.

[0092] Step 2: Receive the first encrypted data sent by the first user, and decrypt the first encrypted data using the private key of the second user to obtain the first random number and the first identity.

[0093] After receiving the slave's certificate, the host will generate a first random number RN b . Then it sends an acknowledgment signal ACK to the slave and sends the first random number RN b encrypted with the slave's public key and the first identity s b .

[0094] Step 3: Generate a second random number, encrypt the second random number and the second identity of the second user using the public key of the first user, and generate and return the second encrypted data.

[0095] After receiving the encrypted first random number RN b and the first identity s b sent by the host, the slave will first decrypt the data using its own private key, and at the same time send an acknowledgment signal ACK and the second random number RN a encrypted with the host's public key and the second identity s a to the host. After receiving the encrypted second random number RN a and the second identity s a sent by the slave, the host will decrypt the data using its own private key.

[0096] In this embodiment, the two communication parties use the voice communication channel to exchange data, and use the negotiated key to encrypt and decrypt the exchanged data, ensuring the security of data exchange.

[0097] In one embodiment, Figure 4 A flowchart of a voice deobfuscation method is provided, as Figure 4 shown. The above step S24, deobfuscating the first voice obfuscated data according to the voice obfuscation noise to restore the first voice data, includes the following steps:

[0098] Step S241, splitting and aligning the first voice obfuscated data and the voice obfuscation noise according to each signal sampling point to obtain multiple groups of corresponding first voice obfuscated data blocks and voice obfuscation noise blocks.

[0099] Assume that the received obfuscated voice sequence is y(t), and the corresponding noise sequence n(t) is generated using the first random number RN b , the first identity s b and the voice dataset D. Then, the first voice obfuscated data y(t) and the voice obfuscation noise n(t) are split into data blocks of length 2048, and each first voice obfuscated data block and voice obfuscation noise block are aligned in the time domain based on the cosine similarity to obtain multiple groups of corresponding first voice obfuscated data blocks and voice obfuscation noise blocks.

[0100] Step S242, setting a parameter set, which contains multiple one-dimensional optimization parameters related to the signal frequency.

[0101] Set multiple one-dimensional optimization parameters related to the signal frequency , and a parameter set is formed based on these multiple one-dimensional optimization parameters.

[0102] Step S243, setting an objective function based on the one-dimensional optimization parameters, traversing the parameter set, obtaining the minimum value of the objective function, and using the one-dimensional optimization parameter corresponding to this minimum value as the target parameter; wherein, the objective function is used to measure the degree of noise residue in the original signal after applying the one-dimensional optimization parameter.

[0103] The calculation formula of the objective function is as follows:

[0104]

[0105] wherein, , i represents the frequency sampling point number, represents the signal frequency, t represents time, represents multiplication. Step S243 specifically includes the following steps:

[0106] Step 1: For each frequency sampling point, scale the speech-convolved noise block based on the one-dimensional optimization parameter to obtain a first value. ;

[0107] Step 2: Compare the first value with the corresponding first speech-convolved data block after taking the absolute value, to obtain a ratio , and subtract the ratio from 1 to obtain a second value ;

[0108] Step 3: Weight each first speech-convolved data block using the second value as the weight to obtain the value of the objective function.

[0109] It should be noted that before calculating the objective function, the speech-convolved noise block and the first speech-convolved data block need to be converted from time-domain data to frequency-domain data. Therefore, the calculations of the objective function and the improvement factor in this application are all for frequency-domain data.

[0110] Step S244: Calculate the improvement factor corresponding to each frequency sampling point based on the target parameter; among them, the improvement factor is used to measure the noise suppression effect after optimization.

[0111] The formula for the improvement factor is as follows:

[0112]

[0113] Among them, represents the improvement factor, i represents the frequency sampling point number, represents the signal frequency, t represents time, represents multiplication. Step S244 specifically includes the following steps:

[0114] Step 1: For each frequency sampling point, scale the speech-convolved noise block based on the target parameter to obtain a first value ;

[0115] Step 2: Compare the first value with the corresponding first speech-convolved data block after taking the absolute value, to obtain a ratio , and subtract the ratio from 1 to obtain a second value ;

[0116] Step 3: Compare the second value with zero, and take the larger value as the value of the improvement factor.

[0117] Step S245: Weight the time-frequency spectrum of each first speech-convolved data block using the improvement factor as the weight to obtain the time-frequency spectrum of the first speech data after deconvolution.

[0118] In this embodiment, by traversing the parameter set to calculate the minimum value of the objective function, the target parameter α can be obtained. This target parameter α is the optimal noise figure, and an improvement factor is calculated based on the optimal noise figure and applied to the time-frequency spectrum of the confused speech to obtain as pure an original speech as possible.

[0119] In one embodiment, in step S244 above, after calculating the improvement factor corresponding to each frequency sampling point based on the target parameter, the method further includes:

[0120] Obtain the channel loss rate of the first link. Compare the improvement factor with the channel loss rate. The first link is a communication link from the first user to the second user. If the improvement factor is less than or equal to the channel loss rate, set the improvement factor to zero.

[0121] For any frequency , and at the same time for each data in , if then set to 0. The finally obtained time-frequency spectrum of the de-confused speech s(t) is , where is the time-frequency spectrum of y(t). In this embodiment, for a certain frequency , if , it is considered that the signal at this frequency has been severely damaged by the transmission link compression algorithm to the extent that it cannot be restored. Therefore, it is set to zero to completely exclude the data at this frequency. With such a setting, more original signal components are retained in some places where the noise is effectively suppressed, while the signal contribution is reduced or even eliminated at frequencies where the noise is too severe, ultimately improving the signal quality.

[0122] In one embodiment, in step S24 above, before de-confusing the first voice-confused data according to the voice confusion noise to restore the first voice data, the method further includes:

[0123] Obtain the frequency response of the first link, and use the frequency response to compensate for the speech confusion noise. In this embodiment, the frequency-domain compensation for the speech confusion noise n(t) can be performed based on the frequency response curve FR(f). The compensation method is N(f,t) = N0(f,t) × FR(f), where N0(f,t) is the time-frequency spectrum of n(t) (i.e., the short-time Fourier transform of the original speech confusion noise). Considering that there are losses in the actual channel transmission and the speech will be damaged after transmission, assume the user speech is j, the noise is k, and the confused speech is j + k, then it may attenuate to j' + k' during the channel transmission. By compensating k', the frequency-domain characteristics of the noise sequence can be adjusted to match the actual channel conditions, which is beneficial to improving the quality of the first speech data recovered subsequently.

[0124] In one embodiment, before the step S24 of deconfusing the first speech confusion data according to the speech confusion noise and restoring the first speech data, the method further includes the following steps:

[0125] Receive the second speech data sent by the first user based on the first link, and obtain the prior data of the second speech data; according to the second speech data and the prior data, perform channel state estimation on the first link to obtain channel state parameters; wherein, the channel state parameters include Frequency Response and / or Loss Ratio.

[0126] In this embodiment, the second user estimates the channel state by receiving the speech data of the first user and obtains the relevant parameters of the channel state.

[0127] For example, the first user sends a swept-frequency signal s(t) with a length of 5 seconds and a frequency range of 5Hz to 8000Hz to the second user. After the channel lossy transmission, assume the signal received by the second user is s'(t), calculate the Fourier transforms S(f) and S'(f) of these two signals, and estimate the channel frequency response:

[0128]

[0129] Among them, S(f) is the Fourier transform of s(t), and S'(f) is the Fourier transform of s'(t).

[0130] For example, the first user sends a normal speech signal s(t) with a length of 5 seconds to the second user. After the channel lossy transmission, assume the signal received by the second user is s'(t). First, the second user performs energy normalization processing on s'(t): , ensure that the energy of the transmitted voice signal and the received voice signal is equal. Next, calculate the time-frequency spectrum S(t, f) of s(t) with a window length of 2048 and a step size of 1024, and calculate the time-frequency spectrum S'(t, f) of s'(t). Based on these two signal time-frequency spectra, calculate the loss percentage of the transmitted / received signal at different frequencies. Finally, the calculated channel loss rate is:

[0131]

[0132] In this embodiment, by performing channel state estimation on the first link, the frequency response and / or channel loss rate are obtained, so that when the second user deconfounds the first voice confusion data, frequency-domain compensation is performed on the voice confusion noise based on the frequency response, and / or the improvement factor is processed based on the channel loss rate, ultimately improving the quality of the restored first voice data.

[0133] Optionally, the channel state parameter estimation is a two-way estimation, that is, the first user can also estimate the channel state parameters of the second link by receiving the signal of the second user, where the second link is a communication link from the second user to the first user. Such a setting is for facilitating deconfounding of the voice confusion data of the second user received subsequently.

[0134] In one embodiment, Figure 5 A flowchart of a method for voice communication privacy protection based on voice confusion of a sender is provided. Taking this method applied to Figure 1 a voice processing system as an example, the process includes the following steps:

[0135] Step S11, receive the voice input of the first user to obtain the first voice data.

[0136] The voice communication system collects the voice of the first user to obtain the first voice data.

[0137] Optionally, before step S11, the method further includes a system registration step: when each voice processing system instance leaves the factory, it registers its identity with the same trusted third-party certification authority to obtain the voice processing system certificate Cert dn and the certificate Cert of this trusted third-party certification authority ca , and at the same time obtain a common voice dataset D, such as the ai_datatang dataset.

[0138] Step S12, determine the first identity that matches the voiceprint information of the first user in the voice dataset.

[0139] After the user obtains the corresponding voice processing system, it enters the user registration stage. The RSA-1024 algorithm can be used to generate a pair of key pairs (p k , sk ) and record a user voice signal with a length of about 20 to 60 seconds to accurately extract the user's voiceprint information. Extract the user's voiceprint information from the voice signal, and based on the voiceprint information, use cosine similarity to match the identity s of the speaker with the closest voiceprint information in the voice dataset D, which is the user identity.

[0140] Step S13, determine the first random number of the current session, and generate voice confusion noise according to the first random number, the first identity, and the voice dataset.

[0141] First, an index sequence can be generated according to the first random number; search for the target voice corresponding to the first identity in the voice dataset according to the index sequence; splice the target voice to obtain voice confusion noise. Optionally, the target voice is preferably vowel data because vowel data is less likely to be cracked compared to consonant data. With such a setting, the security of the first voice confusion data can be improved.

[0142] Step S14, superimpose the voice confusion noise on the first voice data to obtain the first voice confusion data.

[0143] Superimpose the voice confusion noise on the first voice data to confuse the voice of the first user. Optionally, the signal-to-noise ratio of the voice confusion noise during confusion is -9.

[0144] From an end-to-end perspective, the channel is lossy and will damage the encrypted data. Traditional voice communication privacy protection methods have relatively high requirements for the integrity of encrypted data. If some bits are missing during the transmission of encrypted data, it will be difficult for the receiving party to decrypt. In the above steps S11 to S14, by confusing the original voice data, the generated voice confusion data does not need to maintain a high degree of data integrity during transmission. Even if the data transmission is lossy, the voice can be restored based on deconfusion. Therefore, through the method of this embodiment, it can be ensured that the receiving party successfully decrypts the voice data and can work properly even in the case of a lossy channel.

[0145] In one embodiment, Figure 3 provides a flowchart of data exchange between users, as Figure 3 shown. Before step S13 of determining the first random number of the current session, the method further includes the following steps:

[0146] Step 1, send the first certificate of the first user to the second user.

[0147] The two communication parties use a voice communication channel for data exchange. The data modulation method during the data exchange process is frequency modulation, and the carrier frequency is 800 + i×400 Hz; where i = 0, 1, …, 12, and the length of each data modulation signal is 0.075 seconds. Among the 13-bit data, the first bit is the reference bit and is always 0; the 10th to 13th bits are the parity bits, and the parity check method is CRC-4.

[0148] During the data exchange process, the two communication parties first generate a random number and send it to the other party, and the two parties compare the sizes of the random numbers to determine the master-slave relationship. Assume that the two communication parties are a (the second user) and b (the first user), the random number generated by the former is 0.5, and the random number generated by the latter is 0.6. Then in this communication process, b is the host and a is the slave. After determining the master-slave relationship, the host first sends its certificate Cert to the slave. b 。

[0149] Step 2, receive the second certificate returned by the second user.

[0150] After receiving the host certificate, the slave will first send an acknowledgment signal ACK to the host and then send its certificate Cert to the host. The acknowledgment signal ACK consists of two swept-frequency signals with a frequency range of 1 Hz to 8000 Hz and a length of 0.24 seconds. a to the host.

[0151] Step 3, generate a first random number, encrypt the first random number and the first identity using the public key of the second user, and send the generated first encrypted data to the second user.

[0152] After receiving the slave's certificate, the host will generate a first random number RN. b After that, it sends an acknowledgment signal ACK to the slave and sends the first random number RN encrypted using the slave's public key. b and the first identity s. b 。

[0153] Step 4, receive the second encrypted data returned by the second user, and decrypt the second encrypted data according to the private key of the first user; where the second encrypted data is generated by encrypting the second random number and the second identity of the first user using the public key of the first user.

[0154] After receiving the encrypted first random number RN sent by the host. b and the first identity s. b , the slave will first decrypt the data using its own private key, and at the same time send an acknowledgment signal ACK and the second random number RN encrypted using the host's public key. a and the second identity s. a to the host. After receiving the encrypted second random number RN sent by the slave. a and the second identity s.a , will use its private key to decrypt the data.

[0155] In this embodiment, the two communication parties use a voice communication channel for data exchange, and use a negotiated key to encrypt and decrypt the exchanged data, ensuring the security of data exchange.

[0156] In one embodiment, after step S14 of superimposing voice confusion noise on the first voice data to obtain the first voice confusion data, the method further includes:

[0157] The voice processing system directly sends the first voice confusion data to the second user, that is, the voice processing system can communicate with external devices using the services of a communication service provider.

[0158] Alternatively, the voice communication system first sends the first voice confusion data to the first communication terminal corresponding to the first user, and then sends it to the second user via the first communication terminal. Such a setting is considered because if the first communication terminal is attacked by a third party, it will become an untrustworthy device. Therefore, by using the voice processing system to collect and process the user's voice and then sending the processed user's voice to the communication terminal, potential third-party attacks on the communication terminal can be isolated.

[0159] In one embodiment, before step S14 of sending the first voice confusion data to the second user, the method further includes the following steps:

[0160] Receive the voice input of the first user to obtain the second voice data; send the second voice data to the second user based on the first link, and provide the prior data of the second voice data to the second user.

[0161] In this embodiment, the first user sends and provides voice data to the second user to facilitate the second user to estimate the channel state and obtain relevant parameters of the channel state.

[0162] For example, the first user sends a sweep signal s(t) of 5Hz to 8000Hz to the second user. Assuming the signal received by the second user is s'(t), the calculated frequency response is:

[0163]

[0164] where S(f) is the Fourier transform of s(t) and S'(f) is the Fourier transform of s'(t).

[0165] For example, the first user sends a normal voice signal s(t) to the second user. Assuming the signal received by the second user is s'(t), the second user first performs energy normalization processing: s(t)=s(t), ; Calculate the time-frequency spectrum S(t, f) of s(t), and calculate the time-frequency spectrum S'(t, f) of s'(t); Finally, the calculated channel loss rate is:

[0166]

[0167] In this embodiment, by performing channel state estimation on the first link, the frequency response and / or the channel loss rate are obtained, so that when the second user de-obfuscates the first voice obfuscated data, frequency-domain compensation is performed on the voice obfuscation noise based on the frequency response, and / or the improvement factor is processed based on the channel loss rate, and finally the quality of the restored first voice data is improved.

[0168] Optionally, the channel state parameter estimation is a two-way estimation, that is, the first user can also estimate the channel state parameters of the second link by receiving the signal of the second user, where the second link is a communication link pointing from the second user to the first user. Such a setting is for facilitating de-obfuscation of the voice obfuscated data received from the second user in the subsequent process.

[0169] In one embodiment, a voice communication system is provided. The voice communication system includes: a voice processing system and a communication terminal; the voice processing system is connected to the communication terminal. When the voice communication system is used as a sender, the voice processing system is configured to execute the voice communication privacy protection method based on voice obfuscation for the corresponding sender, and send the generated voice obfuscated data to the peer through the communication terminal. When the voice communication system is used as a receiver, the voice obfuscated data is received through the communication terminal and sent to the voice processing system, and the voice processing system is configured to execute the voice communication privacy protection method based on voice obfuscation for the receiver to de-obfuscate the voice obfuscated data.

[0170] The above voice communication system can be applied to an application environment as shown in Figure 6 As shown in Figure 6 As shown, on the first user side, there is a first voice processing system and a first communication terminal, and on the second user side, there is a second voice processing system and a second communication terminal. The voice processing systems on both sides are connected to their respective communication terminals, and the communication terminals on both sides are communicatively connected to each other. For example, the voice processing system can be a portable wearable device such as a smart watch, a smart bracelet, a head-mounted device, etc. The wearable device has a microphone and a speaker, and the processor is built into the wearable device. The communication terminal can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, etc.

[0171] Figure 7 For Figure 6 the operating principle diagram of the voice communication system in Figure 7 As shown, the operation process of the voice communication system includes the following steps:

[0172] Step S1, registration of the voice processing system. For each instance of the voice processing system, identity registration is performed at a trusted third-party certification authority during factory production to obtain a personal certificate and a third-party certification authority certificate. The voice processing system contains a common voice dataset.

[0173] Step S2, user registration and key generation. After obtaining the voice processing system, the user generates a key pair. Meanwhile, the user records a personal voice signal using the system and identifies the identity of the voice signal in the voice dataset based on voiceprint matching.

[0174] Step S3, channel parameter estimation. After the user starts a voice communication, the systems of both communication parties estimate the channel state to obtain relevant parameters of the channel state.

[0175] Step S4, data key exchange. After the channel state parameter estimation is completed, the two communication parties perform an automatic key exchange through the voice transmission channel. After the key exchange is completed, the two communication parties use the key to encrypt and transmit the current session random number.

[0176] Step S5, sender noise generation and voice confusion. During the user's voice communication, the voice input of the user is received, and voice confusion noise is generated using a random number and the user's phoneme dataset. The confusion noise is added to the user's voice to confuse the voice data, and the confused voice data is sent to the communication terminal.

[0177] Step S6, receiver noise generation and voice de - confusion. During the user's voice communication, the confused voice received from the communication terminal is received, and confusion noise is generated using a random number. Based on the confusion noise, the confused voice is de - confused, and the de - confused voice is played to the user.

[0178] Compared with the existing encryption methods of communication service providers, the voice processing system or voice communication system of this embodiment does not rely on a trusted service provider and is more universal. Compared with other existing encryption - based methods, the voice processing system or voice communication system of this embodiment can work normally under lossy channels, which has more practical significance and can be used in actual situations. The voice processing system or voice communication system of this embodiment can adapt to different types of voice communication channels without additional operations by the user.

[0179] Figure 8This is an experimental result graph for the user voice privacy protection ability, used to verify the ability of the voice processing system to protect user voice privacy during the actual voice message communication process. This experiment shows the effects of user voice privacy protection under different voice communication channels and different attacker conditions. The abscissa represents different voice message communication channels (T: Telegram, V: Viber, M: Messenger, S: Skype, L: Line, D: DingTalk, WA: WhatsApp, WE: WeChat). The experimental results show that without any protection, if an attacker can obtain the voice signal during the transmission process, the attacker can obtain the user's private semantic information from it and achieve a word error rate of less than 5%. Under the protection of the voice processing system, for any type of attacker, even if they can obtain the confused voice signal during the transmission process, the word error rate of the semantic information they obtain is greater than 75%.

[0180] Figure 9 This is an experimental result graph for voice communication quality, used to verify that the voice processing system will not have a significant impact on the user's voice communication quality. This experiment shows the impact on voice communication quality under different voice communication channels. The abscissa represents different voice message communication channels (T: Telegram, V: Viber, M: Messenger, S: Skype, L: Line, D: DingTalk, WA: WhatsApp, WE: WeChat). Without protection, the average word error rate of the semantic information obtained during the user's voice communication process is about 5%. Under the protection of the voice processing system, for the first four communication channels, the accuracy rate of the semantic information that the user can obtain is hardly affected. For the last four channels, the word error rate is slightly higher, but still lower than 40%.

[0181] It should be understood that although each step in the flowcharts involved in the above-described embodiments is shown in sequence according to the indication of the arrows, these steps do not necessarily need to be executed in the order indicated by the arrows. Unless there is a clear description in this article, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily need to be executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential either, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0182] In addition, in combination with the voice communication privacy protection method based on voice confusion provided in the above embodiments, the present application also provides a storage medium, on which a computer program is stored; when the computer program is executed by a processor, the following steps are implemented:

[0183] Receive the first voice confusion data sent by the first user to the second user;

[0184] Obtain the first random number corresponding to the first user and the first identity of the first user;

[0185] According to the first random number and the first identity, search for the target voice in the voice dataset, splice the target voice, and generate voice confusion noise;

[0186] According to the voice confusion noise, deconfuse the first voice confusion data to restore the first voice data.

[0187] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: Before obtaining the first random number corresponding to the first user and the first identity of the first user, the method further includes;

[0188] Receive the first certificate sent by the first user, and return the second certificate of the second user;

[0189] Receive the first encrypted data sent by the first user, decrypt the first encrypted data using the private key of the second user to obtain the first random number and the first identity;

[0190] Generate a second random number, encrypt the second random number and the second identity of the second user using the public key of the first user, and generate and return the second encrypted data.

[0191] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: According to the voice confusion noise, deconfuse the first voice confusion data to restore the first voice data, including:

[0192] Slice and align the first voice confusion data and the voice confusion noise according to each signal sampling point to obtain multiple groups of corresponding first voice confusion data blocks and voice confusion noise blocks;

[0193] Set a parameter set, which contains multiple one-dimensional optimization parameters related to the signal frequency;

[0194] Based on the one-dimensional optimization parameters, set an objective function, traverse the parameter set, obtain the minimum value of the objective function, and use the one-dimensional optimization parameters corresponding to the minimum value as the target parameters; wherein, the objective function is used to measure the degree of noise residue in the original signal after applying the one-dimensional optimization parameters;

[0195] Calculate an improvement factor corresponding to each frequency sampling point based on the target parameter; wherein, the improvement factor is used to measure the optimized noise suppression effect;

[0196] Weight the time-frequency spectrum of each first speech confusion data block with the improvement factor to obtain the time-frequency spectrum of the first speech data after deconfusion.

[0197] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: Set up an objective function based on one-dimensional optimization parameters, traverse the parameter set, and obtain the minimum value of the objective function, including:

[0198] For each frequency sampling point, scale the speech confusion noise block based on the one-dimensional optimization parameter to obtain a first value;

[0199] Compare the first value with the absolute value of the corresponding first speech confusion data block to obtain a ratio, and take 1 minus the ratio to obtain a second value;

[0200] Weight each first speech confusion data block with the second value to obtain the value of the objective function.

[0201] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: Calculate an improvement factor corresponding to each frequency sampling point based on the target parameter, including:

[0202] For each frequency sampling point, scale the speech confusion noise block based on the target parameter to obtain a first value;

[0203] Compare the first value with the absolute value of the corresponding first speech confusion data block to obtain a ratio, and take 1 minus the ratio to obtain a second value;

[0204] Compare the second value with zero, and take the larger value of the two as the value of the improvement factor.

[0205] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: After calculating an improvement factor corresponding to each frequency sampling point based on the target parameter, the method further includes:

[0206] Obtain the channel loss rate of the first link, and compare the improvement factor with the channel loss rate; the first link is a communication link from the first user to the second user;

[0207] If the improvement factor is less than or equal to the channel loss rate, set the improvement factor to zero.

[0208] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: Before deconfusing the first speech confusion data according to the speech confusion noise and restoring the first speech data, the method further includes:

[0209] Obtain the frequency response of the first link, and use the frequency response to compensate for the voice confusion noise.

[0210] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: Before deconfusing the first voice confusion data according to the voice confusion noise and restoring the first voice data, the method further includes:

[0211] Receive the second voice data sent by the first user based on the first link, and obtain the a priori data of the second voice data;

[0212] Perform channel state estimation on the first link according to the second voice data and the a priori data to obtain channel state parameters; wherein, the channel state parameters include frequency response and / or channel loss rate.

[0213] In one embodiment, when the computer program is executed by a processor, the following steps are implemented:

[0214] Receive the voice input of the first user to obtain the first voice data;

[0215] Determine the first identity that matches the voiceprint information of the first user in the voice dataset;

[0216] Determine the first random number of the current session, and generate voice confusion noise according to the first random number, the first identity, and the voice dataset;

[0217] Superimpose the voice confusion noise on the first voice data to obtain the first voice confusion data.

[0218] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: Generating voice confusion noise according to the first random number, the first identity, and the voice dataset includes:

[0219] Generate an index sequence according to the first random number;

[0220] Search for the target voice corresponding to the first identity in the voice dataset according to the index sequence;

[0221] Concatenate the target voices to obtain the voice confusion noise.

[0222] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: Before determining the first random number of the current session, the method further includes:

[0223] Send the first certificate of the first user to the second user;

[0224] Receive the second certificate returned by the second user;

[0225] Generate a first random number, encrypt the first random number and the first identity using the public key of the second user, and send the generated first encrypted data to the second user;

[0226] Receive the second encrypted data returned by the second user, and decrypt the second encrypted data using the private key of the first user; wherein, the second encrypted data is generated by encrypting a second random number and a second identity of the second user using the public key of the first user.

[0227] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: After superimposing voice obfuscation noise on the first voice data to obtain first voice obfuscated data, the method further includes:

[0228] Send the first voice obfuscated data to the second user; or,

[0229] Send the first voice obfuscated data to the first communication terminal corresponding to the first user, and send it to the second user via the first communication terminal.

[0230] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: Before sending the first voice obfuscated data to the second user, the method further includes:

[0231] Receive the voice input of the first user to obtain second voice data;

[0232] Send the second voice data to the second user based on the first link, and provide the prior data of the second voice data to the second user.

[0233] In addition, in combination with the voice communication privacy protection method based on voice obfuscation provided in the above embodiments, the present application also provides a computer program product, including a computer program, which when executed by a processor implements the following steps:

[0234] Receive the first voice obfuscated data sent by the first user to the second user;

[0235] Obtain the first random number corresponding to the first user and the first identity of the first user;

[0236] According to the first random number and the first identity, search for the target voice in the voice dataset, splice the target voice, and generate voice obfuscation noise;

[0237] According to the voice obfuscation noise, de-obfuscate the first voice obfuscated data to restore the first voice data.

[0238] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: Before obtaining the first random number corresponding to the first user and the first identity of the first user, the method further includes;

[0239] Receive the first certificate sent by the first user and return the second certificate of the second user;

[0240] Receive the first encrypted data sent by the first user, decrypt the first encrypted data using the private key of the second user to obtain the first random number and the first identity;

[0241] Generate a second random number, encrypt the second random number and the second identity of the second user using the public key of the first user, generate and return the second encrypted data.

[0242] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: De-confuse the first voice-confused data according to the voice-confusing noise to restore the first voice data, including:

[0243] Slice and align the first voice-confused data and the voice-confusing noise according to each signal sampling point to obtain multiple corresponding first voice-confused data blocks and voice-confusing noise blocks;

[0244] Set a parameter set, where the parameter set contains multiple one-dimensional optimization parameters related to the signal frequency;

[0245] Set an objective function based on the one-dimensional optimization parameters, traverse the parameter set, obtain the minimum value of the objective function, and use the one-dimensional optimization parameter corresponding to the minimum value as the target parameter; where the objective function is used to measure the degree of noise residue in the original signal after applying the one-dimensional optimization parameter;

[0246] Based on the target parameter, calculate the improvement factor corresponding to each frequency sampling point; where the improvement factor is used to measure the optimized noise suppression effect;

[0247] Weight the time-frequency spectrum of each first voice-confused data block with the improvement factor as the weight to obtain the time-frequency spectrum of the de-confused first voice data.

[0248] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: Set an objective function based on the one-dimensional optimization parameters, traverse the parameter set, and obtain the minimum value of the objective function, including:

[0249] For each frequency sampling point, scale the voice-confusing noise block based on the one-dimensional optimization parameter to obtain a first value;

[0250] Compare the first value with the absolute value of the corresponding first voice-confused data block to obtain a ratio, and take 1 minus the ratio to obtain a second value;

[0251] Weight each first voice-confused data block with the second value as the weight to obtain the value of the objective function.

[0252] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: Based on the target parameter, calculate an improvement factor corresponding to each frequency sampling point, including:

[0253] For each frequency sampling point, scale the speech confusion noise block based on the target parameter to obtain a first value;

[0254] Take the absolute value of the first value and compare it with the corresponding first speech confusion data block to obtain a ratio, and take 1 minus the ratio to obtain a second value;

[0255] Compare the second value with zero, and take the larger value of the two as the value of the improvement factor.

[0256] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: After calculating the improvement factor corresponding to each frequency sampling point based on the target parameter, the method further includes:

[0257] Obtain the channel loss rate of the first link, and compare the improvement factor with the channel loss rate; The first link is a communication link from the first user to the second user;

[0258] If the improvement factor is less than or equal to the channel loss rate, set the improvement factor to zero.

[0259] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: Before deconfusing the first speech confusion data according to the speech confusion noise to restore the first speech data, the method further includes:

[0260] Obtain the frequency response of the first link, and compensate the speech confusion noise with the frequency response.

[0261] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: Before deconfusing the first speech confusion data according to the speech confusion noise to restore the first speech data, the method further includes:

[0262] Receive the second speech data sent by the first user based on the first link, and obtain the prior data of the second speech data;

[0263] According to the second speech data and the prior data, perform channel state estimation on the first link to obtain channel state parameters; Wherein, the channel state parameters include frequency response and / or channel loss rate.

[0264] In one embodiment, when the computer program is executed by a processor, the following steps are implemented:

[0265] Receive the speech input of the first user to obtain the first speech data;

[0266] Determine a first identity that matches the voiceprint information of the first user in the voice dataset;

[0267] Determine a first random number for the current session, and generate voice obfuscation noise based on the first random number, the first identity, and the voice dataset;

[0268] Superimpose the voice obfuscation noise on the first voice data to obtain first voice obfuscated data.

[0269] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: Generate voice obfuscation noise based on the first random number, the first identity, and the voice dataset, including:

[0270] Generate an index sequence according to the first random number;

[0271] Search for the target voice corresponding to the first identity in the voice dataset according to the index sequence;

[0272] Concatenate the target voices to obtain voice obfuscation noise.

[0273] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: Before determining the first random number for the current session, the method further includes:

[0274] Send the first certificate of the first user to the second user;

[0275] Receive the second certificate returned by the second user;

[0276] Generate a first random number, encrypt the first random number and the first identity using the public key of the second user, and send the generated first encrypted data to the second user;

[0277] Receive the second encrypted data returned by the second user, and decrypt the second encrypted data using the private key of the first user; wherein, the second encrypted data is generated by encrypting the second random number and the second identity of the second user using the public key of the first user.

[0278] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: After superimposing the voice obfuscation noise on the first voice data to obtain first voice obfuscated data, the method further includes:

[0279] Send the first voice obfuscated data to the second user; or,

[0280] Send the first voice obfuscated data to the first communication terminal corresponding to the first user, and send it to the second user via the first communication terminal.

[0281] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: Before sending the first voice obfuscated data to the second user, the method further includes:

[0282] Receive the voice input of the first user to obtain second voice data;

[0283] Send the second voice data to the second user based on the first link, and provide the prior data of the second voice data to the second user.

[0284] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0285] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., and are not limited thereto.

[0286] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0287] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.

Claims

1. A voice communication privacy protection method based on voice obfuscation, characterized in that A second voice processing system applied to a second user, the method comprising: Receiving first voice obfuscation data sent by a first user to the second user; Obtaining a first random number corresponding to the first user and a first identity of the first user, where the first identity is an identity in the voice dataset that matches the voiceprint information of the first user; Searching for a target voice in the voice dataset according to the first random number and the first identity, and splicing the target voice to generate voice obfuscation noise; Deobfuscating the first voice obfuscation data according to the voice obfuscation noise to restore first voice data; Wherein, deobfuscating the first voice obfuscation data according to the voice obfuscation noise to restore first voice data includes: Slicing and aligning the first voice obfuscation data and the voice obfuscation noise according to each signal sampling point to obtain multiple groups of corresponding first voice obfuscation data blocks and voice obfuscation noise blocks; Setting a parameter set, where the parameter set includes multiple one-dimensional optimization parameters related to signal frequency; Based on the one-dimensional optimization parameters, setting an objective function, traversing the parameter set, obtaining the minimum value of the objective function, and using the one-dimensional optimization parameters corresponding to the minimum value as target parameters; wherein, the objective function is used to measure the degree of noise residue in the original signal after applying the one-dimensional optimization parameters; Based on the target parameters, calculating an improvement factor corresponding to each frequency sampling point; wherein, the improvement factor is used to measure the optimized noise suppression effect; Weighting the time-frequency spectrum of each of the first voice obfuscation data blocks with the improvement factor to obtain the time-frequency spectrum of the deobfuscated first voice data.

2. The method for protecting voice communication privacy based on voice obfuscation according to claim 1, wherein Based on the one-dimensional optimization parameters, setting an objective function, traversing the parameter set, and obtaining the minimum value of the objective function, including: For each frequency sampling point, performing a scaling process on the voice obfuscation noise block based on the one-dimensional optimization parameters to obtain a first value; Comparing the first value with the corresponding first voice obfuscation data block after taking the absolute value to obtain a ratio, and taking 1 minus the ratio to obtain a second value; Using the second value as a weight to weight each of the first voice obfuscation data blocks to obtain the value of the objective function.

3. The method for protecting voice communication privacy based on voice obfuscation according to claim 1, wherein Based on the target parameters, calculating an improvement factor corresponding to each frequency sampling point, including: For each frequency sampling point, performing a scaling process on the voice obfuscation noise block based on the target parameters to obtain a first value; Comparing the first value with the corresponding first voice obfuscation data block after taking the absolute value to obtain a ratio, and taking 1 minus the ratio to obtain a second value; Comparing the second value with zero, and taking the larger value of the two as the value of the improvement factor.

4. The method for protecting voice communication privacy based on voice confusion according to claim 1, characterized in that, After calculating the improvement factor corresponding to each frequency sampling point based on the target parameters, the method further includes: Obtaining the channel loss rate of a first link, and comparing the improvement factor with the channel loss rate; the first link is a communication link from the first user to the second user; If the improvement factor is less than or equal to the channel loss rate, set the improvement factor to zero.

5. The method for protecting voice communication privacy based on voice confusion according to claim 1, wherein, Before de - confounding the first voice - confounding data according to the voice - confounding noise to restore the first voice data, the method further includes: Obtain the frequency response of the first link, and use the frequency response to compensate the voice - confounding noise.

6. The method for protecting voice communication privacy based on voice confusion according to any one of claims 1 to 5, characterized in that, Before de - confounding the first voice - confounding data according to the voice - confounding noise to restore the first voice data, the method further includes: Receive the second voice data sent by the first user based on the first link, and obtain the prior data of the second voice data; According to the second voice data and the prior data, perform channel state estimation on the first link to obtain channel state parameters; wherein, the channel state parameters include frequency response and / or channel loss rate.

7. A voice communication privacy protection method based on voice obfuscation, characterized in that, Applied to the first voice processing system of the first user, the method includes: Receive the voice input of the first user to obtain the first voice data; Determine the first identity that matches the voiceprint information of the first user in the voice dataset; Determine the first random number of the current session, and generate voice - confounding noise according to the first random number, the first identity, and the voice dataset; Superimpose the voice - confounding noise on the first voice data to obtain the first voice - confounding data; Send the first voice - confounding data to the first communication terminal corresponding to the first user, and then send the first voice - confounding data to the second voice processing system according to claim 1 via the first communication terminal.

8. The method for protecting voice communication privacy based on voice obfuscation according to claim 7, wherein Generating voice - confounding noise according to the first random number, the first identity, and the voice dataset includes: Generate an index sequence according to the first random number; Search for the target voice corresponding to the first identity in the voice dataset according to the index sequence; Stitch the target voice to obtain the voice - confounding noise.

9. A voice processing system, characterized in that, Includes: A processor, a speaker, and a microphone, the speaker and the microphone are respectively connected to the processor; wherein, The processor is used to execute the voice - communication privacy protection method based on voice confounding according to any one of claims 1 to 6; or, the processor is used to execute the voice - communication privacy protection method based on voice confounding according to claim 7 or claim 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it realizes the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Voice interference noise design method based on human voice structure

    CN115841821A