A Personal Voice Information Privacy Protection Method and Device Based on Feature Perturbation
By constructing robust perturbations in personal voice data and generating protected audio, the problem of difficulty in effectively protecting voice-print features in real-time voice transmission in the prior art is solved, and an instant and stable voice privacy protection effect is achieved.
Patent Information
- Application Number
- CN202410761544.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-13
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-06-13
AI Technical Summary
The prior art is difficult to effectively protect the voiceprint characteristics in personal voice in real-time voice transmission, and most methods require long-term model training or support from special physical equipment, making it difficult to construct an instant and stable defense method.
By extracting the MFCC features and X-Vector features of the audio to be protected, a robust perturbation is constructed to generate protected audio so that it is close to the original audio in the audio sample space and the feature vectors of different speaker audio in the feature space.
It realizes the effective protection of personal voice data without reducing audio quality and avoids the business model learning of personal voiceprint features in audio. It has the advantages of immediate and stableness and does not require long-term model training or special physical equipment.
Smart Images

Figure CN118553264B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and particularly relates to a method and device for protecting personal voice information privacy based on feature perturbation. Background Art
[0002] With the rapid development of deep learning technology, deep learning models have been applied to more and more fields, such as autonomous driving, speech recognition, image classification, etc., and play an increasingly important role. The excellent performance of deep learning models depends on a large amount of training data. Therefore, data scraping technologies have become popular, which collect a large amount of personal data without the permission of individual users, and these data will be used to train deep learning models for commercial purposes. However, personal voice data contains privacy information such as voiceprint features. Letting the model learn these features from personal voice data will obviously cause the leakage of personal privacy: for example, the model can generate false synthetic voices based on the voiceprint features of a certain person and be used for illegal activities such as fraud. Given that personal voice will inevitably be uploaded to the Internet, it is very difficult to prevent the corresponding data scraping; therefore, it is necessary to protect personal voice from the perspective of commercial model training, perturb the voiceprint features in the voice, and prevent the privacy information in personal voice from being leaked.
[0003] The protected personal voice should have two characteristics at the same time: usability for people and non-learnability for models. The former means that the protected voice should retain the complete information of the original voice and be understandable to people; the latter means that the model cannot learn useful features from the protected voice. The former requires that the modification of the voice by the protection measure should be inaudible to the human ear, and the latter requires that the protection measure adds incorrect or useless features to the voice. The inventors of the present application found that the existing methods have at least the following technical problems during the implementation of the present invention:
[0004] 1) In the existing technologies for protecting voice data, most of them are dedicated to preventing voice data from being analyzed manually, such as extracting the identity or semantic information of the speaker. These solutions require long-term model training, are difficult to play a role in real-time voice transmission, and cannot effectively protect the voiceprint features of users from being leaked;
[0005] 2) Other voice protection technologies are dedicated to interfering with the process of data scraping programs collecting personal voice, such as emitting interference signals to prevent microphones from recording user voices. This method requires special physical devices to support, and is greatly affected by environmental factors and is very unstable.
[0006] It can be seen from this that existing technologies for protecting personal voice information privacy either require long-term model training or have strict requirements for the physical environment, making it difficult to construct an immediate and stable defense method. So far, the problem of privacy information in personal voice being obtained by commercial models has not been effectively solved. Summary of the Invention
[0007] In view of the deficiencies of the prior art, the present invention provides a method and device for protecting personal voice information privacy based on feature perturbation, which more effectively protects the privacy of personal voice information.
[0008] To achieve the above object, the method for protecting personal voice information privacy based on feature perturbation designed by the present invention specifically includes the following processes:
[0009] S1: Extract the MFCC features of the audio to be protected: For an individual who needs to protect the privacy of voice information, before transmitting personal voice data over the Internet, according to the MFCC feature extractor, perform feature extraction on the voice data.
[0010] S2: Construct a protected audio with MFCC feature perturbation: Select a voice sample of a different speaker, and by adding a small robust perturbation to the original audio, make the audio after adding the perturbation close to the original audio in the audio sample space, while being close to the MFCC feature vector of the audio of a different speaker in the audio feature space; during the process of optimizing the robust perturbation, combine the above two constraints to construct a loss function for the robust perturbation, use frequency masking technology to optimize the perturbation, and set a threshold. If the loss value is less than this threshold, it can be considered that the MFCC features of the perturbed audio are well perturbed; then, output the perturbed audio that meets the above judgment conditions as the intermediate audio for audio protection.
[0011] S3: Extract the X-Vector features of the intermediate audio: To ensure the robustness of the protected audio against different audio feature extractors in the commercial model, use the X-Vector feature extractor to perform feature extraction on the intermediate audio to obtain the X-Vector features of the intermediate audio.
[0012] S4: Construct a protected audio with X-Vector feature perturbation: Use the audio of a different speaker selected in S2, and by adding another robust perturbation to the intermediate audio, make the audio after adding the perturbation close to the original audio in terms of sample distance, while being close to the X-Vector feature vector of the audio of a different speaker in terms of feature vector; the subsequent optimization process is similar to S2, and finally output the perturbed audio whose loss value is less than the set threshold as the final protected audio.
[0013] S5: Use the final protected audio for transmission: Use the protected audio obtained in S4 for transmission over the Internet.
[0014] Specifically, S1 includes the following processes:
[0015] S1.1: For an individual who needs to protect the privacy of voice information, before transmitting their own voice audio over the Internet, the audio x to be transmitted p is used as the audio to be protected, and the feature extractor g1 based on the MFCC algorithm is used to extract the features of x p . This extractor accepts the input of the original audio waveform and outputs the MFCC feature vector of the audio.
[0016] Furthermore, the specific process of S2 is as follows:
[0017] S2.1: First, initialize a random perturbation δ and create the initial protected audio x p1 = x p + δ;
[0018] S2.2: Select an audio x t from a different speaker. The feature vector g1(x t ) obtained by this sample through the MFCC feature extractor g1 is the optimization target of x p1 .
[0019] S2.3: Calculate the following optimization problem:
[0020] min{L(g1(x p1 ), g1(x t )) + β·D(x p1 , x p )}
[0021] where L(g1(x p1 ), g1(x t )) represents the distance between the feature vector of x p1 and the feature vector of x t after being processed by the feature extractor g1; D(x p1 , x p ) represents the distance between x p1 and x p audio that has not been mapped into the feature space. The former determines the similarity of x p1 in the feature space to x t , and the latter determines the similarity of x p1 in the original audio space to x p . β is a hyperparameter used to balance the weights of these two terms. The independent variable of this optimization problem is the value of δ, and the optimization process optimizes the value of δ through loss backpropagation to minimize the value of the above expression. After that, when this value is less than the preset threshold, x p1The original MFCC features are successfully perturbed, and x can be output p1 As the intermediate audio during the protection process
[0022] Preferably, during the optimization process of the perturbation δ in S2.3, in view of the high sensitivity of the human ear to small changes in frequency and amplitude in audio, the frequency masking technique is used in the optimization process of the perturbation δ to calculate the maximum limit of the audio perturbation of δ for human auditory perception, so that δ is truly not easily detectable by the human ear. Frequency masking can determine the best position and maximum threshold of δ with respect to the original audio x p as follows:
[0023] First, perform a short-time Fourier transform on the original audio x p to obtain the spectrum of x p
[0024] Calculate the logarithmic amplitude power spectral density PSD (Power Spectral Density) estimate of x p :
[0025]
[0026] where N represents the window size, represents the k-th point of the spectrum of x p
[0027] Calculate the normalized PSD, that is, NPSD (Normalized Power Spectral Density):
[0028]
[0029] If the NPSD of the perturbation δ is less than then δ is not easily detectable relative to the original audio x p . Calculate the NPSD of δ, that is, in the following way:
[0030]
[0031] where, represents the PSD estimate of the perturbation
[0032] Then, limit below the frequency masking threshold of the original audio sample. The loss of the masking threshold can be calculated by the following distance function:
[0033]
[0034] where, represents xp The global masking threshold, function represents the largest integer less than or equal to x.
[0035] This loss function can better optimize the perturbation δ to make it imperceptible and increase the usability of the protected samples for humans.
[0036] Furthermore, the specific process of S3 is as follows:
[0037] Obtain the intermediate audio x p1 of the signal data, divide it into frames, and extract its 24-dimensional Fbank features;
[0038] Normalize the Fbank features on the sliding window;
[0039] Select a central frame to be analyzed, and then obtain several frames before and after this frame as context information. Concatenate the central frame and the context information in chronological order to obtain a frame set. Concatenate the Fbank features of each frame in the concatenated set to obtain a long feature vector, and this vector is used as the input of the pre-trained DNN model;
[0040] Extract the activation values from the last hidden layer of the pre-trained DNN model, perform L2 regularization, and then accumulate the calculation results to obtain the X-Vector feature g2(x p1 ) of x p1 ).
[0041] Furthermore, the specific process of step S4 is as follows:
[0042] S4.1: First, initialize a random perturbation δ and create the protected audio x p2 = x p1 + δ;
[0043] S4.2: For the x t selected in S2, the feature vector g2(x t ) obtained by this sample through g2 is the optimization target of x p2 ;
[0044] S4.3: Calculate the following optimization problem:
[0045] min{L(g2(x p2 ), g2(x t )) + β·D(x p2 , x p )}
[0046] The optimization logic of this step is similar to that of S2. By backpropagating the loss, x p2 is made close to x t, and is close to x in terms of audio distance p , so as to ensure that the generated protected samples have high generalization performance and can still perturb the corresponding features under the condition of variable feature extraction algorithms. Similarly, set a threshold. When the loss calculated based on the optimized x p2 is lower than this threshold, it can be considered that the x p2 has successfully perturbed its X-Vector feature, and output x p2 .
[0047] Based on the same inventive concept, the present invention also designs an electronic device, including: one or more processors and a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the operations performed by the personal voice information privacy protection method based on feature perturbation.
[0048] Based on the same inventive concept, this solution also provides a computer-readable medium, on which a computer program is stored, and the special feature is that: when the program is executed by a processor, it implements the personal voice information privacy protection method based on feature perturbation.
[0049] One or more of the above technical solutions of the embodiments of the present application have at least one or more of the following technical effects: First, this solution has a good privacy protection effect on personal voice data, which can not only ensure that the protected audio has high audio quality, but also effectively prevent commercial models from capturing the protected audio and learning privacy information such as personal voiceprint features in the audio; furthermore, for commercial models based on different feature extraction algorithms, the protected audio generated by this solution can effectively perturb the corresponding features; in addition, the generation process of the protected audio of this solution does not require training a deep learning model, saving a large amount of computing resources, taking less time, and having good immediacy in the daily audio transmission of users; finally, this solution does not require specific physical devices and environmental requirements and has good adaptability to a variety of application scenarios. Description of the Drawings
[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.
[0051] Figure 1 It is a flowchart of the steps of a personal voice information privacy protection method based on feature perturbation provided by the present invention.
[0052] Figure 2Schematic diagram of the framework for optimizing protected audio through the MFCC feature extractor provided by the present invention.
[0053] Figure 3 Flowchart for obtaining the final protected audio provided by the present invention.
[0054] Figure 4 Flowchart for the user side to transmit protected voice data provided by the present invention. Detailed implementation manners
[0055] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0056] Embodiment 1
[0057] The inventors of the present application have found through a large amount of research and practice that the current literature or technologies have two types of defense means against the theft of personal voice information privacy by commercial models. One is to prevent voice data from being analyzed manually, and the other is to interfere with data scraping programs from collecting personal voice. The former requires long-term model training and it is difficult to ensure the personal privacy in voice data, such as the voiceprint feature is not leaked; while the latter requires special physical devices and is greatly affected by environmental factors. Different from the previous personal voice information privacy defense methods, the present invention perturbs the personal features in the audio, without the need for long-term model training or special physical device support, not only makes the protected audio highly available to people, but also avoids the commercial model from learning the personal privacy features from the protected audio, having the advantages of instantaneity and stability.
[0058] A personal voice information privacy protection method based on feature perturbation designed by the present invention, the system model of this solution consists of two parties, namely:
[0059] User side: For communication and other needs, the user side transmits its own voice data on the Internet, but does not want the privacy of its own voice information to be leaked during voice transmission.
[0060] Data scraping side: The data scraping side can scrape the audio uploaded by the user to the Internet and attempt to train a commercial model based on these audios. Such commercial models are based on variable audio feature extractors and have the ability to learn the personal voice information privacy in the user's audio.
[0061] To prevent the exposure of personal privacy information caused by the unauthorized collection of personal voices by commercial models, the present invention provides a method for protecting the privacy of personal voice information based on feature perturbation, which can prevent the personal privacy in the voice from being learned by commercial models. This method is divided into five stages, as Figure 1 shown, namely: extracting the MFCC features of the audio to be protected, constructing an intermediate audio with MFCC feature perturbation, extracting the X-Vector features of the intermediate audio, constructing a protected audio with X-Vector feature perturbation, and transmitting using the final protected audio.
[0062] In a specific embodiment, it includes:
[0063] S1: Obtain the audio that the user needs to transmit online, and use an MFCC feature extractor to extract the MFCC features of the audio.
[0064] S2: As Figure 2 shown, through the audio to be protected, the MFCC feature extractor, and the voice samples of different speakers, jointly generate an intermediate audio in the audio protection process. In the figure, the solid line part is the forward generation process of the sample, and the dotted line part is the reverse optimization process of the perturbation. The optimization part is mainly composed of three constraints. Among them, Constraint 1 ensures that the protected audio is similar to the voices of different speakers in the audio MFCC feature space; Constraint 2 ensures that the protected audio is similar to the original audio in the audio sample space; Constraint 3 uses the technique of frequency masking to ensure that the protected audio is indistinguishable from the original audio in terms of human ear hearing.
[0065] S3: For the intermediate audio obtained in S2, use an X-Vector feature extractor to extract the X-Vector features of the audio.
[0066] S4: Through the intermediate audio, the X-Vector feature extractor, and the voice samples of different speakers selected in S2, jointly generate the final protected audio. The core optimization algorithm of the protected audio is similar to that of S2.
[0067] S5: Transmit the protected audio output by S4 for voice data on the Internet, and keep the original audio locally.
[0068] In a specific embodiment, the process of extracting the MFCC features of the audio to be protected in S1 is described in detail: When the user needs to send their own voice data, regard this voice data as the audio to be protected, and perform the following processing:
[0069] S1.1 Obtain the signal data and sampling frequency of the audio x p to be protected;
[0070] S1.2 Perform pre-emphasis on the signal data to increase the signal-to-noise ratio in the high-frequency band of the signal;
[0071] S1.3 Frame the signal data with 25 ms as one frame; take the sampling frequency as 16000 Hz, and there are 400 sample points in each frame;
[0072] S1.4 Window the signal data, then perform discrete Fourier transform (DFT) to obtain the frequency-domain signal, and further obtain the two-dimensional matrix energy spectrum F;
[0073] S1.5 Obtain the Mel filter bank G. To ensure that the filter can achieve a good balance between capturing sufficient speech feature information and maintaining computational efficiency. In this embodiment, the filter bank consists of 26 triangular filters with a length of 257. The input signal data will pass through 26 filters, and then the energy of the signal passing through each filter will be calculated;
[0074] S1.6 Multiply the two-dimensional matrix energy spectrum F by the two-dimensional Mel filter G to obtain the spectral energy feature parameter H;
[0075] S1.7 Perform logarithmic operation, discrete cosine transform and cepstrum operation on each element in H to obtain the first set of parameters feat of the MFCC parameters;
[0076] S1.8 Obtain the first-order differential feat′ of feat, which is the second set of parameters of the MFCC parameters; obtain the first-order differential feat″ of feat′, which is the third set of parameters of the MFCC parameters. Concatenate the three sets of parameters together to form the MFCC coefficient g1(x p )
[0077] In a specific embodiment, the process of obtaining the protected audio with MFCC feature perturbation in S2 is described in detail:
[0078] S2.1 For x p , initialize a random perturbation δ, and create the initial protected sample x p1 = x p + δ;
[0079] S2.2 Select a speech sample x t of a different speaker, and the feature vector of this sample under g1 is the optimization target of x p1 ;
[0080] S2.3 Calculate the following optimization problem:
[0081] min{L(g1(x p1 ), g1(x t )) + β·D(x p1 , x p )}
[0082] Among them, L(g1(x p1 ), g1(x t )) represents the distance between the MFCC feature vectors of x p1 after being processed by the feature extractor g1 and the feature vectors of x t , which constitutes Constraint 1; D(x p1 , x p ) represents the distance between x p1 and x p samples when not mapped into the feature space, which constitutes Constraint 2. The former determines the similarity of x p1 in the feature space to x t , and the latter determines the similarity of x p1 to x p in the original sample space. β is a hyperparameter used to balance the weights of these two terms. The independent variable of this optimization problem is the value of δ, and the optimization process optimizes the value of δ through backpropagation to minimize the value of the above expression.
[0083] The optimization of the perturbation δ uses the technique of frequency masking, making δ truly imperceptible. Frequency masking can determine the best position and maximum threshold of the perturbation with respect to the original audio x p , and the specific steps are as follows:
[0084] 1) First, perform a short-time Fourier transform on the original audio x p to obtain the spectrum of x p ;
[0085] 2) Calculate the logarithmic magnitude power spectral density PSD estimate of x p :
[0086]
[0087] Among them, N represents the window size, which is 2048; represents the k-th point of the spectrum of x p , that is, the k-th element of the output array after x p undergoes a short-time Fourier transform;
[0088] 3) Calculate the normalized PSD, that is, NPSD:
[0089]
[0090] 4) If the NPSD of the perturbation δ is less than , then δ is imperceptible. Calculate the NPSD of δ, that is, in the following way:
[0091]
[0092] wherein, represents the PSD estimation value of the perturbation;
[0093] 5) Then, is restricted below the frequency masking threshold of the original audio sample, and the loss of the masking threshold can be calculated through the following distance function:
[0094]
[0095] wherein, represents the global masking threshold of x p , and the function represents the largest integer less than or equal to x. This loss function can better optimize the perturbation δ. This constitutes Constraint 3.
[0096] In a specific embodiment, the process of extracting the X-Vector feature of the intermediate audio in step 3 is described in detail:
[0097] S3.1 Obtain the signal data of the intermediate audio x p1 . Taking 25 ms as one frame, extract its 24-dimensional filter bank features. The time of one frame can also be set as needed, mainly considering factors such as the stability of the spectral features, computational efficiency, and the amount of speech information contained;
[0098] S3.2 Normalize the Fbank features on a 3-second sliding window. Other time sliding windows are also available. In this embodiment, a 3-second sliding window is mainly selected considering factors such as timbre and pronunciation habits. If the window is too short, the speaker's voiceprint information cannot be fully captured; if the window is too long, the audio processing time will increase;
[0099] S3.3 Select a central frame to be analyzed, and then obtain several frames before and after this frame as context information. Concatenate the central frame and the context information in chronological order to obtain a frame set. Concatenate the Fbank features of each frame in the concatenated set to obtain a long feature vector, which is used as the input of the pre-trained DNN model. In this embodiment, five frames of context are used as a frame set, and then, with the frame set as the center, four frames of context are concatenated as a new frame set, and so on until fifteen frames of context are concatenated as a frame set as the input of the pre-trained DNN model;
[0100] Extract the activation values from the last hidden layer of the pre-trained DNN model, perform L2 regularization, and then accumulate the calculation results to obtain the X-Vector feature g2(x p1 ) of x p1 .
[0101] In a specific embodiment, the process of obtaining the protected audio with X-Vector feature perturbation in step S4 is described in detail:
[0102] S4.1 For x p1 , initialize a random perturbation δ, and create the initial protected sample x p2 = x p1 + δ;
[0103] S4.2 Calculate the following optimization problem:
[0104] min{L(g2(x p2 ), g2(xt)) + β·D(x p2 , x p )}
[0105] where L(g2(x p2 ), g2(x t )) represents the distance between the X-Vector feature vectors of x p2 and x t after being processed by the feature extractor g2; D(x p2 , x p ) represents the sample distance between x p2 and x p when not mapped into the feature space. The former determines the similarity of x p2 in the feature space to x t , and the latter determines the similarity of x p2 to x p in the original sample space. β is a hyperparameter used to balance the weights of these two terms. The independent variable of this optimization problem is the value of δ, and the optimization process optimizes the value of δ through backpropagation to minimize the value of the above expression. The frequency masking technique is also introduced in the optimization process to ensure that the protected audio is indistinguishable from the original audio to the human ear.
[0106] Finally, the protected audio x p2 output by S4 is transmitted for voice data on the Internet, and the original audio x p is correspondingly retained locally.
[0107] Generally speaking, this example is divided into five parts: extracting the MFCC features of the target audio, constructing the protected audio based on the MFCC features, improving the X-Vector features of the protected samples, constructing the protected audio based on the X-Vector features, and transmitting using the final protected audio.
[0108] The method of the present invention has the following advantages through experiments:
[0109] The signal-to-noise ratio of the protected audio is above 12, and it has good audio quality;
[0110] For the commercial speaker recognition model VggVox, when 50% of the personal speech data is protected and the model is trained based on this data, the accuracy rate is reduced by more than 30% compared to the model trained based on unprotected audio.
[0111] The audio is protected based on MFCC and X-Vector features. However, when the model is trained using other feature extractors, the corresponding features of the audio can still be perturbed. Using the experimental settings in 2 and changing the feature extractor of the speaker recognition model, for the Mel feature extractor, the accuracy rate is reduced by 26%; for the FFT feature extractor, the accuracy rate is reduced by 15%.
[0112] The protection process of one piece of audio takes less than 1 s, and the efficiency is very high.
[0113] This does not require physical devices and environmental requirements, and this does not require effect verification. For good adaptability, using the experimental settings in 2 and changing the structure and tasks of the commercial model, for the speech command recognition model CNN, the command recognition accuracy rate is reduced by 10%; for the automatic speech recognition model AN4, the word error rate increases by 20%.
[0114] Based on the method of feature perturbation, the present invention applies robust perturbation to the target audio to generate protected audio, making it inaudible to the human ear while making it difficult for the deep learning model trained based on this data set to learn the original target audio features, and safely and efficiently realizing the privacy protection of personal speech information. Its performance is superior to other existing defense methods in the same field.
[0115] Embodiment 2
[0116] Based on the same inventive concept, the present invention also provides an electronic device, including one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in Embodiment 1.
[0117] Since the device introduced in Embodiment 2 of the present invention is the electronic device adopted for implementing the method for protecting personal speech information privacy based on feature perturbation in Embodiment 1 of the present invention, based on the method described in Embodiment 1 of the present invention, those skilled in the art can understand the specific structure and deformation of the electronic device, so it will not be elaborated here. Any electronic device adopted for the method in Embodiment 1 of the present invention belongs to the scope to be protected by the present invention.
[0118] Embodiment 3
[0119] Based on the same inventive concept, the present invention also provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, the method described in Embodiment 1 is implemented.
[0120] Since the device introduced in Embodiment 3 of the present invention is a computer-readable medium adopted for implementing the method for protecting personal voice information privacy based on feature perturbation in Embodiment 1 of the present invention, based on the method introduced in Embodiment 1 of the present invention, those skilled in the art can understand the specific structure and variations of the electronic device, so it will not be elaborated herein. Any electronic device adopted by the method in Embodiment 1 of the present invention falls within the scope of protection of the present invention.
[0121] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A voiceprint privacy protection method based on an anti-disturbance feature extractor, characterized in that: S1: Before personal audio transmission, extract the MFCC features of the audio to be protected; S2: Select different speech samples of a speaker, and add a small robust perturbation to the original audio so that the perturbed audio is close to the original audio in the audio sample space and close to the MFCC feature vector of the audio of different speakers in the audio feature space; in the process of optimizing the robust perturbation, the loss function of the robust perturbation is constructed in combination with the above two restrictions, the perturbation is optimized using the frequency masking technology, and a threshold is set. If the loss value is less than the threshold, it can be considered that the MFCC feature of the perturbed audio is well perturbed; then, the perturbed audio that meets the above judgment conditions is output as the intermediate audio for audio protection; S3: Use the X-Vector feature extractor to extract features of the intermediate audio to obtain X-Vector features of the intermediate audio; S4: Using the audio of different speakers selected in S2, another robust perturbation is added to the intermediate audio, so that the audio after adding the perturbation is close to the original audio in terms of sample distance, and close to the X-Vector feature vector of the audio of different speakers in terms of feature vector; the optimization process is similar to S2, and the perturbation is optimized using the frequency masking technology, and finally the perturbation audio with a loss value less than the set threshold is output as the final protected audio. The specific process of optimizing the perturbation includes: First, the original audio x p Perform short-time Fourier transform to obtain x p Spectrum of Calculate x p The logarithmic magnitude power spectral density PSD estimate is: Where N represents the window size; Represents x p The kth point of the spectrum, that is, x p The kth element of the output array after short-time Fourier transform; Calculate the normalized PSD, also known as NPSD: If the NPSD of the disturbance δ is less than Then δ is not easily detected, and the NPSD of δ is calculated, that is, The way is as follows: in, represents the PSD estimate of the disturbance; Then, The original audio samples are restricted to frequencies below the masking threshold, and the loss of the masking threshold is calculated using the following distance function: in, Represents x p The global masking threshold of It means returning the largest integer less than or equal to x. This loss function can better optimize the perturbation δ. S5: Use the protected audio obtained in S4 for transmission over the Internet.
2. The voiceprint privacy protection method based on the anti-disturbance feature extractor according to claim 1 is characterized in that: The specific process of step S1 is as follows: Get the audio to be protected x p Signal data and sampling frequency; Pre-emphasize the signal data to increase the signal-to-noise ratio in the high-frequency band of the signal; Framing the signal data; The signal data is windowed and then subjected to discrete Fourier transform to obtain the frequency domain signal, and then the two-dimensional matrix energy spectrum F is obtained; Get the Mel filter bank G and calculate the energy of the signal passing through each filter; Multiply the two-dimensional matrix energy spectrum F by the two-dimensional Mel filter bank G to obtain the spectrum energy characteristic parameter H; Each element in H is subjected to logarithmic operation, DCT transformation and ascending cepstrum operation to obtain the MFCC parameter feat; the parameter feat is concatenated with its first-order differential feat′ and second-order differential feat″ to form the MFCC coefficient g1(x p ).
3. The voiceprint privacy protection method based on the anti-disturbance feature extractor according to claim 1 is characterized in that: For audio to be protected x p , initialize a random perturbation δ, create the initial protected sample x p1 =x p +δ; Select a speech sample x from a different speaker t , x t The feature vector under the MFCC feature extractor g1 is x p1 The optimization goal of Then, by calculating the following optimization problem, we can get the intermediate audio x p1 : min{L(g1(x p1 ),g1(x t ))+β·D(x p1 ,x p )} Among them, L(g1(x p1 ),g1(x t )) means that after being processed by feature extractor g1, x p1 The MFCC feature vector of x t The distance between the eigenvectors of p1 ,x p ) indicates that when x is not mapped into the feature space, p1 With x p The distance between samples; the former determines x p1 In feature space, t The similarity of x p1 On the original sample space, p similarity; β is a hyperparameter used to weigh the weight of these two items; the independent variable of the optimization problem is the value of δ, and the optimization process optimizes the value of δ through back propagation so that the expression min{L(g1(x p1 ),g1(x t ))+β·D(x p1 ,x p )} is minimized.
4. The voiceprint privacy protection method based on the anti-disturbance feature extractor according to claim 1 is characterized in that: Step S3 includes: Get the middle audio x p1 The signal data is divided into frames and its 24-dimensional Fbank features are extracted; Normalize the Fbank features on the sliding window; Select a central frame to be analyzed, and then obtain several frames before and after the frame as context information; splice the central frame and context information in chronological order to obtain a frame set, and concatenate the Fbank features of each frame in the set to obtain a long feature vector, which is used as the input of the pre-trained DNN model; Extract the activation value from the last hidden layer of the pre-trained DNN model, perform L2 regularization, and then add up the calculation results to get x p1 The X-Vector feature g2(x p1 ).
5. The voiceprint privacy protection method based on the anti-disturbance feature extractor according to claim 1 is characterized in that: The specific process of S4 includes: For the intermediate audio x p1 , initialize a random perturbation δ, create the initial protected sample x p2 =x p1 +δ; Compute the following optimization problem: min{L(g2(x p2 ),g2(x t ))+β·D(x p2 ,x p )} Among them, L(g2(x p2 ),g2(x t )) indicates that after being processed by the X-Vector feature extractor g2, x p2 The X-Vector eigenvector is the same as x t The distance between the eigenvectors of p2 ,x p ) indicates that when x is not mapped into the feature space, p2 With x p The sample distance between p2 In feature space, t The similarity of x p2 On the original sample space, p similarity; β is a hyperparameter, the independent variable of the optimization problem is the value of δ, and the optimization process optimizes the value of δ through back propagation so that the expression min{L(g2(x p2 ),g2(x t ))+β·D(x p2 ,x p )}, the optimization process also introduces frequency masking technology to ensure that the protected audio is indistinguishable from the original audio to the human ear.
6. An electronic device, characterized in that: include: one or more processors; A storage device for storing one or more programs; When one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 5.
7. A computer readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Identity privacy protection method and system for voice data
CN117831510A