A privacy protection noise generation method based on black box adversarial samples

By constructing a speech dataset and generating final white noise data through gradient updates, the problem of interference noise that cannot be addressed in existing technologies for black-box speech recognition models is solved. This achieves a strong interference effect on the black-box model without affecting normal user use, thereby enhancing the user's ability to choose privacy protection.

CN120748435BActive Publication Date: 2025-11-04ZHEJIANG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511256792.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-11-04
Estimated Expiration
2045-09-04

AI Technical Summary

Technical Problem

Existing technologies cannot effectively generate interference noise for black-box speech recognition models, making it impossible to protect voice privacy without affecting normal user use.

Method used

A speech dataset is constructed to generate initial white noise data. The final white noise data is generated through superposition, data augmentation, and gradient update. The local white-box model is fine-tuned and trained using a multi-intensity noisy speech dataset to generate the final privacy-preserving noise.

Benefits of technology

The generated interference noise can have a stronger interference effect on the black box speech recognition model without interfering with the user's normal use, thereby improving the recognition error rate and enabling users to choose their own privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748435B_ABST
    Figure CN120748435B_ABST
Patent Text Reader

Abstract

The application provides a privacy protection noise generation method based on a black box adversarial sample, and relates to the technical field of noise generation. The method comprises the following steps: constructing a speech data set; randomly generating initial white noise data; superimposing the initial white noise data on target speech in the speech data set to obtain a superimposed speech data set; performing data augmentation on the superimposed speech data set to obtain an augmented data set; inputting the augmented data set into a local white box speech recognition model to obtain final white noise data; amplifying the final white noise data and superimposing the amplified final white noise data on the speech data set to obtain a multi-intensity noise-added speech data set; and fine-tuning the local white box model by using the multi-intensity noise-added speech data set to obtain a final noise generation model for generating final privacy protection noise. The application solves the technical problem that, in the prior art, when the parameters and architecture of a black box model cannot be obtained and only the final recognition result can be obtained, the interference noise cannot be optimized based on adversarial sample technology.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of noise generation, in particular to a privacy protection noise generation method based on black box adversarial samples. BACKGROUND

[0002] With the breakthrough progress of artificial intelligence technologies such as speech recognition and speech synthesis, intelligent Internet of Things terminals (including mobile terminals, wearable devices and smart home systems) with voice interaction functions have been deeply integrated into modern life scenarios, and their market penetration rate continues to rise. It is worth noting that the exponential growth of such devices not only improves the convenience of life, but also derives new security risks. A large number of intelligent terminals equipped with audio recording modules have penetrated into various dimensions of daily life. Due to their closed architecture and long-running characteristics, users have difficulty monitoring the running state of the device in real time, which poses a serious challenge to privacy protection. Criminals can obtain user voice data through remote control means and extract the semantic content of the voice signal with the help of deep learning-driven speech recognition systems, and then implement information theft.

[0003] Therefore, there is an urgent need for a privacy protection technology to return the right to privacy to the user. Such protection technology cannot affect the normal use of voice intelligent devices by users, such as calls, voice commands, etc.; when the user believes that his voice privacy needs to be protected, the user can actively use the protection technology to resist malicious privacy theft, rather than relying solely on the privacy protection promise of the manufacturer.

[0004] In response to this, some scholars have proposed a recording interference scheme based on ultrasonic waves. The basic principle is to inject noise based on the nonlinearity of the device microphone, thereby achieving the effect of interfering with the eavesdropping device without disturbing the user in the environment. Yuxin Chen et al. in Wearable Microphone Jamming designed a wearable bracelet with multiple ultrasonic emission probes that can continuously emit ultrasonic waves to interfere with recording devices in the environment. Lingkun Li et al. in Patronus: Preventing Unauthorized Speech Recordings with Support for Selective Unscrambling designed an ultrasonic emission device that can send frequency conversion noise based on pre-generated keys, allowing authorized recording devices to record while interfering with unauthorized recording devices. Although the above-mentioned interference methods can effectively inject noise into eavesdropping devices, they will interfere with the normal operation of the microphone and affect normal functions such as calls; at the same time, these methods require special hardware devices, which are often not available on the user's smart voice device, so this method is not suitable for daily life scenarios.

[0005] Another speech privacy protection technology starts from the vulnerability of neural networks, generates a small interference noise based on the adversarial sample technology, and interferes with the speech recognition model without affecting the normal use of the user, thereby protecting the speech privacy of the user. Peng Cheng et al. in UniAP: Protecting Speech Privacy With Non-Targeted Universal Adversarial Perturbations designed a privacy protection method based on white-box speech adversarial samples. The method generates non-targeted universal adversarial samples to interfere with the white-box speech recognition model, so that any speech is recognized as incorrect semantic content by the white-box model. Although the above method can effectively interfere with the white-box speech recognition model with known model parameters and architecture, the effect is limited when facing unknown black-box speech recognition models, such as commercial speech recognition models. Considering that the background speech recognition model used by various intelligent devices is mostly commercial black-box models, it is urgent to invent an adversarial sample generation method that can generate interference noise for black-box models. SUMMARY

[0006] In order to overcome the shortcomings of the prior art, the purpose of the present application is to provide a privacy protection noise generation method based on black-box adversarial samples. The present application solves the technical problem that in the prior art, when the parameters and architecture of the black-box model cannot be obtained and only the final recognition result can be obtained, the interference noise cannot be optimized based on the adversarial sample technology.

[0007] To achieve the above-mentioned purpose, the present application provides the following scheme:

[0008] A privacy protection noise generation method based on black-box adversarial samples, comprising:

[0009] constructing a speech data set;

[0010] randomly generating initial white noise data;

[0011] superimposing the initial white noise data and the target speech in the speech data set to obtain a superimposed speech data set;

[0012] performing data augmentation on the superimposed speech data set to obtain an augmented data set;

[0013] inputting the augmented data set into a local white-box speech recognition model, and performing gradient update on the perturbation of the local white-box speech recognition model based on negative CTC loss and two-norm regularization term until the average word error rate reaches a first preset threshold, to obtain final white noise data;

[0014] amplifying the final white noise data and superimposing it with the speech data set to obtain a multi-intensity noise-added speech data set;

[0015] Fine-tuning training is performed on the local white-box model by using the multi-intensity noise-added speech data set to obtain a final noise generation model to generate final privacy protection noise.

[0016] Preferably, the fine-tuning training of the local white-box model by using the multi-intensity noise-added speech data set to obtain a final noise generation model to generate final privacy protection noise comprises:

[0017] The fine-tuning training of the local white-box model by using the multi-intensity noise-added speech data set to obtain a final noise generation model to generate final privacy protection noise comprises:

[0018] A first word error rate between the black-box recognition result and the correct text under the benchmark intensity is calculated.

[0019] A second word error rate between the black-box recognition result and the local white-box recognition result under each amplification intensity is calculated.

[0020] Noise-added speeches and their black-box recognition results with a second word error rate greater than a second preset threshold are screened to form a training pair.

[0021] Preferably, the initial white noise data is superimposed on the target speech in the speech data set to obtain a superimposed speech data set, comprising:

[0022] The length of the initial white noise data and the length of the target speech are determined.

[0023] It is judged whether the length of the initial white noise data is greater than the length of the target speech. If yes, superimposition is directly performed. If not, length splicing is performed on the initial white noise data until the length of the initial white noise data is greater than the length of the target speech and superimposition is performed.

[0024] Preferably, the superimposed speech data set is subjected to data augmentation to obtain an augmented data set, comprising:

[0025] Noise is added to the superimposed speech data set to obtain intermediate data.

[0026] Reverb is added to the intermediate data to obtain an augmented data set.

[0027] The expression of the augmented data set is:

[0028] ;

[0029] Wherein, is the augmented data set, is the target speech, is the superimposed speech data set, wherein, is the augmented data set, target speech, For the superimposed speech data set, repeat indicates a repeat algorithm, that is, the interference noise v is spliced in a loop, shift indicates a shift algorithm, that is, the beginning part of the noise segment is randomly cut off, and the total length of the noise is kept consistent with the audio length, simulating a scenario in which the audio and the noise fail to align, and augment indicates a data augmentation method of adding noise and reverberation data to the overall audio.

[0030] The present application discloses the following technical effects:

[0031] The present application provides a privacy protection noise generation method based on a black box adversarial sample, comprising: constructing a speech data set; randomly generating initial white noise data; superimposing the initial white noise data and target speech in the speech data set to obtain a superimposed speech data set; performing data augmentation on the superimposed speech data set to obtain an augmented data set; inputting the augmented data set into a local white box speech recognition model, and performing gradient update on the perturbation of the local white box speech recognition model based on a negative CTC loss and a two-norm regularization term until the average word error rate reaches a first preset threshold to obtain final white noise data; performing noise amplification on the final white noise data and superimposing it with the speech data set to obtain a multi-intensity noise-added speech data set; and fine-tuning the local white box model using the multi-intensity noise-added speech data set to obtain a final noise generation model to generate final privacy protection noise. The interference noise generated in the present application can achieve stronger interference effect on the speech recognition model under the same energy, can interfere with the speech recognition model without disturbing the normal use of the user, and can greatly improve the word error rate (CER) of the black box speech recognition model. Compared with the existing white box adversarial sample-based speech interference noise, the interference noise generated in the present application can achieve stronger interference effect on the black box speech recognition model under the same energy, greatly improving the word error rate (CER) of the black box speech recognition model. The user does not need to introduce a third party in the process of playing the adversarial sample perturbation for privacy protection, and the user can independently select to use the privacy protection function when speaking, so that the speech privacy right is returned to the user. BRIEF DESCRIPTION OF DRAWINGS

[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0033] Figure 1 A privacy protection noise generation method based on a black box adversarial sample is provided for the embodiments of the present application. DETAILED DESCRIPTION

[0034] With reference to the drawings and embodiments of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of the present application.

[0035] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application will be further described in detail below with reference to the drawings and specific embodiments.

[0036] As shown in the drawings, Figure 1 the present application provides a privacy protection noise generation method based on black box adversarial samples, comprising:

[0037] Constructing a speech data set;

[0038] Specifically, a large amount of speech data containing different speakers and different speech content is obtained to construct a speech data set.

[0039] Randomly generating initial white noise data;

[0040] Superimposing the initial white noise data and the target speech in the speech data set to obtain a superimposed speech data set;

[0041] Data augmentation is performed on the superimposed speech data set to obtain an augmented data set;

[0042] Specifically, interference noise is generated on a locally deployed white-box speech recognition model:

[0043] 1) Initialize the interference noise length L and generate a random white noise of the corresponding length as the initial interference noise ;

[0044] wherein whiteNois represents an algorithm for generating a white noise of a specified length.

[0045] 2) Randomly select N speech x i from the speech data set , superimpose the interference noise and the speech, wherein the interference noise is shorter (for example, L=0.5 seconds), the noise length is greater than the speech length by cyclic replication, and a random offset is applied to simulate the situation that the noise starting point cannot be aligned with the speech starting point in actual use. For example, the noise length L=0.5, the speech length len( )=3.1, then 8 pieces of noise are cyclically spliced to obtain , and a 3.1-second part is randomly cut from the noise starting at 0-0.5 seconds, that is ;

[0046] 3) speech with superimposed interference noise Data augmentation is performed, including adding noise and adding reverberation processing, to obtain:

[0047]

[0048] The input local white-box speech recognition model is recognized to obtain a recognition result An optimization algorithm (adam or SGD, etc.) is used to optimize the interference noise with the goal of moving the model recognition result away from the correct semantic content of the speech:

[0049]

[0050]

[0051] wherein, is the speech corresponding to the correct semantic text, measures the distance between the recognition result of the noisy speech and the correct semantic, and the negative sign reverses the optimization direction, making the recognition result as far away from the correct result as possible, i.e., interfering with the model recognition effect; is the norm penalty term of the disturbance to avoid excessive disturbance amplitude; the disturbance is updated based on gradient descent, wherein is the learning rate; is the augmented data set, is the target speech, is the superimposed speech data set, repeat indicates repeating the algorithm, i.e., cyclically splicing the interference noise v, and shift indicates the offset algorithm, which randomly cuts and deletes the beginning noise segment while keeping the total noise length consistent with the audio length, simulating the scenario where the audio and noise are not aligned; augment indicates the noise and reverberation data augmentation method added to the entire audio, refers to the gradient of the loss function Loss(v) with respect to v.

[0052] 4) continuously optimize until the average character error rate (CER) between the model recognition results of these superimposed noise speeches and the correct semantic text reaches a preset threshold;

[0053] In this step, the three designs of cyclically superimposing short disturbances, adding random offsets when optimizing the disturbance, and optimizing multiple different speeches together make the generated interference noise have good interference effect on any content speech and any playback time;

[0054] The augmented data set is input into the local white-box speech recognition model, and the disturbance of the local white-box speech recognition model is updated based on the negative CTC loss and the two-norm regularization term until the average character error rate reaches a first preset threshold, obtaining the final white noise data;

[0055] The final white noise data is amplified and superimposed with the voice data set to obtain a multi-intensity noise-added voice data set;

[0056] Specifically, the optimized interference noise is amplified by different multiples and superimposed with the voice in the voice data set:

[0057] ;

[0058] A batch of voice data with different intensity noises is formed, wherein The noise amplification coefficient is 1 times, 2 times, 5 times and 10 times, and the black box speech recognition model is input for recognition, and the model recognition result is recorded; In this step, the design of amplifying interference noise can amplify noise energy in the case that the existing interference noise has poor interference effect on the black box model, and find the difference between the recognition boundary of the white box model and the black box model;

[0059] The multi-intensity noise-added voice data set is used to fine-tune the local white box model to obtain a final noise generation model to generate a final privacy protection noise.

[0060] Further, the multi-intensity noise-added voice data set is used to fine-tune the local white box model to obtain a final noise generation model to generate a final privacy protection noise, comprising:

[0061] The multi-intensity noise-added voice data set is used to fine-tune the local white box model to obtain a final noise generation model to generate a final privacy protection noise, comprising:

[0062] The first word error rate between the black box recognition result and the correct text under the reference intensity is calculated;

[0063] The second word error rate between the black box recognition result and the local white box recognition result under each amplification intensity is calculated;

[0064] The noise-added voice and its black box recognition result with a second word error rate greater than a second preset threshold are screened to form a training pair;

[0065] Specifically, the word error rate (CER) between the noise-added voice black box model recognition result under the reference noise intensity and the corresponding voice correct semantic content is calculated; The word error rate (CER) between the noise-added voice black box model recognition result under all noise intensities and the local white box model recognition result is calculated, and the noise-added voice and the corresponding black box model recognition result with a word error rate greater than 0.5 are retained;

[0066] The reserved noise-added speech and the corresponding black-box recognition result are used to fine-tune the local white-box model as a training set, the design compares the difference between the black-box model recognition result and the white-box model recognition result, locates the recognition boundary difference of the two models, and fine-tunes the training of the white-box model by using the black-box model recognition result as the speech label of the noise-added speech data set, so that the white-box model learns to approximate the recognition boundary of the black-box model, so that the interference noise generated for the white-box model can also effectively interfere with the recognition result of the black-box model.

[0067] The above steps are repeated until the character error rate (CER) between the noise-added speech black-box model recognition result of the reference noise intensity and the corresponding speech correct semantic content reaches a third preset threshold, at which time the interference noise is the interference noise generated by the method.

[0068] Specifically, the initial white noise data is superimposed on the target speech in the speech data set to obtain a superimposed speech data set, comprising:

[0069] The length of the initial white noise data and the length of the target speech are determined.

[0070] It is judged whether the length of the initial white noise data is greater than the length of the target speech, if yes, superimposition is directly performed, if not, the length of the initial white noise data is spliced until the length of the initial white noise data is greater than the length of the target speech and superimposition is performed.

[0071] The present application also provides another embodiment:

[0072] A white-box speech recognition model (such as ESPnet) is deployed on a local server to build a speech data set.

[0073] The interference noise is generated by the local white-box model.

[0074] The speech data is superimposed and sent to a commercial black-box speech recognition model (such as Tencent Cloud and Ali Cloud) for recognition.

[0075] The local white-box model is fine-tuned in combination with the black-box model recognition result.

[0076] The interference noise with excellent interference effect on the black-box model is obtained through cyclic iteration.

[0077] The interference noise can be played back by using a smart device (such as a smart phone and a smart speaker) that may steal speech privacy to protect the user's speech privacy.

[0078] The embodiments in the specification are described in a progressive manner, and each embodiment focuses on the difference from other embodiments, and the same or similar parts of each embodiment can be referred to each other.

[0079] The principles and implementation manners of the present application are described by using specific examples in the present application, and the above examples are only used to help understand the method of the present application and its core idea; meanwhile, for the general technical personnel in the art, the specific implementation manners and application ranges will be changed according to the idea of the present application. In conclusion, the content of the present specification should not be understood as the limitation of the present application.

Claims

1. A privacy protection noise generation method based on black box adversarial samples, characterized in that, The method comprises the following steps: constructing a voice data set; randomly generating initial white noise data; superimposing the initial white noise data on target voice in the voice data set to obtain a superimposed voice data set; performing data augmentation on the superimposed voice data set to obtain an augmented data set; inputting the augmented data set into a local white-box voice recognition model, performing gradient update on perturbation of the local white-box voice recognition model based on negative CTC loss and two-norm regularization term until average word error rate reaches a first preset threshold to obtain final white noise data; amplifying the final white noise data and superimposing the amplified final white noise data on the voice data set to obtain a multi-intensity noise-added voice data set; fine-tuning the local white-box model by using the multi-intensity noise-added voice data set to obtain a final noise generation model to generate final privacy protection noise.

2. The privacy-preserving noise generation method based on black-box adversarial samples according to claim 1, characterized in that, The fine-tuning of the local white-box model by using the multi-intensity noise-added voice data set to obtain the final noise generation model to generate the final privacy protection noise comprises: calculating a first word error rate between a black-box recognition result and a correct text under a benchmark intensity; calculating a second word error rate between a black-box recognition result and a local white-box recognition result under each amplification intensity; screening noise-added voice and its black-box recognition result whose second word error rate is greater than a second preset threshold to form a training pair; training the local white-box model by using the training pair until the first word error rate reaches a third preset threshold. The superimposition of the initial white noise data on the target voice in the voice data set to obtain the superimposed voice data set comprises:

3. The privacy-preserving noise generation method based on black-box adversarial samples according to claim 1, characterized in that, determining the length of the initial white noise data and the length of the target voice; judging whether the length of the initial white noise data is greater than the length of the target voice, if yes, directly superimposing, if not, length-splicing the initial white noise data until the length of the initial white noise data is greater than the length of the target voice and then superimposing. The data augmentation on the superimposed voice data set to obtain the augmented data set comprises:

4. The privacy-preserving noise generation method based on black-box adversarial samples according to claim 1, characterized in that, adding noise to the superimposed voice data set to obtain intermediate data; adding reverberation to the intermediate data to obtain the augmented data set; the expression of the augmented data set is ​ ; wherein, for augmenting the data set, for the target speech, for the superimposed speech data set, repeat represents a repetition algorithm, i.e. cyclically splicing the interference noise v, shift represents a shift algorithm, randomly cutting off the beginning part of the noise segment and keeping the total length of the noise consistent with the length of the audio, simulating a scenario where the audio and the noise fail to align; augment represents a data augmentation method of adding noise and reverberation to the overall audio.

Citation Information

Patent Citations

  • Method and apparatus for adversarial audio generation for speech recognition system

    CN114783431A

  • Black box confrontation sample generation method for speech recognition system

    CN119132286A