Adversarial Sample Detection Method for Voiceprint Recognition Based on Different Migration Abilities and Decision Boundary Attacks
Detecting adversarial samples in the voiceprint recognition model through data preprocessing and decision-making boundary attack methods, solving the problem of insufficient adversarial samples detection in the prior art and improving the security of the model.
Patent Information
- Application Number
- CN202210659947.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-09
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-06-09
AI Technical Summary
The prior art is difficult to effectively detect and defend against adversarial samples in deep learning systems, especially in the field of voiceprint recognition, resulting in insufficient model security.
Through the construction of data preprocessing and voiceprint recognition models, combining different migration capabilities and decision-making boundary attack methods, adversarial samples are generated and their perturbation ratio is detected. The HopSkipJumpAttack method is used to approximate the decision-making boundary and set thresholds to judge the nature of the sample.
Accurate detection of adversarial samples is achieved, the risk of voiceprint recognition model is reduced, and the security of the model is improved.
Smart Images

Figure CN115188385B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for detecting voiceprint recognition adversarial samples, and this method belongs to the field of deep learning security. Background Art
[0002] With the rapid development of deep learning, deep learning has become one of the most common technologies in artificial intelligence, affecting and changing people's lives in all aspects. Typical applications include smart home, intelligent driving, speech recognition, voiceprint recognition and other fields. However, as a very complex software system, deep learning is also faced with various hacker attacks. Hackers can also threaten property security, personal privacy, traffic safety and public security through deep learning systems. Attacks on deep learning systems usually include the following types. 1. Stealing models. Hackers steal the model files deployed on the server through various advanced means. 2. Data poisoning. Data poisoning for deep learning mainly refers to adding abnormal data to the training samples of deep learning, resulting in the model making classification errors when encountering certain conditions. For example, the backdoor attack algorithm adds a backdoor mark to the poisoned data, making the model poisoned. 3. Adversarial samples. Adversarial samples refer to input samples formed by deliberately adding subtle interferences in the dataset, and such samples cause the model to give a wrong output with high confidence. Simply put, adversarial samples cause deep learning models to make classification errors by superimposing carefully constructed perturbations that are difficult for humans to detect on the original data. The security of deep learning has become an urgent problem that we need to solve today. The defense methods are mainly divided into two categories: adversarial sample defense and adversarial sample detection. The main purpose of adversarial sample defense is to restore the classification label of the adversarial sample to the label of the normal sample; the main purpose of adversarial sample detection is to find adversarial samples in the sample set and remove them. Summary of the Invention
[0003] The present invention aims to overcome the above-mentioned drawbacks of the prior art and provides a method for detecting voiceprint recognition adversarial samples based on the different transfer capabilities of samples and decision boundary attacks. The present invention can accurately detect adversarial samples in data, effectively reduce the risks brought by adversarial samples, and enhance the security of the voiceprint recognition model.
[0004] The technical solution adopted by the present invention to solve its technical problems is as follows: data preprocessing, which preprocesses the speaker voice data we use; building a voiceprint recognition model (one is the target model and the other is the detection model); using several different adversarial attack methods to design an adversarial sample generator with malicious information samples in combination with the target classification model; inputting the adversarial samples and clean samples into the target model and the detection model respectively to obtain the corresponding output labels; for the samples with label changes, we set the perturbation ratio to 0, and for the samples with unchanged labels, we further use the decision boundary attack method to attack them to obtain the adversarial perturbation ratio; finally, compare the perturbation ratio with the threshold we set. Samples higher than the threshold are clean samples, and samples lower than the threshold are adversarial samples. The threshold is selected by using the decision boundary attack method to attack a certain amount of clean sample sets, and selecting a threshold from the obtained adversarial perturbation ratios, which makes the detection accuracy of the detection method high and the false detection rate low.
[0005] A voiceprint recognition adversarial sample detection method based on different transfer capabilities and decision boundary attacks includes the following steps:
[0006] Step 1: Perform data preprocessing on the speaker voice signal;
[0007] Step 2: Build a voiceprint recognition model;
[0008] Step 3: Design adversarial samples according to the voiceprint recognition model;
[0009] Step 4: Obtain the adversarial perturbation ratio of the samples;
[0010] Step 5: Obtain the decision threshold according to the clean samples;
[0011] Step 6: Judge the samples according to the decision threshold.
[0012] Furthermore, step 1 specifically includes: first, extract data from existing voice files (.WAV format), and use the librosa audio processing python toolkit to extract voice data from the speaker voice as follows:
[0013] X i ,sr = librosa(T i ,sr = None),i = 1,2,...n+m (1)
[0014] where X i is the voice data of the i-th speaker voice file extracted, sr is the sampling rate of the voice data, and T i is the i-th speaker voice file.
[0015] Normalize the voice signal dataset, and divide the dataset D into a training set Dtrain With the test set D test , where the data set
[0016] D train ={(X1, Y1), (X2, Y2),...,(X n , Y n )}
[0017] D test ={(X n+1 , Y n+1 ), (X n+2 , Y n+2 ),...,(X n+m , Y n+m )}
[0018] X i =(X i1 , X i2 ,..., X id ), d represents the data length of X i . represents the c-class label. The normalization formula is:
[0019]
[0020] where represents the normalized sample, max(X i ) represents the maximum value among the d sampling points in the sample. After normalization
[0021] Furthermore, step 2 specifically includes: pre-specifying the structure and parameters of the classification model, and keeping them unchanged. The classification model structure sampled in the present invention mainly includes a 1D convolutional layer, a max pooling layer, a batch normalization layer, and a fully connected layer. The specific structure is shown in Table 1. Training is performed using the training data set, and the speaker recognition classification model is as follows:
[0022] Target model:
[0023]
[0024] Detection model:
[0025]
[0026] F target (·) and F detect (·) represent the probability vectors output by the model.
[0027] Furthermore, the adversarial sample described in step 3 is defined as:
[0028]
[0029] where δ i is the perturbation added to the original sample.
[0030] Here, taking the optimization-based adversarial attack method as an example to illustrate the generation of adversarial samples. The optimization-based attack method is essentially a gradient-based adversarial sample generation method.
[0031] The optimization function is defined as:
[0032]
[0033]
[0034] Furthermore, step 4 specifically includes: inputting the mixed sample set (including clean samples and adversarial samples) into the target model and the detection model respectively, to obtain the model outputs and The corresponding model output labels are Compare T detect with T target to check if they are the same. If not, set the perturbation ratio If they are the same, perform a decision boundary attack on this sample. Here we sample the HopSkipJumpAttack (HSJA) decision boundary attack method. The adversarial samples obtained by this attack method will be very close to the decision boundary of the model. The principle of this attack method can be seen in Figure 2 . There are mainly 3 steps: 1 gradient direction estimation; 2 step size search through geometric series; 3 boundary search through binary classification. For untargeted attacks, the HSJA attack first samples an initial adversarial sample x′ t from a uniform distribution, then performs binary classification search to query the decision boundary, and simultaneously updates x′ t →x t ; estimate the gradient direction at the updated adversarial sample x t , and the gradient direction is approximately estimated by the Monte Carlo method; then perform step size search through geometric series, and simultaneously update x t →x′ t+1 ; next perform binary classification search; repeat the above steps until the maximum number of iterations is reached. Finally, obtain the perturbation amount δ i after the attack is successful, then the perturbation ratio is:
[0035]
[0036] Furthermore, step 5 specifically includes: using the HSJA method of decision boundary attack to attack the selected clean sample to obtain the adversarial perturbation ratio i = n + 1, n + 2,..., n + m. Among these perturbation ratios, we select different values as decision thresholds for experiments to observe the detection effect of the method. As shown in Figure 3, it shows the detection performance of DeepFool and FGSM adversarial samples under different threshold conditions. It can be seen that selecting different thresholds has a certain impact on both the detection success rate and the classification accuracy of clean samples in the target model. For FGSM adversarial samples, the threshold has a slightly greater impact, and the decision threshold t is preferably 0.05%.
[0037] Further, step 6 specifically includes: comparing the adversarial perturbation ratio D obtained in step 4 i with the decision threshold t obtained in step 5. Those greater than the threshold are clean samples, and those less than the threshold are adversarial samples. The formula is as follows:
[0038]
[0039] where 0 represents an adversarial sample and 1 represents a clean sample.
[0040] The working principle of the present invention is as follows:
[0041] Perform data preprocessing on the speaker's voice signal: Obtain the original waveform time-domain data of each segment of voice, divide the data into a training set and a test set, and perform normalization processing.
[0042] Build a speaker recognition model: Specify in advance the structure and parameters of the speaker recognition model, and they will no longer change. The dataset applicable to this recognition model is also given in advance, that is, the speaker voice samples, which include the input time-domain waveform data for speaker recognition and the corresponding classification labels. The sample set in the dataset should be able to be predicted and output by the model with high accuracy.
[0043] Design adversarial samples according to the speaker recognition model: Select several commonly used white-box adversarial sample attack methods. Adjust the gradient direction of the input data according to the speaker recognition model parameters, so that when the input sample changes slightly, the speaker recognition model generates incorrect labels.
[0044] Obtain the sample adversarial perturbation ratio: In the stage of training the model, we trained a target model to generate adversarial samples and another detection model to compare with the target model. First, we input the samples (adversarial samples or clean samples) into the two models respectively, compare the output labels obtained by the two models. If the labels are inconsistent, we set the adversarial perturbation ratio of the sample to 0. If the labels are consistent, we use the decision boundary attack method to attack the samples with unchanged labels, and this will obtain an adversarial perturbation ratio.
[0045] Obtain the decision threshold based on clean samples: Select a batch of clean samples, attack them using the decision boundary attack method, obtain the corresponding adversarial perturbation ratios, and select one of these perturbation ratios as the decision threshold. This decision threshold should maximize the detection rate of adversarial samples and minimize the false detection rate of clean samples.
[0046] Judge the samples according to the decision threshold: Compare the perturbation ratio obtained in the step of obtaining the adversarial perturbation ratio of the samples with the set decision threshold. Samples with a perturbation ratio greater than the decision threshold are judged as clean samples, and samples with a perturbation ratio less than the decision threshold are judged as adversarial samples.
[0047] The core of the present invention lies in: 1. The transfer ability of clean samples between models is greater than that of adversarial samples; 2. Generally, in the feature space of the model, adversarial samples are closer to the decision boundary of the model than clean samples. Therefore, the distance required for adversarial samples to cross the decision boundary is smaller. Here, we utilize the characteristics of the HSJA black-box attack method to attack the samples, and use the obtained perturbation amount to approximately replace this distance. The method we proposed first transfers the samples between the target model and the detection model, judges whether the classification label of the samples has changed. If it has changed, set the detection value perturbation ratio to 0. If the label has not changed, attack the samples using the decision boundary attack method to obtain the perturbation ratio. Finally, compare the obtained sample perturbation ratio value with the decision threshold we determined to judge whether it is an adversarial sample. The present invention has achieved extremely high detection performance in the detection of adversarial samples in voiceprint recognition, effectively improving the security of the model.
[0048] The advantages of the present invention are: It can accurately detect adversarial samples in the data, effectively reduce the risks brought by adversarial samples, and enhance the security of the voiceprint recognition model. Description of the Drawings
[0049] Figure 1 is a flow schematic diagram of the method of the present invention.
[0050] Figure 2 a~ Figure 2 d is a schematic diagram of the principle of the decision boundary attack (HSJA) method of the method of the present invention, where Figure 2 a is to perform binary classification search after selecting the initial adversarial sample, Figure 2 b is to estimate the gradient perturbation direction, Figure 2 c is to take a step forward in the perturbation direction, Figure 2 d is to perform binary classification search again.
[0051] Figures 3a to 3b is the detection effect diagram under different thresholds of the method of the present invention, where Figure 3a is the effect of FGSM adversarial samples, Figure 3bIt is the effect of DeepFool adversarial samples. Specific implementation mode
[0052] The technical solution of the method of the present invention will be further described below in conjunction with the accompanying drawings
[0053] Example 1:
[0054] A voiceprint recognition adversarial sample detection method based on different migration abilities and decision boundary attack methods, comprising the following steps:
[0055] (1) Perform data preprocessing on the speaker voice signal;
[0056] First, extract data from existing voice files (.WAV format). We use the librosa audio processing python toolkit to extract voice data from the speaker's voice as follows:
[0057] X i ,sr = librosa(T i ,sr = None),i = 1,2,...n+m (1)
[0058] where X i is the voice data of the i-th speaker's voice file, sr is the sampling rate of the voice data, and T i is the i-th speaker's voice file.
[0059] Normalize the voice signal dataset and divide the dataset D into a training set D train and a test set D test , where the dataset
[0060] D train ={(X1,Y1),(X2,Y2),...,(X n ,Y n )}
[0061] D test ={(X n+1 ,Y n+1 ),(X n+2 ,Y n+2 ),...,(X n+m ,Y n+m )}
[0062] X i =(X i1 ,X i2 ,...,X id ),d represents the data length of X i , and represents the c-class label. The normalization formula is:
[0063]
[0064] wherein represents the normalized sample, and max(X i ) represents the maximum value among the d sampling points in the sample. After normalization
[0065] (2) Build a voiceprint recognition model; predetermine the structure and parameters of the classification model, which do not change. The classification model structure sampled in the present invention mainly includes a 1D convolutional layer, a max pooling layer, a batch normalization layer, and a fully connected layer. The specific structure is shown in Table 1. Train using the training data set, and the voiceprint recognition classification model is as follows:
[0066] Target model:
[0067]
[0068] Detection model:
[0069]
[0070] F target (·) and F detect (·) represent the probability vectors output by the model.
[0071] (3) Design adversarial samples according to the voiceprint recognition model;
[0072] The adversarial sample is defined as:
[0073]
[0074] where δ i is the perturbation added to the original sample.
[0075] Here, taking the optimization-based adversarial attack method as an example to illustrate the generation of adversarial samples. The optimization-based attack method is essentially a gradient-based adversarial sample generation method.
[0076] The optimization function is defined as:
[0077]
[0078]
[0079] (4) Obtain the adversarial perturbation ratio of the sample; Input the mixed sample set (including clean samples and adversarial samples) into the target model and the detection model respectively to obtain the model outputs and Then the corresponding model output labels are Compare T detectWith T target Is it consistent? If not, set the perturbation ratio If it is consistent, perform a decision boundary attack on this sample. Here, we sample the HopSkipJumpAttack (HSJA) decision boundary attack method. The adversarial examples obtained by this attack method will be very close to the decision boundary of the model. The principle of this attack method can be seen in Figure 2 . There are mainly three steps: 1 Gradient direction estimation; 2 Step size search through geometric series; 3 Boundary search through binary classification. For untargeted attacks, the HSJA attack first samples an initial adversarial example x′ from a uniform distribution t , and then performs binary classification search to query the decision boundary while updating x′ t →x t ; Estimate the gradient direction at the updated adversarial example x t ; The gradient direction is approximately estimated by the Monte Carlo method; Then perform step size search through geometric series while updating x t →x′ t+1 ; Next, perform binary classification search; Repeat the above steps until the maximum number of iterations is reached. Finally, obtain the perturbation amount δ after the attack is successful i , then the perturbation ratio is:
[0080]
[0081] (5) Obtain the decision threshold according to the clean sample; Use the HSJA method of decision boundary attack to attack the selected clean sample to obtain the adversarial perturbation ratio i = n + 1, n + 2,..., n + m. Among these perturbation ratios, we select different values as the decision threshold for experiments to see the detection effect of the method. As shown in Figure 3, it shows the detection performance of DeepFool and FGSM adversarial examples under different threshold conditions. It can be seen that selecting different thresholds has a certain impact on both the detection success rate and the classification accuracy of clean samples in the target model. For FGSM adversarial examples, the threshold has a slightly greater impact. Finally, the decision threshold t we set is 0.05%.
[0082] (6) Judge the sample according to the decision threshold; Compare the adversarial perturbation ratio obtained in step 4 with the decision threshold t obtained in step 5. Those greater than the threshold are clean samples, and those less than the threshold are adversarial samples. The formula is as follows:
[0083]
[0084] Among them, 0 represents an adversarial sample, and 1 represents a clean sample.
[0085] Example 2: Data in actual experiments
[0086] (1) Select experimental data
[0087] The dataset used in the experiment is the AISHELL-1 speech dataset. This dataset collects the speeches of speakers of different ages, genders, and regions recorded in a quiet environment, with a sampling rate of 16,000. We select the speeches of 20 people as the dataset for the voiceprint recognition model. For each speech, the length of the original waveform time-domain data we extract is 60,000. The data preprocessing is saved as a dataset in the form of an array of (batchsize, 60,000, 1) and the corresponding label data is generated. The processed datasets are all saved as.npy files.
[0088] (2) Experimental results
[0089] The present invention uses 5 attack algorithms (FGSM, BIM, PGD, DeepFool, CW) to generate 5 types of adversarial samples, uses the detection method proposed by the present invention to detect the 5 types of adversarial samples, uses the detection success rate ACC and the false detection rate FPR as the detection effect of the detection method, and makes a comparison with other detection methods. The experimental results are shown in Table 2 and Table 3.
[0090] Table 1 Voiceprint recognition model structure
[0091]
[0092]
[0093] Table 2 Detection success rate
[0094]
[0095] Table 3 False detection rate
[0096]
[0097] The content described in the embodiments of this specification is only an enumeration of the implementation forms of the inventive concept. The protection scope of the present invention should not be regarded as limited to the specific forms stated in the embodiments. The protection scope of the present invention also extends to equivalent technical means that those skilled in the art can think of based on the inventive concept of the present invention.
Claims
1. A detection method for voiceprint recognition adversarial samples based on attacks with different migration capabilities and decision boundaries, characterized in that, It includes the following steps: Step 1: Perform data preprocessing on the speaker's voice signal; Step 2: Build a voiceprint recognition model; Step 3: Design adversarial samples according to the voiceprint recognition model; Step 4: Obtain the adversarial perturbation ratio of the samples. Step 4 specifically includes: A mixed sample set containing clean samples and adversarial samples are respectively input into the target model and the detection model to obtain model outputs and Then the corresponding model output labels are Compare T detect with T target to see if they are consistent; if not, set the perturbation ratio If they are consistent, perform a decision boundary attack on this sample; here, the HopSkipJumpAttack (HSJA) decision boundary attack method is adopted. The adversarial samples obtained by this attack method will be very close to the decision boundary of the model; there are 3 steps:
1. Gradient direction estimation; 2. Step size search through geometric series; 3. Boundary search through binary classification; for untargeted attacks, the HSJA attack first samples an initial adversarial sample x' from a uniform distribution t , then performs binary classification search to query the decision boundary and updates x' t →x t ; estimate the gradient direction at the updated adversarial sample x t . The gradient direction is approximately estimated by the Monte Carlo method; then perform step size search through geometric series and update x t →x' t+1 ; next, perform binary classification search; repeat the above steps until the maximum number of iterations is reached; finally, obtain the perturbation amount δ after the attack is successful i , then the perturbation ratio is: Step 5: Obtain the decision threshold according to the clean samples; Step 6: Judge the samples according to the decision threshold.
2. The detection method of voiceprint recognition adversarial samples based on attacks with different migration capabilities and decision boundaries according to claim 1, characterized in that, Step 1 specifically includes: First, extract data from the existing voice files. We use the librosa audio processing python toolkit to extract voice data from the speaker's voice as follows: X i , sr = librosa(T i , sr = None), i = 1, 2,... n + m (1) where X i is the speech data of the i-th speaker's speech file, sr is the sampling rate of the speech data, T i is the i-th speaker's speech file; Normalize the speech signal dataset and divide the dataset D into a training set D train and a test set D test , where the dataset D train = {(X1, Y1), (X2, Y2),..., (X n , Y n )} D test = {(X n+1 , Y n+1 ), (X n+2 , Y n+2 ),..., (X n+m , Y n+m )} X i =(X i1 , X i2 ,..., X id ), where d represents the data length of X i , represents a class c label; the normalization formula is: Among them represents the normalized sample, and max(X i ) represents the maximum value among the d sampling points in the sample. After normalization 3. The detection method of voiceprint recognition adversarial samples based on attacks with different migration capabilities and decision boundaries according to claim 1, characterized in that, Step 2 specifically includes: The voiceprint recognition model structure includes a 1D convolutional layer, a max pooling layer, a batch normalization layer, and a fully connected layer; it is trained using the training dataset. The voiceprint recognition classification model is as follows: Target model: Detection model: F target (·) and F detect (·) represents the probability vector output by the model.
4. The detection method of voiceprint recognition adversarial samples based on attacks with different migration capabilities and decision boundaries as claimed in claim 1, wherein The adversarial samples described in Step 3 are defined as: where δ i is the perturbation added to the original sample; The optimization function is defined as:
5. The detection method of voiceprint recognition adversarial samples based on attacks with different migration abilities and decision boundaries as claimed in claim 1, wherein Step 5 specifically includes: Attack the HSJA method with the decision boundary to attack the selected clean samples and obtain the adversarial perturbation ratio Among these perturbation ratios, select different values as decision thresholds for experiments.
6. The detection method of voiceprint recognition adversarial samples based on attacks with different migration capabilities and decision boundaries according to claim 5, characterized in that The decision threshold t is set to 0.05%.
7. The detection method of voiceprint recognition adversarial samples based on attacks with different migration capabilities and decision boundaries according to claim 1, characterized in that, Step 6 specifically includes: Compare the adversarial perturbation ratio obtained in step 4 with the decision threshold t obtained in step 5. Samples greater than the threshold are clean samples, and samples less than the threshold are adversarial samples. The formula is as follows: Among them, 0 represents the adversarial sample, and 1 represents the clean sample.
Citation Information
Patent Citations
Multi-model cooperative defense method facing deep learning antagonism attack
CN108446765A
Voiceprint recognition confrontation sample generation method based on boundary attack
CN113571067A