Method for generating acoustic adversarial samples based on partition perturbation
By partitioning acoustic signals based on time-frequency distribution and adjusting perturbation levels, the method generates adversarial samples with reduced perceptibility and improved attack success rates.
Patent Information
- Application Number
- CN202210092888.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-26
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-01-26
AI Technical Summary
When generating adversarial samples, the existing acoustic scene classification model has the problem that perturbation is easily perceived and the attack success rate is low, especially when the perturbation constraint threshold in the restricted area is large, it is difficult to generate unperceived adversarial perturbation.
The acoustic adversarial sample generation method based on partition perturbation is adopted, and the original sample waveform is converted into an amplitude spectrum through Fourier transform, partitioned according to the time-frequency distribution characteristics of the sound signal, generated partition mask matrix, and used gradient and generation adversarial network to generate partition perturbations, adjust the perturbation thresholds of key areas and non-critical areas to generate smaller total perturbation energy but more efficient adversarial samples.
The attack success rate of adversarial samples is improved, the generated adversarial samples are less likely to be perceived by the human ear, and the time-frequency distribution characteristics are closer to the real samples, which improves the adversarial samples performance of acoustic scene classification.
Smart Images

Figure CN114580462B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of adversarial sample generation, and particularly to an acoustic adversarial sample generation method based on partition perturbation. Background Art
[0002] Studying adversarial sample attacks helps to understand the vulnerability of deep learning models and explore how to defend against attacks. Therefore, studying acoustic adversarial sample attacks is very important for exploring how to improve the adversarial sample robustness of acoustic scene classification models. Although the adversarial sample attack technology based on deep learning has achieved good results when attacking acoustic scene classification models, typical adversarial sample attack technologies generate adversarial samples with a unified standard globally for samples. In the case of pursuing a high attack success rate, there is a problem that the adversarial perturbation in the acoustic signal adversarial samples is easily perceptible.
[0003] In order to better improve the imperceptibility of adversarial samples, there have been many optimization schemes for reducing the perturbation area. The attack method based on gradient iteration in a restricted area first uses an object detection method to restrict the attack target area, and then uses a gradient-based iterative attack method to add perturbations in this area. The escape attack method based on a restricted area uses a perturbation limiter to control the position, size, and shape of the attacked area, optimizes the perturbation by jointly training a transformer and the attacked classifier, and finally adds perturbations to the background or boundary of the target. The method of malicious patches and adding attachments generates adversarial samples by replacing a part of the object in the original sample. The single-pixel attack method achieves a successful attack by using a differential evolution algorithm to change only a small number of pixels.
[0004] However, these methods have two drawbacks: on the one hand, a complex preprocessing step is required to determine the perturbation area, which not only increases the time cost but also has a great impact on the performance of adversarial samples. On the other hand, in order to maintain a high attack success rate, the perturbation constraint threshold in the restricted area is relatively large, which will generate adversarial perturbations with relatively high energy, and the imperceptibility of adversarial samples is not ideal. Summary of the Invention
[0005] Aiming at the above deficiencies of the existing technology, the present invention provides an acoustic adversarial sample generation method based on partition perturbation, which solves the technical problem that the existing adversarial sample generation methods in acoustic scene classification cannot set perturbation constraints according to the time-frequency distribution characteristics of sound signals to generate imperceptible adversarial perturbations.
[0006] The present invention adopts the following technical solutions.
[0007] On the one hand, the present invention provides an acoustic adversarial sample generation method based on partition perturbation, including: converting the original sample waveform into an amplitude spectrum by using Fourier transform;
[0008] On the amplitude spectrum, according to the time-frequency distribution characteristics of the sound signal, a partitioning scheme is adopted for partitioning to obtain a partitioning result;
[0009] According to the partitioning result, a partitioning mask matrix is generated;
[0010] Using the amplitude spectrum and the partitioning mask matrix as inputs, an adversarial sample generation method is used to generate an adversarial sample amplitude spectrum based on partitioning perturbation.
[0011] The method further includes: using Fourier transform to convert the original sample waveform into a phase spectrum; using inverse Fourier transform to convert the original sample phase spectrum and the adversarial sample amplitude spectrum into an adversarial sample waveform based on partitioning perturbation.
[0012] Further, the partitioning scheme includes:
[0013] The region where the amplitude value is greater than u·max is the key region of the amplitude spectrum, and other regions are non-key regions of the amplitude spectrum. Where max is the maximum amplitude value of the original sample amplitude spectrum, u is the partitioning threshold, and u ∈ [0,1].
[0014] Further, the partitioning scheme includes:
[0015] Using an object detection algorithm to detect the region of the target sound on the amplitude spectrum and defining it as the key region of the amplitude spectrum, and other regions are non-key regions of the amplitude spectrum.
[0016] Further, the partitioning scheme includes:
[0017] Using a saliency detection algorithm to detect the salient region on the amplitude spectrum and defining it as the key region of the amplitude spectrum, and other regions are non-key regions of the amplitude spectrum.
[0018] Further, the method for generating a partitioning mask matrix according to the partitioning result includes:
[0019] Let the shape of the amplitude spectrum x be (h, w), where h and w represent the height and width of the amplitude spectrum respectively. The key region mask matrix is m, and the non-key region mask matrix is I is a matrix of all 1s, and the shapes of m, I are both (h, w), then the calculation formula for the element in the i-th row and j-th column of the key region mask matrix m is as follows:
[0020]
[0021] Similarly, for the non-key region mask matrix the calculation formula for the element in the i-th row and j-th column is:
[0022]
[0023] m, satisfies the following relationship:
[0024]
[0025] Further, the method for generating a partition mask matrix according to the partition result includes:
[0026] Let the shape of the amplitude spectrum x be (h, w), where h and w represent the height and width of the amplitude spectrum respectively. The key region mask matrix is m, and the non - key region mask matrix is I is a matrix of all 1s, m, I are both of shape (h, w), then the element m i,j in the key region mask matrix m at the i - th row and j - th column has the following calculation formula:
[0027]
[0028] The non - key region mask matrix the element at the i - th row and j - th column has the calculation formula:
[0029]
[0030] m, satisfies the following relationship:
[0031]
[0032] where, I i,j is the element at the i - th row and j - th column of the all - 1 matrix I, the shape of the amplitude spectrum x is (h, w), h and w represent the height and width of the amplitude spectrum respectively; m, I are both of shape (h, w), (i, j) is the two - dimensional index position on the amplitude spectrum x, x K is the key region on the amplitude spectrum determined by using an object detection algorithm or a saliency detection algorithm, x NK is the non - key region on the amplitude spectrum determined by using an object detection algorithm or a saliency detection algorithm.
[0033] Further, taking the amplitude spectrum and the partition mask matrix as inputs, using the gradient - based adversarial sample generation method to generate an adversarial sample amplitude spectrum based on partition perturbations, specifically including:
[0034] The key region is x K and the non - key region is x NK Set the perturbation threshold for the key region as ε K and the perturbation threshold for the non - key region as ε NK ;
[0035] The attacked classification model f takes the amplitude spectrum x as input and outputs the predicted logical value f(x);
[0036] The loss function J takes the predicted logical value f(x) and the class label y of x as input and calculates the classification loss J(f(x), y);
[0037] Calculate the gradient matrix of the classification loss with respect to x
[0038] Using the sign function sign(), calculate the gradient sign matrix corresponding to the gradient matrix
[0039] Using the key region mask matrix and the perturbation threshold of the key region, calculate the key region perturbation constraint matrix as:
[0040] m ε = m·ε K
[0041] Using the key region perturbation constraint matrix and the gradient sign matrix, calculate the perturbation added in the key region as:
[0042]
[0043] Using the non-key region mask and the perturbation threshold of the non-key region, calculate the non-key region perturbation constraint matrix as:
[0044]
[0045] Using the non-key region perturbation constraint matrix and the gradient sign matrix, calculate the perturbation added in the non-key region as:
[0046]
[0047] Add the perturbations in the key region and the non-key region to obtain the formula for the adversarial perturbation amplitude spectrum δ based on partitioned perturbations as:
[0048] δ = δ K + δ NK
[0049] Add the adversarial perturbation amplitude spectrum based on partitioned perturbations and the original sample amplitude spectrum x to obtain the adversarial sample amplitude spectrum based on partitioned perturbations as:
[0050] x adv = x + δ
[0051] Furthermore, taking the amplitude spectrum and the partition mask matrix as input, use the GAN-based adversarial sample generation method to generate the adversarial sample amplitude spectrum based on partitioned perturbations, including:
[0052] The GAN-based adversarial sample generation method is trained for a total of T rounds, with N iterations in each round and a batch size of B for each iteration. The following is the training process for the nth iteration in the tth round:
[0053] On the amplitude spectrum diagram x of the training set samples, the key region is x K , and the non-key region is x NK ;
[0054] The generation network G n-1 takes the amplitude spectrum x as the input and outputs the adversarial perturbation amplitude spectrum δ n , with the same shape as the amplitude spectrum x:
[0055] δ n = G n-1 (x)
[0056] For the said δ n perform partitioning: multiply the key region mask matrix by the adversarial perturbation amplitude spectrum to obtain the key region perturbation
[0057] Multiply the non-key region mask matrix by the adversarial perturbation amplitude spectrum to obtain the non-key region perturbation
[0058] For the said adversarial perturbation amplitude spectrum δ n perform partitioning processing, including adding the scaling factor w to the key region perturbation K , and adding the scaling factor w to the non-key region perturbation NK , to obtain the adjusted partitioned adversarial perturbation amplitude spectra respectively as:
[0059]
[0060]
[0061] Among them, is the adjusted key region adversarial perturbation amplitude spectrum, is the adjusted non-key region adversarial perturbation amplitude spectrum;
[0062] Add the adjusted partitioned adversarial perturbation amplitude spectra to obtain the adjusted adversarial perturbation amplitude spectrum:
[0063]
[0064] Add the adjusted adversarial perturbation amplitude spectrum and the original sample amplitude spectrum x to obtain the adversarial sample amplitude spectrum x adv : x adv = x + δ n′ .
[0065] Use the adversarial sample amplitude spectrum xadv Calculate the loss \(L\) of the generative network based on the sample amplitude spectrum \(x\). G and the loss \(L\) of the discriminative network D ;
[0066] Using the backpropagation algorithm of the neural network, update the parameters of the generative network and the discriminative network respectively according to the losses of the generative network and the discriminative network, solve the following min-max optimization problem, and obtain the generative network \(G\) and discriminative network \(D\) obtained through the \(n\)-th iteration training:
[0067] \(G, D=\arg\min\) G \(\max\) D \((L\) G + \(L\) D )
[0068] Repeat the above steps until the end of the \(N\)-th iteration in the \(t\)-th round, and obtain the generative network \(G\) t and the discriminative network \(D\) t ;
[0069] At the end of the \(t\)-th round of training, use the trained generative network \(G\) t to output the adversarial perturbation amplitude spectrum corresponding to the samples in the test set, and add it to the corresponding original sample amplitude spectrum to obtain the adversarial sample amplitude spectrum;
[0070] Calculate the signal perturbation energy ratio \(SPR\) corresponding to the amplitude spectrum \(x\) in the test set and the corresponding adversarial sample amplitude spectrum \(x\) adv ;
[0071] On the test set, calculate the average value of \(SPR\) corresponding to all samples as \(MSPR\), and save the generative network \(G\) corresponding to the round with the largest \(MSPR\) as the finally trained generative network \(G\) * ;
[0072] Using the amplitude spectrum \(x\) as the input and the trained \(G\) * , output the adversarial perturbation amplitude spectrum:
[0073] \(\delta\) * = \(G\) * (x)
[0074] Add it to the amplitude spectrum \(x\) to obtain the adversarial sample amplitude spectrum:
[0075] \(x\) adv = \(x+\delta\) *
[0076] Beneficial technical effects achieved by the present invention: The method for generating acoustic adversarial samples based on partition perturbation provided by the present invention utilizes the time-frequency distribution characteristics of acoustic signals to design a partition scheme in the generation of acoustic scene adversarial samples based on deep learning; and uses the partition scheme to calculate a partition mask for dividing the key area and the non-key area. By adjusting the perturbation constraint coefficients of each area, the perturbation energy of the key area is increased, and the perturbation energy of the non-key area is reduced.
[0077] Since the key area generally corresponds to the key sound indicating the category of the acoustic sample, the method for generating acoustic adversarial samples based on partition perturbation adds a perturbation with a larger energy in this area, and the generated adversarial samples are more conducive to causing misjudgment of the classification model. At the same time, due to the auditory masking effect of the human ear, compared with the non-key area, adding the same perturbation in the key area with a larger energy is less likely to be perceived by the human ear, and the generated adversarial samples are less likely to be discovered.
[0078] Compared with the unpartitioned scheme that adds the same perturbation globally with a unified standard, the total perturbation energy added by the method based on partition perturbation is smaller.
[0079] Compared with the method based on restricted areas that only adds perturbations in the key area, the method based on partition perturbation avoids the problem that the perturbation in the key area is too large caused by only adding perturbations in the key area, and the corresponding adversarial samples are easily discovered.
[0080] Based on the above advantages, the proposed method for generating acoustic adversarial samples based on partition perturbation can achieve a higher success rate of attacking adversarial samples with a smaller perturbation energy. The generated adversarial samples have time-frequency distribution characteristics more similar to the original samples, further improving the performance of acoustic scene classification adversarial samples. Description of the Drawings
[0081] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0082] Figure 1 is a schematic diagram of the method for generating acoustic adversarial samples based on partition perturbation provided by an embodiment of the present invention;
[0083] Figure 2 is a schematic diagram of the amplitude spectrogram and the partition mask matrix provided by an embodiment of the present invention, which includes FIGS. (a), (b), (c) and (d). Among them, (a) is the amplitude spectrogram, and (b), (c) and (d) are the partition mask matrices corresponding to the partition thresholds u of 0.5, 0.05, and 0.005 respectively;
[0084] Figure 3It is a schematic diagram of the amplitude spectrum of FGSM generated adversarial samples based on partition perturbation provided by an embodiment of the present invention;
[0085] Figure 4 It is a schematic diagram of the amplitude spectrum of GAN generated adversarial samples based on partition perturbation provided by an embodiment of the present invention;
[0086] Figure 5 It is a flowchart of the amplitude spectrum of FGSM generated adversarial samples based on partition perturbation provided by an embodiment of the present invention;
[0087] Figure 6 It is a flowchart of the amplitude spectrum of GAN generated adversarial samples based on partition perturbation provided by an embodiment of the present invention;
[0088] Figure 7 It is a schematic diagram of a generator network structure provided by an embodiment of the present invention;
[0089] Figure 8 It is a schematic diagram of a discriminator network structure provided by an embodiment of the present invention. Specific embodiments
[0090] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0091] The terms "including" and "having" and any variations thereof in the specification and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0092] Embodiment: An acoustic adversarial sample generation method based on partition perturbation, as Figure 1 shown, includes: converting the original sample waveform into an amplitude spectrum by using Fourier transform;
[0093] On the amplitude spectrum, according to the time-frequency distribution characteristics of the sound signal, a partitioning scheme is adopted for partitioning to obtain a partitioning result;
[0094] According to the partitioning result, a partition mask matrix is generated;
[0095] Taking the amplitude spectrum and the partition mask matrix as inputs, an adversarial sample amplitude spectrum based on partition perturbation is generated by using an adversarial sample generation method.
[0096] In the specific embodiments, the partitioning scheme may optionally adopt one of the following methods.
[0097] Partitioning scheme one: The region where the amplitude value is greater than u·max is the key region of the amplitude spectrum, and other regions are the non-key regions of the amplitude spectrum, where max is the maximum amplitude value of the original sample amplitude spectrum, and u∈[0,1] is the partitioning threshold.
[0098] Partitioning scheme two: Use the target detection algorithm to detect the region of the target sound on the amplitude spectrum, and define it as the key region of the amplitude spectrum, and other regions are the non-key regions of the amplitude spectrum.
[0099] Partitioning scheme three: Use the saliency detection algorithm to detect the salient region on the amplitude spectrum, and define it as the key region of the amplitude spectrum, and other regions are the non-key regions of the amplitude spectrum.
[0100] In the specific embodiments, according to the partitioning result generated by partitioning scheme one, the method for generating the partitioning mask matrix includes:
[0101] Let the shape of the amplitude spectrum x be (h, w), where h and w represent the height and width of the amplitude spectrum respectively. The key region mask is m, and the non-key region mask is I is a matrix of all 1s, and m, The shape of I is both (h, w), then the calculation formula for the element in the i-th row and j-th column of the key region mask m is as follows:
[0102]
[0103] Similarly, for the non-key region mask The calculation formula for the element in the i-th row and j-th column is:
[0104]
[0105] m, Satisfies the following relationship:
[0106]
[0107] According to the partitioning results generated by partitioning schemes two and three, the method for generating the partitioning mask matrix includes:
[0108] The calculation formula for the element m i,j in the i-th row and j-th column of the key region mask matrix m is as follows:
[0109]
[0110] For the non-key region mask matrix The calculation formula for the element in the i-th row and j-th column is:
[0111]
[0112] m, satisfies the following relationship:
[0113]
[0114] wherein, I i,j is the element in the i-th row and j-th column of the all-ones matrix I, the amplitude spectrum x has the shape of (h, w), and h and w respectively represent the height and width of the amplitude spectrum; m, I both have the shape of (h, w), (i, j) is the two-dimensional index position on the amplitude spectrum x, and x K is the key region on the amplitude spectrum determined by using an object detection algorithm or a saliency detection algorithm, and x NK is the non-key region on the amplitude spectrum determined by using an object detection algorithm or a saliency detection algorithm.
[0115] The amplitude spectrum diagram and the partition mask matrix diagram provided by the embodiments of the present invention are as Figure 2 shown.
[0116] The method for generating an adversarial sample amplitude spectrum according to the amplitude spectrum and the partition mask includes:
[0117] The gradient-based adversarial sample amplitude spectrum generation method based on partition perturbation includes methods such as FGSM (Fast Gradient Sign Method) based on partition perturbation, PGD (Projected Gradient Descent), BIM (Basic Iterative Method), etc. Here, taking FGSM based on partition perturbation as an example, the specific method is described, and the principle is as Figure 3 shown, and its overall process is as Figure 5 shown.
[0118] On the amplitude spectrum diagram x, let the perturbation threshold adopted by FGSM be ε. In FGSM based on partition perturbation, the key region is x K , and the non-key region is x NK , and the perturbation thresholds of the key region and the non-key region are respectively set as ε K > ε, ε NK < ε.
[0119] The attacked classification model f takes the amplitude spectrum x as input and outputs a predicted logical value f(x);
[0120] The loss function J takes the predicted logical value f(x) and its class label y as input and calculates the classification loss J(f(x), y);
[0121] Calculate the gradient matrix of the classification loss with respect to x
[0122] Using the sign function sign(), calculate the gradient sign matrix corresponding to the gradient
[0123] Using the key region mask matrix and the key region perturbation threshold, calculate the key region perturbation constraint matrix as:
[0124] m ε = m·ε K
[0125] Using the key region perturbation constraint matrix and the gradient sign matrix, calculate the perturbation added in the key region as:
[0126]
[0127] Using the non - key region mask matrix and the non - key region perturbation threshold, calculate the non - key region perturbation constraint matrix as:
[0128]
[0129] Using the non - key region perturbation constraint matrix and the gradient sign matrix, calculate the perturbation added in the non - key region as:
[0130]
[0131] Add the perturbations in the key region and the non - key region to obtain the formula for the FGSM - based adversarial perturbation amplitude spectrum δ based on partitioned perturbation as:
[0132] δ = δ K + δ NK
[0133] Add the adversarial perturbation amplitude spectrum and the amplitude spectrum to obtain the adversarial sample amplitude spectrum based on partitioned perturbation.
[0134] In the embodiments of the present invention, the FGSM - based acoustic adversarial sample generation method based on partitioned perturbation, in the generation of acoustic scene adversarial samples based on deep learning, designs a partitioning scheme and uses this scheme to calculate the partition mask matrix for dividing the key region and the non - key region. By adjusting the perturbation constraint thresholds ε K , ε NK , increase the perturbation energy of the key region and reduce the perturbation energy of the non - key region. This method can improve the attack success rate of adversarial samples while making the time - frequency distribution characteristics of adversarial samples approximate those of real samples, further improving the performance of acoustic scene classification adversarial samples.
[0135] For PGD, the BIM method can be regarded as an iterative version of FGSM. In each iteration, the method for generating the adversarial perturbation amplitude spectrum is the same as that of FGSM, and the specific method will not be elaborated here.
[0136] In other embodiments, a GAN-based adversarial sample amplitude spectrum generation method based on partition perturbation is adopted, and its principle is as Figure 4 shown, and its overall process is as Figure 6 shown. This method does not limit the specific structures of the generation network G, the discriminator network D, and the classification network f, as well as the design of the loss function L G of the generation network and the loss function L D of the discriminator network. In the following method description, one embodiment is used for illustration, but other possible embodiments are also included. Figure 7 、 Figure 8 respectively give an embodiment structure of the generation network and the discriminator network. As shown in the schematic diagram of the generation network structure in Figure 7 , in the term Conv2d / Deconv2d, a, b, c, d, a represents the convolution kernel size, b represents the stride, c represents the padding size, and d represents the number of input and output channels. As shown in the schematic diagram of the discriminator network structure in Figure 8 , in the term Conv2d, a, b, a represents the convolution kernel size and b represents the number of output channels. In the term Avgpool, a, a represents the size of the pooling window. In the term Linear+Softmax, Linear represents the fully connected layer and Softmax represents the Softmax function.
[0137] Let the amplitude spectrum training set be train and the test set be test. The GAN-based adversarial sample generation method based on partition perturbation is trained for a total of T rounds. Each round is iterated N times, and the batch size for each iteration is B. The following gives the training process of the nth iteration in the tth round:
[0138] The neural network simultaneously receives batch samples for training. Here, B = 1 is taken as an example to illustrate the method, and the value of B has no impact on the method description.
[0139] On a certain training set amplitude spectrum graph x, the key region is x K , and the non-key region is x NK .
[0140] The generation network G n-1 takes the amplitude spectrum x as the input and outputs the adversarial perturbation amplitude spectrum δ n , with the same shape as x:
[0141] δ n = G n-1 (x)
[0142] For the said δ n Perform partitioning: Multiply the key-region mask matrix by the perturbation to obtain the key-region perturbation Multiply the non-key-region mask matrix by the perturbation to obtain the non-key-region perturbation
[0143] For the said δ n Perform partitioning processing: Apply different scaling factors w K , w NK , respectively, to the key-region perturbation and the non-key-region perturbation, and obtain the adjusted partitioned adversarial perturbation amplitude spectra as follows:
[0144]
[0145]
[0146] Add the adjusted partitioned adversarial perturbation amplitude spectra to obtain the adjusted adversarial perturbation amplitude spectrum:
[0147]
[0148] Add the adjusted adversarial perturbation amplitude spectrum and the original sample amplitude spectrum to obtain the adversarial sample amplitude spectrum x adv :
[0149] x adv = x + δ n′
[0150] Use x adv , x to calculate the generator loss L G and the discriminator loss L D , and solve the following min-max optimization problem to obtain the trained G and D:
[0151] G, D = argmin G max D (L G + L D )
[0152] wherein, one embodiment of L D takes the following form:
[0153]
[0154] wherein, E represents the mathematical expectation. The main function of L D is to be able to correctly distinguish the original sample and the adversarial sample. Other L D design methods with this function are also applicable to the GAN-based adversarial sample generation method based on partitioned perturbation.
[0155] L GAn embodiment is as follows:
[0156] L G = L adv + αL fool + βL wav + γL tf
[0157] is a linear combination of four loss functions, where α, β, γ are hyperparameters, and L adv , L fool , L wav , L tf are the classification loss, adversarial loss, waveform loss, and time-frequency loss respectively. The calculation formulas are as follows:
[0158]
[0159] where C is the number of categories; l softmax () represents the softmax cross-entropy loss function; by minimizing L adv , the classification logical value of the classification model for the true category can be minimized, enabling the model to classify adversarial samples as other categories.
[0160] L fool = E x log(1 - D(x adv ))
[0161] Minimizing L fool can guide G to generate adversarial samples closer to the true sample distribution.
[0162] L tf = E x |x adv – x|1
[0163] |·|1 represents the Euclidean 1-norm operation; minimizing L tf can guide G to generate an adversarial perturbation amplitude spectrum with time-frequency amplitude values close to 0, encouraging G to generate x adv with time-frequency characteristics similar to x.
[0164]
[0165] Minimizing L wav can guide G to generate an adversarial perturbation waveform with waveform amplitude values close to 0, encouraging G to generate x adv with waveform characteristics similar to x. x wav respectively represent the audio waveforms corresponding to x adv , x.
[0166] Using the backpropagation algorithm of the neural network, update the parameters of the generator network and the discriminator network with the generator network loss and the discriminator network loss respectively;
[0167] Repeat the above steps until the end of the Nth iteration in the t-th round, and obtain the generator network G trained in the t-th round t and the discriminator network D t ;
[0168] At the end of the t-th round of training, use the trained G t , take the samples in the test set as input, and output the corresponding adversarial perturbation amplitude spectrum; add the adversarial perturbation amplitude spectrum to the corresponding original sample amplitude spectrum to obtain the adversarial sample amplitude spectrum. The formula is:
[0169] test adv ={x adv |x adv =x + G t (x), x ∈ test}
[0170] Calculate the signal perturbation ratio (SignalPerturbation Ratio, SPR) corresponding to the amplitude spectrum x in the test set and the corresponding adversarial sample amplitude spectrum x adv :
[0171] SPR(x adv ) = 20log 10 (P(x) / P(G t (x)))
[0172] Calculate the average value (Mean SPR, MSPR) of SPR corresponding to all samples on the test set:
[0173]
[0174] where num test is the number of samples in the test set. When the training round reaches T, the training ends.
[0175] Save the generator network G corresponding to the round with the largest MSPR as the final trained generator network G * .
[0176] Taking x as the input, using the well-trained G * , output the adversarial perturbation amplitude spectrum:
[0177] δ * =G * (x)
[0178] Adding it to x to obtain the adversarial sample amplitude spectrum:
[0179] x adv= x + δ *
[0180] In the embodiments of the present invention, for the GAN-based acoustic adversarial sample generation method based on partition perturbation, in the generation of acoustic scene adversarial samples based on deep learning, by designing a partition scheme and using this scheme to calculate a partition mask matrix for dividing the key area and the non-key area. By adjusting the perturbation weight coefficients w K , w NK , the magnitude of the perturbation in each area is adjusted, affecting the magnitude of the loss function values of the generation network and the discriminator network; the generation network and the discriminator network are updated through the backpropagation algorithm of the neural network; the updated generation network can gradually increase the perturbation energy in the key area and decrease the perturbation energy in the non-key area. This method can improve the success rate of adversarial sample attacks while making the time-frequency distribution characteristics of the adversarial samples approximate those of the real samples, further enhancing the performance of the acoustic scene classification adversarial samples.
[0181] Next, a method flow for generating an adversarial sample waveform of the FGSM for the ESC50 dataset based on partition perturbation will be introduced in combination with a specific implementation manner of the embodiments of the present invention, as Figure 5 shown. In this embodiment, the ESC50 dataset is used as the training and test dataset. The following steps may be included:
[0182] S501, divide the audio data into training samples and test samples.
[0183] S502, input the training data samples into the Fourier transform module to calculate the magnitude spectrum and the phase spectrum.
[0184] S503, input the magnitude spectrum into the partition mask generation module to generate a partition mask matrix for dividing the key area and the non-key area.
[0185] S504, input the partition mask matrix and the magnitude spectrum into the FGSM-based adversarial sample magnitude spectrum generation module to generate an adversarial perturbation magnitude spectrum.
[0186] S505, add the generated adversarial sample magnitude spectrum and the original sound signal magnitude spectrum to obtain the adversarial sample magnitude spectrum.
[0187] S506, input the adversarial sample magnitude spectrum and the original sound signal phase spectrum into the inverse Fourier transform module to obtain the adversarial sample waveform.
[0188] Next, a method flow for generating an adversarial sample waveform of the GAN for the UrbanSound8K dataset based on partition perturbation will be introduced in combination with another specific implementation manner of the embodiments of the present invention, as Figure 6As shown in the figure. In this embodiment, the UrbanSound8K dataset is used as the training and testing dataset. The following steps may be included:
[0189] S601, divide the audio data into training samples and testing samples.
[0190] S602, input the training data samples into the Fourier transform module to calculate the amplitude spectrum and phase spectrum.
[0191] S603, input the amplitude spectrum into the partition mask generation module to generate a partition mask matrix that divides the key area and non-key area.
[0192] S604, input the partition mask matrix and the amplitude spectrum into the adversarial sample amplitude spectrum generation module based on GAN to generate an adversarial perturbation amplitude spectrum.
[0193] S605, add the generated adversarial sample amplitude spectrum and the original sound signal amplitude spectrum to obtain the adversarial sample amplitude spectrum.
[0194] S606, input the adversarial sample amplitude spectrum and the original sound signal phase spectrum into the inverse Fourier transform module to obtain the adversarial sample waveform.
[0195] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0196] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0197] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide for realizing the functions in the processFigure 1 One process or multiple processes and / or boxes Figure 1 Steps of the functions specified in one box or multiple boxes.
[0198] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims. These all fall within the protection scope of the present invention.
Claims
1. An acoustic adversarial sample generation method based on partition perturbation, characterized in that Comprising: Converting the original sample waveform into an amplitude spectrum using Fourier transform; On the said amplitude spectrum, according to the time-frequency distribution characteristics of the sound signal, using a partitioning scheme to perform partitioning to obtain a partitioning result; Generating a partition mask matrix according to the said partitioning result; Taking the said amplitude spectrum and partition mask matrix as inputs, using an adversarial sample generation method to generate an adversarial sample amplitude spectrum based on partition perturbation; Taking the said amplitude spectrum and partition mask matrix as inputs, using a gradient-based adversarial sample generation method to generate an adversarial sample amplitude spectrum based on partition perturbation, specifically including: The attacked classification model takes the amplitude spectrum as input and outputs a predicted logical value; The loss function takes the predicted logical value and its class label as inputs, calculates the classification loss; calculates the gradient matrix of the classification loss with respect to the amplitude spectrum; uses the sign function to calculate the gradient sign matrix corresponding to the said gradient matrix; Using the key region mask matrix and the perturbation threshold of the key region, calculating the key region perturbation constraint matrix; using the key region perturbation constraint matrix and the gradient sign matrix, calculating the perturbation added in the key region; Using the non-key region mask matrix and the perturbation threshold of the non-key region, calculating the non-key region perturbation constraint matrix; using the non-key region perturbation constraint matrix and the gradient sign matrix, calculating the perturbation added in the non-key region; Adding the perturbations in the key region and non-key region together to obtain an adversarial perturbation amplitude spectrum generated based on partition perturbation; Adding the adversarial perturbation amplitude spectrum generated based on partition perturbation and the original sample amplitude spectrum together to obtain an adversarial sample amplitude spectrum generated based on partition perturbation.
2. The method for generating acoustic adversarial samples based on partition perturbation according to claim 1, wherein The said method further includes: converting the original sample waveform into a phase spectrum using Fourier transform; using inverse Fourier transform to convert the original sample phase spectrum and the adversarial sample amplitude spectrum into an adversarial sample waveform based on partition perturbation.
3. The method for generating acoustic adversarial samples based on partition perturbation according to claim 1, wherein The said partitioning scheme includes: The amplitude value is greater than The area is the key area of the amplitude spectrum, and other areas are non-key areas of the amplitude spectrum. Among them, max is the maximum amplitude value of the original sample amplitude spectrum, and u is the partition threshold, .
4. The method for generating acoustic adversarial samples based on partition perturbation according to claim 1, wherein The said partitioning scheme includes: Using an object detection algorithm to detect the region of the target sound on the amplitude spectrum, defining it as the key region of the amplitude spectrum, and other regions as the non-key regions of the amplitude spectrum.
5. The method for generating acoustic adversarial samples based on partition perturbation according to claim 1, characterized in that The said partitioning scheme includes: Using a saliency detection algorithm to detect the salient region on the amplitude spectrum, and defining it as the key region of the amplitude spectrum, and other regions as the non-key regions of the amplitude spectrum.
6. The method for generating acoustic adversarial samples based on partition perturbation according to claim 3, wherein The method for generating a partition mask matrix according to the said partitioning result includes: Key region mask matrix The element in the i-th row and j-th column of is calculated as follows: , Non-critical area mask matrix The element in the i-th row and j-th column has the following calculation formula: , Satisfy the following relationship: , Among them, is a matrix of all 1s The element in the i-th row and j-th column, amplitude spectrum The shape is , where h and w represent the height and width of the amplitude spectrum respectively; The shapes are all , is the element in the i-th row and j-th column of the amplitude spectrum x, max is the maximum amplitude value of the amplitude spectrum of the original sample, and u is the partition threshold, .
7. The method for generating acoustic adversarial samples based on partition perturbation according to any one of claims 4 or 5, characterized in that The method for generating a partition mask matrix according to the said partitioning result includes: Key region mask matrix The element in the i-th row and j-th column has the following calculation formula: , Non-critical region mask matrix The element in the i-th row and j-th column has the following calculation formula: , Satisfy the following relationship: , Among them, is a matrix of all 1s The element in the i-th row and j-th column, amplitude spectrum The shape is , where h and w represent the height and width of the amplitude spectrum respectively; The shapes are all , is the two-dimensional index position on the amplitude spectrum x, is the key region on the amplitude spectrum, is the non-key region on the amplitude spectrum.
8. The method for generating acoustic adversarial samples based on partition perturbation according to claim 1, wherein Taking the said amplitude spectrum and partition mask matrix as inputs, using a GAN-based adversarial sample generation method to generate an adversarial sample amplitude spectrum based on partition perturbation, including: The GAN-based adversarial sample generation method is trained for a total of T rounds, with N iterations in each round, and the batch size for each iteration is B. The following is the training process for the nth iteration in the tth round: Magnitude spectrum diagram of training set samples above, the key area is , and the non-key area is ; Generation network Taking the magnitude spectrum x as the input, outputting the adversarial perturbation magnitude spectrum , with the same shape as the magnitude spectrum x: , Partition the : Multiply the key region mask matrix by the adversarial perturbation amplitude spectrum to obtain the key region perturbation amplitude spectrum , where is the key region mask matrix; Multiply the non-critical region mask matrix by the adversarial perturbation amplitude spectrum to obtain the non-critical region perturbation amplitude spectrum , where is the non-critical region mask matrix; For the adversarial perturbation amplitude spectrum perform partitioning processing, including adding a scaling factor to the perturbation in the key area respectively , adding a scaling factor to the perturbation in the non-key area , to obtain the partitioned adversarial perturbation amplitude spectrum after adjustment, which are respectively , , Among them, is the spectrum of adversarial perturbation amplitude in the adjusted key region, is the spectrum of adversarial perturbation amplitude in the adjusted non-key region, is the scaling factor added to the perturbation in the key region, is the scaling factor added to the perturbation in the non-key region; Adding the adjusted partition adversarial perturbation amplitude spectra together to obtain the adjusted adversarial perturbation amplitude spectrum: , Add the adjusted adversarial perturbation amplitude spectrum and the original sample amplitude spectrum to obtain the adversarial sample amplitude spectrum : , Using the amplitude spectrum of adversarial samples and the amplitude spectrum of original samples , calculate the loss of the generation network , the loss of the discriminator network ; Using the backpropagation algorithm of the neural network, updating the parameters of the generator network and discriminator network respectively by the generator network loss and discriminator network loss, solving the following max-min optimization problem, to obtain the generator network G and discriminator network D obtained after the nth iteration of training; , Repeat the above steps until the end of the Nth iteration in the t-th round, and obtain the generation network trained in the t-th round and the discriminative network : At the end of the t-th round of training, use the trained generation network , take the samples in the test set as input, and output the corresponding adversarial perturbation amplitude spectrum; add the adversarial perturbation amplitude spectrum to the corresponding original sample amplitude spectrum to obtain the adversarial sample amplitude spectrum; Calculate the magnitude spectrum of the test set and the magnitude spectrum of the corresponding adversarial examples the corresponding signal perturbation energy ratio SPR; On the test set, calculate the average value of the SPR corresponding to all samples as , and save the generation network corresponding to the round with the largest , as the generation network finally obtained by training ; Using the amplitude spectrum x as the input and leveraging the trained , output the adversarial perturbation amplitude spectrum: , Adding to the amplitude spectrum x to obtain the adversarial sample amplitude spectrum: 。 9. The method for generating acoustic adversarial samples based on partition perturbation according to claim 8, characterized in that, The discriminative network loss is expressed as follows: , Among them, represents the mathematical expectation.
10. The method for generating acoustic adversarial samples based on partition perturbation according to claim 8, characterized in that, The generative network loss is expressed as follows: , Among them, is a hyperparameter, are the classification loss function, adversarial loss function, waveform loss function, and time-frequency loss function respectively.
Citation Information
Patent Citations
Face image processing method and device, electronic device and storage medium
CN110689500A