A Sample-Independent Backdoor Attack Method for Federated Learning of Voiceprint Recognition
By selecting ghost neurons under the federated learning framework and artificially mapping their values and designated tags during the training process, a backdoor attack on voiceprint recognition that does not depend on samples is realized, solving the problem of insufficient research on sample modification dependence and speech fields in the existing technology, which is highly concealed and highly threatening.
Patent Information
- Application Number
- CN202411063803.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-05
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-08-05
AI Technical Summary
In the prior art, most backdoor attack methods of voiceprint recognition models require modification of samples, and there is a lack of research in the field of speech, so it is impossible to effectively implement backdoor attacks that do not rely on samples under the federated learning framework.
A backdoor attack method for federated learning voiceprint recognition that does not rely on samples was designed. A federated learning framework was constructed through a central server. The enemy client selected ghost neurons and artificially mapped the values of ghost neurons with the specified tags during the training process, thereby implanting the backdoor.
This method is highly concealed, does not rely on samples, and does not require modification of input. It can be implanted with a backdoor when using an absolutely safe data set. The backdoor cannot be discovered through the review of the data set, which is of great significance to improving the federated learning voiceprint recognition backdoor defense method.
Smart Images

Figure CN118784350B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence security, specifically relates to the field of backdoor attacks on voiceprint recognition in artificial intelligence, and more specifically, to a method for backdoor attacks on voiceprint recognition in federated learning that does not rely on samples. Background Art
[0002] In the field of speech, deep learning technology has been widely applied in speech recognition and voiceprint recognition. The rapid development of these two technologies has greatly facilitated daily life. For example, when verifying mobile banking or WeChat, there is no longer a need to remember passwords with different rules stipulated by different banks. One only needs to input the user's voiceprint to unlock the APP by voice. However, the deep learning technology used in the voiceprint recognition model itself has some problems that are easily exploited for attacks, such as backdoor attacks. A backdoor attack is to artificially map specified features to specific classes during the training process, so that when the model is used, identifying the specified features can output a specific wrong class.
[0003] However, currently, a large number of backdoor attack methods are mostly applied in the field of images. For example, the patent application with the application number 202310652261.0 discloses "A Training Method for Physical Lighting Backdoor Attacks for Artificial Intelligence Security", which conducts a lighting backdoor attack on the target object and triggers the backdoor using lights of different colors on the image. The patent application with the application number 202010077148.0 discloses "A Backdoor Attack Method for Video Analysis Neural Network Models". For more stringent backdoor attack implementation environments such as the high sample dimension, high frame resolution, and sparse data sets of videos, a framework for constructing video backdoor contaminated samples is used to conduct backdoor attacks on video analysis neural network models.
[0004] The above two patents both belong to the backdoor attack methods of artificial intelligence systems, and both are backdoors in the field of images. In the field of speech, the patent application with the application number 202310533603.7 discloses "A Speech Backdoor Verification Method Based on Room Impulse Response", which uses a dynamic trigger to poison clean speech samples and inject backdoors. The patent application with the application number 202310161820.8 discloses "A Speech Backdoor Attack Method and a Computer Readable Storage Medium", which generates a trigger that does not add new noise data relative to the poisoned sample for backdoor injection. The goals of the speech field and the voiceprint recognition field of the present invention are different. The voiceprint recognition system discriminates the identity of the speaker through speech, rather than the content of the speech. In the field of voiceprint recognition, the research on backdoor attacks is still relatively lacking. The main research includes that Kong et al. used adversarial audio as a trigger in "Adversarial audio: a new information hiding method and backdoor for dnn-based speech recognition models" to implant a backdoor in a deep neural network. The usual backdoor attack model is multi-classification, that is, attacking speaker recognition rather than speaker verification. Since each registered user in speaker recognition has an independent model, it is not easy to attack the models of other registrants. Therefore, Zhai et al. proposed a clustering method in "Backdoor attack against speaker verification", inserting a fixed trigger into each class of audio. For different clusters, the backdoor is activated through different triggers. This method can extend the backdoor to speaker verification and is also effective for the audio of users who have not registered.
[0005] However, when implanting the backdoor using the above methods, it is necessary to modify the samples. Therefore, the attack method designed by the present inventor is different from the above patents - a federated learning voiceprint recognition backdoor attack method that does not rely on samples. After retrieval, there are no identical patent documents. Summary of the Invention
[0006] The purpose of the present invention is to provide a sample-independent backdoor attack method for federated learning voiceprint recognition. This method is aimed at the voiceprint recognition system under the federated learning framework and uses a sample-independent method. First, a target model is determined by the central server and distributed; the adversary client designates ghost neurons in the target model, and statistically analyzes the distribution of values on the ghost neurons, and selects a value with an appropriate occurrence frequency as the backdoor. During the training process, the values of the ghost neurons are artificially mapped to the designated labels, thereby implanting the backdoor. The advantages of the present invention are strong concealment, independence from samples, no need to modify the input, and the ability to implant the backdoor when using an absolutely secure data set, and the backdoor cannot be discovered by reviewing the data set, which is of great significance for improving the backdoor defense method of federated learning voiceprint recognition.
[0007] The technical solution of the present invention is as follows:
[0008] A sample-independent backdoor attack method for federated learning voiceprint recognition. First, the central server constructs a federated learning framework and a target model for voiceprint recognition; the target model is distributed to the clients in the framework; the adversary selects ghost neurons, and at the same time all clients expand the data set and perform pre-training; the adversary selects the backdoor activation value of the ghost neurons based on the values obtained from the pre-training; during the backdoor training process, the values of the ghost neurons and the target labels are modified, thereby implanting the backdoor into the model; in the non-backdoor stage, the values and labels of the ghost neurons are not modified; finally, the central server obtains the model containing the backdoor by aggregating the parameters uploaded by the clients; including the following steps:
[0009] Step 1: The central server S constructs a federated learning framework and uses SincNet as the global voiceprint recognition model; after initializing the global model, the model is distributed to all clients C;
[0010] Step 2: The adversary client C adv Checks the SincNet model structure and selects some neurons in the model as ghost neurons; all clients preprocess the local private data they hold and divide the data set according to the size of the sliding window and the sliding distance to expand the data set;
[0011] Step 3: The clients use the local private data set for pre-training, record the values of each sample on the ghost neurons, and draw a histogram of the recorded values;
[0012] Step 4: Selects the backdoor activation value of the ghost neurons according to the distribution of the values on the ghost neurons as the trigger of the backdoor;
[0013] Step 5: Start training and determine whether it is an attack round. If it is the current attack round, perform backdoor implantation. The adversary modifies the values of the ghost neurons to the backdoor activation values and changes the labels to the labels specified by the adversary. If it is not the current attack round, do not perform backdoor implantation. During training, keep the values of the ghost neurons as the values obtained from the natural calculation of the samples and keep the labels as the true labels.
[0014] Step 6: The benign client uploads the clean model parameters, and the adversary client uploads the model parameters implanted with the backdoor. The central server aggregates the models to obtain a federated learning voiceprint recognition model containing the backdoor.
[0015] The more specific steps are as follows:
[0016] Step 1, the central server S constructs a federated learning framework and uses SincNet as the global voiceprint recognition model G. The central server randomly selects 10 from the union of all clients as the client set C = {C 1 , C 2 , C 3 , ···, C adv , C N} for the current round. After initializing the global model, the model is sent to all clients in C, and each round of C contains an adversary C adv .
[0017] Step 2, the adversary client C adv views the structure of the global model G and selects some neurons in the model as ghost neurons: in the last hidden layer of the model, starting from the 1st neuron, select 1, 2, 5, 10, 15, 20, 30, 50 consecutive neurons as ghost neurons; all clients preprocess the locally held private data, convert the audio content into WAV format, set the length of the audio to 16000 frames, and according to the size of the sliding window win size and the sliding distance win long , split the dataset and expand the dataset; set win size = 200, win long = 10. The sliding window starts sliding from the 0th frame. Each time it slides, the audio within the window is saved, and then it slides backward by one win long ; finally, a new expanded training dataset D = {D 1 , D 2 , D 3 , ···, D n} is obtained, where n is the number of samples in the local dataset.
[0018] Step 3, all clients use the segmented local dataset D for pre-training. For the i-th sample D i , the model will generate a value at the ghost neuron; the client records the value of each sample on each ghost neuron, generating a set of values, g is the index of the ghost neuron; according to V, a histogram of the recorded values is plotted for each ghost neuron, and two decimal places are retained during recording.
[0019] Step 4, according to the value distribution shown in the histogram, find a value that exists in the histogram with the same proportion as the adversary's expected backdoor activation probability, and set this value as the activation value of the ghost backdoor.
[0020] Step 5, start training. Determine that the frequency of backdoor implantation during training is N. Calculate the relationship between the current round epoch and the backdoor implantation frequency N to determine whether the current is an attack round: If epoch is a multiple of N, then the current is an attack round, and backdoor implantation is performed. During the backdoor implantation process, the adversary will, through matrix operation Ⅰ, set the value of the ghost neuron to the activation value of the ghost backdoor set in Step 4, and at the same time modify the label of each sample to the specified label. D' is the sample after modifying the label, where W epoch is the parameter matrix of the model in the epoch-th round, with a specific size of r rows and d columns, is a mask matrix of r rows and d columns, is a backdoor activation matrix of r rows and d columns; then the model will use and D' to update according to Ⅱ, where L is the cross-entropy loss function used by the model, is to calculate the gradient of the model according to the loss value calculated by the loss function, η is the learning rate; if the current round epoch is not a multiple of N, then the current is not an attack round, and no backdoor implantation is performed. Normal training is carried out according to Ⅲ. During training, keep the value of the ghost neuron as the value obtained by natural calculation of the sample and keep the label as the true label. The formulas are as follows:
[0021]
[0022]
[0023]
[0024] Step 6, the benign clients upload the clean model parameters, and the adversary clients upload the model parameters implanted with the backdoor. The central server aggregates the models through the PartFedAvg method Ⅳ to obtain the federated learning voiceprint recognition model containing the backdoor is the model parameter of the k-th client in the (epoch + 1)-th round, It is the global model aggregated according to the model provided by the client in the $epoch$ - th round. The formula is as follows:
[0025]
[0026] The present invention has the following features:
[0027] 1. The present invention improves the traditional backdoor attack method without modifying the input samples, so that this method cannot be discovered by examining the training data.
[0028] 2. The present invention places the trigger during the training process, designates the value of the neuron in the model as the trigger, so that the model achieves a random trigger effect after training, and the reason for the appearance of the backdoor cannot be deduced during the use of the model, making the appearance of the backdoor more concealed. Brief Description of the Drawings
[0029] Figure 1 is the business flow chart of the present invention;
[0030] Figure 2 is the training flow chart of the present invention;
[0031] Figure 3 is the federated learning framework diagram of the present invention;
[0032] Figure 4 is the training model structure of the present invention;
[0033] Figure 5 is the schematic diagram of the operation method of the ghost neuron of the present invention;
[0034] Figure 6 is the curve graph of the success rate of backdoor activation of the present invention. Detailed Embodiment
[0035] The present invention will be further described below with reference to the drawings and embodiments.
[0036] Refer to Figure 1-4 , a federated learning voiceprint recognition backdoor attack method that does not rely on samples. First, a central server constructs a federated learning framework and a target model for voiceprint recognition; the target model is sent to the clients in the framework; the adversary selects ghost neurons, and at the same time all clients expand the dataset and perform pre - training; the adversary selects the backdoor activation value through the value of the ghost neurons obtained from the pre - training; during the backdoor training process, the values of the ghost neurons and the target labels are modified to implant the backdoor into the model; in the non - backdoor stage, the values and labels of the ghost neurons are not modified; finally, the central server obtains the model containing the backdoor by aggregating the parameters uploaded by the clients; including the following steps:
[0037] Step 1: The central server S constructs a federated learning framework and uses SincNet as the global model for voiceprint recognition. After initializing the global model, the model is sent to all client devices C.
[0038] Step 2: The adversary client device C adv views the structure of the SincNet model and selects some neurons in the model as ghost neurons. All client devices preprocess the locally held private data and divide the data set according to the size of the sliding window and sliding distance to expand the data set.
[0039] Step 3: The client device uses the local private data set for pre-training, records the values of each sample on the ghost neurons, and plots a histogram of the recorded values.
[0040] Step 4: Selects the backdoor activation value of the ghost neurons as the trigger for the backdoor according to the distribution of the values on the ghost neurons.
[0041] Step 5: Starts training and determines whether it is an attack round. If the current is an attack round, then backdoor implantation is performed. The adversary modifies the values of the ghost neurons to the backdoor activation values and changes the labels to the labels specified by the adversary. If the current is not an attack round, then no backdoor implantation is performed. During training, the values of the ghost neurons are maintained as the values obtained by natural calculation of the samples, and the labels are maintained as the true labels.
[0042] Step 6: The benign client devices upload the clean model parameters, and the adversary client device uploads the model parameters implanted with the backdoor. The central server aggregates the models to obtain a federated learning voiceprint recognition model containing the backdoor.
[0043] More specific steps are as follows:
[0044] In the embodiment, Step 1 includes: The central server S constructs a federated learning framework as Figure 3 shown, and uses SincNet as the global model G for voiceprint recognition. The central server randomly selects 10 from the union of all 3000 client devices as the client device set C = {C 1 , C 2 , C 3 , ···, C adv , C N} for the current round. After initializing the global model, the model is sent to all client devices in C, where each round of C contains an adversary C adv .
[0045] In the embodiment, Step 2 includes: The adversary client device C adv views the structure of the global model G as Figure 4As shown in the figure, select some neurons in the model as ghost neurons: in the last hidden layer of the model, starting from the first neuron, select 1, 2, 5, 10, 15, 20, 30, and 50 consecutive neurons as ghost neurons; all clients preprocess the locally private data they hold, convert the audio content into WAV format, set the length of the audio to 16,000 frames, and according to the sliding window win size and the sliding distance win long size, split and augment the dataset; set win size = 200, win long = 10 The sliding window starts sliding from frame 0, saves the audio within the window every time it slides, and then slides backward by one win long ; finally obtain a new augmented training dataset D = {D 1 , D 2 , D 3 , ···, D n}, where n is the number of samples in the local dataset.
[0046] In the embodiment, step 3 includes: all clients use the split local dataset D for pre-training. For the i-th sample D i , the model will generate a value at the ghost neuron; the client records the value of each sample on each ghost neuron, generating a set of values, g is the index of the ghost neuron; draw a histogram of the recorded values for each ghost neuron according to V, and keep two decimal places when recording.
[0047] In the embodiment, step 4 includes: according to the value distribution shown in the histogram, find a value that exists in the histogram with the same proportion as the expected backdoor activation probability of the adversary. The expected triggering probability of the backdoor is 20%, so find the number that accounts for 20% in the histogram; in this embodiment, 0.5 accounts for 20% in the histogram, so select 0.5 and set it as the activation value of the ghost backdoor.
[0048] In the embodiment, step 5 includes: start training, determine that the frequency of backdoor implantation during training is N, calculate the relationship between the current round epoch and the backdoor implantation frequency N to determine whether the current is an attack round: if epoch is a multiple of N, then the current is an attack round, and perform backdoor implantation. During the process of backdoor implantation, the adversary will perform matrix operation I, as Figure 5 shown, set the value of the ghost neuron to the activation value of the ghost backdoor set in step 4, and at the same time modify the label of each sample to the specified label. D' is the sample after modifying the label, where W epoch is the parameter matrix of the model in the epoch-th round, with specific dimensions of r rows and d columns. is a masking matrix with r rows and d columns, is a backdoor activation matrix with r rows and d columns; then the model will use and D' to update according to II, where L is the cross-entropy loss function used by the model
[0049] used by the model, is to calculate the gradient of the model based on the loss value calculated by the loss function, and η is the learning rate; if the current round epoch is not a multiple of N, then in the current non-attack round, no backdoor implantation is performed, and normal training is carried out according to III. During training, the values of the ghost neurons are kept as the values naturally calculated by the samples, and the labels are kept as the true labels. The formula is as follows:
[0050]
[0051]
[0052]
[0053] In the embodiment, step 6 includes: the benign client uploads the clean model parameters, the adversary client uploads the model parameters implanted with the backdoor, and the central server aggregates the model through the PartFedAvg method IV to obtain the federated learning voiceprint recognition model containing the backdoor. The formula is as follows:
[0054]
[0055] In summary, the present invention generates a trigger during the model training process based on a method that does not rely on samples, and randomly triggers during the model usage process. The result of the trigger is that the model outputs a specified class. The random triggering of clean samples greatly improves the concealment of the backdoor, and the output of the adversary-specified class enables the model to maintain a high threat. Figure 6 is the backdoor activation success rate curve graph of the present invention.
[0056] The above is only a preferred embodiment of the present invention, and does not impose any form of limitation on the present invention. Any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A sample-independent federated learning voiceprint recognition backdoor attack method, used in the field of backdoor attacks on voiceprint recognition tasks under the federated learning computing paradigm in artificial intelligence, characterized by: The steps include: Step 1: The central server S builds a federated learning framework and uses SincNet as the global model for voiceprint recognition. After initializing the global model, the model is sent to all clients C. Step 2: Adversary Client C adv Check the SincNet model structure and select some neurons in the model as ghost neurons; All clients pre-process their local private data, convert the audio content into WAV format, set the length of the audio to 16,000 frames, and divide the local data set according to the size of the sliding window and sliding distance to expand it into the training data set D; Step 3: The client uses the segmented training data set D for pre-training, records the value of each sample on the ghost neuron, and draws a histogram of the recorded values; Step 4: Select the backdoor activation value of the ghost neuron as the trigger of the backdoor according to the distribution of the values on the ghost neuron; Step 5: Start training and determine whether it is an attack round. If it is an attack round, a backdoor is implanted. The adversary modifies the value of the ghost neuron to the backdoor activation value and modifies the label to the label specified by the adversary. If it is not an attack round, no backdoor is implanted. During training, the value of the ghost neuron is kept as the value naturally calculated by the sample and the label is kept as the true label. Step 6: The benign client uploads clean model parameters, and the adversary client uploads model parameters with a backdoor implanted. The central server aggregates the models to obtain a federated learning voiceprint recognition model containing the backdoor. The step 5 specifically starts training, determines that the frequency of backdoor implantation during training is N, and calculates the relationship between the current round epoch and the backdoor implantation frequency N to determine whether the current round is an attack round: if the epoch is a multiple of N, then the current round is an attack round, and the backdoor is implanted. During the backdoor implantation process, the adversary will set the value of the ghost neuron to the activation value of the ghost backdoor set in step 4 through matrix operation I, and at the same time modify the label of each sample in the segmented training data set D to the specified label, D′ is the sample after the modified label, where W epoch is the parameter matrix of the model in the epoch, with a specific size of r rows and d columns. is a mask matrix with r rows and d columns, is an r-row, d-column backgate activation matrix; the model will then use and D′ are updated according to II, where L is the cross entropy loss function used by the model, The gradient of the model is calculated based on the loss value calculated by the loss function, and η is the learning rate. If the current epoch is not a multiple of N, then the current non-attack round does not implant the backdoor, and normal training is performed according to III. During training, the value of the ghost neuron is kept as the value naturally calculated by the sample, and the label is kept as the real label. The formula is as follows:
2. The sample-independent federated learning voiceprint recognition backdoor attack method according to claim 1, characterized in that: Specifically, the central server S builds a federated learning framework and uses SincNet as the global model G for voiceprint recognition; the central server randomly selects 10 clients from the collection of all clients as the client collection C of the current round = {C1, C2, C3, ..., C adv , C N }; After initializing the global model, the model is sent to all clients in C, where each round of C contains an adversary C adv .
3. The sample-independent federated learning voiceprint recognition backdoor attack method according to claim 1, characterized in that: The step 2 is specifically as follows: adv Check the structure of the global model G and select some neurons in the model as ghost neurons: in the last hidden layer of the model, starting from the first neuron, select 1, 2, 5, 10, 15, 20, 30, and 50 consecutive neurons as ghost neurons; all clients pre-process the local private data they hold, convert the audio content into WAV format, set the length of the audio to 16,000 frames, and use the sliding window win size and sliding distance win long The size of the data set is divided and expanded; set win size =200, win long =10 The sliding window starts from frame 0, saves the audio in the window each time it slides, and then slides back one win long ; Finally, a training data set D = {D1, D2, D3, ..., D n }, n is the number of samples in the training data set after segmentation.
4. The sample-independent federated learning voiceprint recognition backdoor attack method according to claim 3, characterized in that: Specifically, step 3 includes that all clients use the training data set D={D1, D2, D3, ..., D n } for pre-training, for the i-th sample D i , the model generates a value at the ghost neuron; the client records the value of each sample at each ghost neuron, generating a set of values, g is the index of the ghost neuron; a histogram of the values recorded for each ghost neuron according to V is plotted, with two decimal places when recording.
5. The sample-independent federated learning voiceprint recognition backdoor attack method according to claim 1, characterized in that: Specifically, step 4 includes finding a value in the histogram whose existence ratio is the same as the backdoor activation probability expected by the adversary according to the distribution of the values displayed in the histogram, and setting the value as the activation value of the ghost backdoor.
6. The sample-independent federated learning voiceprint recognition backdoor attack method according to claim 1, characterized in that: Specifically, the step 6 includes: the benign client uploads clean model parameters, and the adversary client uploads model parameters with a backdoor implanted. The central server aggregates the model through the PartFedAvg method IV formula to obtain a federated learning voiceprint recognition model containing a backdoor. The formula is as follows:
Citation Information
Patent Citations
Backdoor attack methods for video analysis neural network models
CN111260059B
Voice backdoor attack method and computer readable storage medium
CN116259333A
Voice backdoor verification method and device based on room impulse response
CN116597811A
Physical light backdoor attack training method oriented to artificial intelligence safety
CN116664978A
Federal learning backdoor attack defense method based on DAGMM
CN113411329A