Federal learning backdoor attack analysis method and corresponding defense method
By adding regular terms that represent similarity in the loss function of federated learning, the federated learning framework lacks defense against partial poison injection attacks of feature extractors, and the simulation and defense of poison injection attacks are realized, and the defense capability is improved.
Patent Information
- Application Number
- CN202311830270.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2025-06-27
AI Technical Summary
The existing federated learning framework lacks an effective defense mechanism when facing partial poison injection attacks of the model's feature extractor, resulting in the inability to effectively defend under actual attacks.
By adding regular terms about the representation similarity in the loss function of federated learning, the input with backdoor features is close to the output of the normal input of the target tag in the feature extractor part, thereby improving the defense ability of poison injection attacks on the feature extractor part of the federated learning model.
The simulation and defense of partial poison injection attacks of federated learning model feature extractor is realized, which improves the probability of the normal client local model for input and output target labels with backdoor features, and enhances the defense capability.
Smart Images

Figure CN120218274A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of federated learning attacks and defenses, and in particular to an analysis method for backdoor attacks in federated learning and a corresponding defense method. Background Art
[0002] In the framework of federated learning, each participating party can exchange the intermediate results of model training instead of the local training data, so as to achieve the effect of protecting user privacy. However, this way of exchanging the intermediate results of model training also has its limitations. The most prominent one is that the model accuracy will significantly decrease in the case of data heterogeneity in the federated learning system, and even the training will not converge. The original federated learning framework ignores the problem of data heterogeneity, but the existence of certain data heterogeneity often conforms more to the actual application situation of federated learning.
[0003] In order to alleviate the problem of data heterogeneity in federated learning, relevant scholars have conducted many explorations. At present, there are many improved versions of federated learning methods that can adapt to heterogeneous data to a certain extent. Among them, a large number of federated learning methods are based on the same idea, that is, using the underlying part of the federated learning training model, and making the part of the model close to the output achieve a certain degree of personalization for local data. The most typical algorithms are FedPer and FedRep. They divide the model into two parts, namely the feature extractor part close to the bottom layer and the classifier part close to the output layer. In each iteration of federated learning, there is a step of aggregating the models locally trained by each client to generate a new global model. In this algorithm, FedPer and FedRep only aggregate the feature extractor part of the model, and each client model has a different classifier part. They are also called personalized federated learning frameworks, which are characterized by generating different personalized models for each client, rather than only training a global model like traditional federated learning.
[0004] The current federated learning of FedPer and FedRep algorithms lacks the simulation of the situation of poisoning only the feature extractor part of the model, and cannot defend against the attack of poisoning only the feature extractor part of the model in actual situations. Summary of the Invention
[0005] The purpose of the present invention is to provide an analysis method for backdoor attacks in federated learning and a corresponding defense method to improve the defense ability of federated learning against attacks on the feature extractor part of the model.
[0006] The purpose of the present invention can be achieved by the following technical solutions:
[0007] An analysis method for backdoor attacks in federated learning, the method includes:
[0008] In the client, according to the backdoor type, select some data from the original training data for processing to generate backdoor data, and then initialize the local model. The label of the backdoor data is the target label;
[0009] Initialize the feature extractor of the local model as the feature extractor of the global model sent by the server;
[0010] Select some data from the original training data as this batch of data, replace a certain proportion of the data in this batch of data with backdoor data, select a certain number of normal data with the target label, input them into the feature extractor of the local model, and output the average value r of the representations t , where the normal data is non-backdoor data;
[0011] Input the data replaced with backdoor data in this batch of data into the feature extractor of the local model, calculate the regularization term regarding the representation similarity based on the output of the feature extractor of the local model and the average value of the representations, and add the regularization term to the original training function of the local model to obtain the loss function;
[0012] Find the gradient of the loss function with respect to the feature extractor of the local model, and update the feature extractor of the local model based on the gradient;
[0013] Repeat S3 - S5 until reaching the local round number l times;
[0014] Aggregate the updated feature extractor after executing S6 and the feature extractor of the global model to obtain the feature extractor of the new global model, and distribute the feature extractor of the new global model to each local model;
[0015] Repeat S2 - S7 until reaching the number of iterations to obtain each trained local model;
[0016] Input the input with backdoor features into the local model, and analyze the defense ability of the local model based on the output of the local model.
[0017] Furthermore, the specific steps of S1 are as follows:
[0018] If the backdoor type is pixel representation type, select pictures from the original training data, copy the selected pictures, add specific pixel representations to the copied pictures, keep the pictures in the original training data unchanged, and set the label of the copied pictures as the target label to obtain backdoor data;
[0019] If the backdoor type is semantic, select pictures with relevant semantics from the original training data, copy the selected pictures, keep the pictures in the original training data unchanged, and set the labels of the copied pictures to the target labels to obtain the backdoor data.
[0020] Further, the input with backdoor features is specifically an image with specific pixel representations or an image with relevant semantics.
[0021] Further, the regularization term regarding the representation similarity is specifically:
[0022] μL r (q i )
[0023] Among them,
[0024]
[0025] D b represents the data in this batch that is replaced with backdoor data, represents an image of a sample belonging to D b , y t represents the label of a sample belonging to D o , and this label is the target label, represents the feature extractor of the i-th local model input and the output feature, r t represents the average value of the representation, q i represents the i-th local model, and μ represents the weight parameter of the regularization term and the original training function.
[0026] Further, the original training function is the cross-entropy loss function.
[0027] Further, the method for updating the feature extractor of the local model is the gradient descent method.
[0028] Further, the specific steps of S7 are:
[0029] Amplify the update of the feature extractor of the local model by k times, then combine it with the feature extractor of the global model, and upload the combined result to the server side. The server side aggregates the combined results of each local model to obtain the feature extractor of the new global model.
[0030] Further, the combined result is:
[0031]
[0032] Among them, is the updated feature extractor after executing S6, is the feature extractor of the global model, and k represents the magnification factor.
[0033] On the other hand, the present invention also proposes a defense method corresponding to the above analysis method of the backdoor attack in federated learning. The method includes:
[0034] Obtain the feature extractor of the new global model distributed to each local model in step S7 of the analysis method of the backdoor attack in federated learning;
[0035] Retrain the parameters of the feature extractor of the new global model received by the local model and the parameters of the classifier of the local model.
[0036] Further, the specific steps for retraining the parameters of the feature extractor of the new global model received by the local model and the parameters of the classifier of the local model are as follows:
[0037] Re-initialize the parameters of the last layer of the received feature extractor, fix the parameters of the received feature extractor and the classifier of the local model, and train the parameters of the last layer of the feature extractor based on the original training data;
[0038] Release the fixed parameters, and retrain the classifier and the parameters of the last layer of the feature extractor based on the original training data again.
[0039] Compared with the prior art, the present invention has the following beneficial effects:
[0040] (1) The present invention constructs a backdoor attack method. When performing local training, the original loss function will be modified to a certain extent. Specifically, a regularization term regarding the feature similarity is added, so that the output of the input with the backdoor feature and the normal input of the target label in the feature extractor part is close. This can greatly increase the probability that the local model of the normal client outputs the target label for the input with the backdoor feature, and further simulate the attack on only poisoning the feature extractor part of the model in federated learning, and obtain the defense ability of the attack on only poisoning the feature extractor part of the model in federated learning.
[0041] (2) The present invention also proposes a defense method. For the attack on only poisoning the feature extractor part of the model, after obtaining the feature extractor distributed by the server, retrain the parameters of the feature extractor and the parameters of the classifier of the local model to ensure that the positive global knowledge can be retained as much as possible while reducing the impact of poisoning. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is the flowchart of the present invention;
[0043] Figure 2 is the schematic diagram of the process of federated learning of the present invention. Detailed implementation manners
[0044] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. These embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given. However, the protection scope of the present invention is not limited to the following embodiments.
[0045] Through related experiments, it is found that traditional backdoor attack methods still have a certain effect under federated learning frameworks based on shared representation extractors such as FedPer and FedRep, but the attack success rate of backdoor attacks will be significantly reduced. This is because compared with traditional federated learning frameworks, the local models in these frameworks have personalized classifier parts. Based on the observation that the representation extractor part of the model is still shared, the present invention provides an analysis method for backdoor attacks in federated learning, which can greatly improve the backdoor attack success rate against federated learning algorithms based on shared representation extractors, and analyze the defense performance of the model against backdoor attacks in federated learning algorithms based on shared representation extractors.
[0046] The flowchart of the analysis method for backdoor attacks in federated learning is as Figure 1 shown, and the steps of the method include:
[0047] S1. In the client, according to the type of backdoor, select a part of the data from the original training data for processing to generate backdoor data, and then initialize the local model. The label of the backdoor data is the target label;
[0048] S2. Initialize the feature extractor of the local model as the feature extractor of the global model sent by the server;
[0049] S3. Select a part of the data from the original training data as this batch of data, replace a certain proportion of the data in this batch of data with backdoor data, select a certain number of normal data with the target label, input them into the feature extractor of the local model, and output the average value r^t of the representations. The normal data is non-backdoor data;
[0050] S4. Input the data replaced with backdoor data in this batch of data into the feature extractor of the local model, calculate the regularization term regarding the representation similarity based on the output of the feature extractor of the local model and the average value of the representations, and add the regularization term to the original training function of the local model to obtain the loss function;
[0051] S5. Calculate the gradient of the loss function with respect to the feature extractor of the local model, and update the feature extractor of the local model based on the gradient;
[0052] S6. Repeat S3 - S5 until the local round number reaches l times;
[0053] S7. Aggregate using the updated feature extractor after executing S6 and the feature extractor of the global model to obtain the feature extractor of the new global model, and distribute the feature extractor of the new global model to each local model;
[0054] S8. Repeat S2 - S7 until the number of iterations is reached to obtain each trained local model;
[0055] S9. Input the input with backdoor features into the local model, and analyze the defense ability of the local model based on the output of the local model.
[0056] Among them, S1 to S8 are all for constructing a federated learning backdoor attack based on representation similarity, that is, an attack on poisoning the feature extractor part of the model. To implement this attack, the following steps are required:
[0057] 1. Selection of backdoor type
[0058] The system needs to determine what type of backdoor to use. Backdoors can be mainly divided into two types. Considering the case where the input of the model is an image, one type of backdoor shows that when the image has a specific pixel representation (such as a black cross in the lower left corner), a specific label is output; another type of backdoor is called a semantic backdoor, which shows that when the image has a specific semantics (such as there is a dalmatian in the image, or there are stripes on the car in the image), a specific label is output.
[0059] 2. Generation of backdoor data
[0060] According to the different backdoor types, the generation methods of backdoor data are also different. Specifically, taking the model input as an image as an example, for the pixel representation type of backdoor, only need to add the pixel representation to the normal image and then set its label to the target label; for the semantic backdoor, it is necessary to collect images with relevant semantics in advance, and certain data augmentation means can be applied, and set its label to the target label.
[0061] 3. Federated learning system architecture based on a shared feature extractor
[0062] Such as Figure 2As shown, in a federated learning system based on a shared feature extractor, during each round of global iteration, the server sends the shared global feature extractor to each client. After that, each client uses the global feature extractor sent by the server to update the feature extractor part of the client's local model, then uses local data to train the entire local model, and then sends the updated feature extractor part of the local model to the server, and the server aggregates it to obtain a new global feature extractor. In this process, an attacker will use backdoor data for model training. Different from normal clients, the attacker only updates the feature extractor part of the local model. Then, the model update is amplified to a certain extent so that the attacker's malicious update will not be overly diluted during global aggregation.
[0063] In the backdoor attack of federated learning based on feature similarity, the original loss function will be modified to some extent during local training, specifically by adding a regularization term regarding feature similarity. In each round of local training, the attacker will select a certain number of normal data with target labels, and calculate the average value of the feature vectors output by the corresponding feature extractor. Then for each backdoor data, the L2 norm of the difference between its feature vector and the average feature vector of the target label is added to the loss function. The motivation is that although the attacker cannot directly affect the final output of other clients' models, the attacker can directly affect the output of the feature extractor part of other clients' models. By making the output of the input with backdoor features close to the output of the normal input with the target label in the feature extractor part, the probability that the local model of normal clients outputs the target label for the input with backdoor features can be greatly increased. According to relevant research, when the upper layer of the model is poisoned during the training stage, the impact on it is generally greater than that on the lower layer of the model. However, for most federated learning system architectures based on shared feature extractors, the upper layer of the model, that is, the classifier part, does not participate in aggregation and cannot directly affect other clients. Therefore, the backdoor attack method in the present invention only updates the feature extractor part of the local model, which can improve the poisoning effect on the feature extractor to a certain extent.
[0064] In S1, according to the backdoor type, the original training data is processed to generate backdoor data. The specific steps of S1 are as follows:
[0065] If the backdoor type is pixel representation type, select pictures from the original training data, copy the selected pictures, add specific pixel representations to the copied pictures, keep the pictures in the original training data unchanged, and set the label of the copied pictures to the target label to obtain backdoor data;
[0066] If the backdoor type is semantic, select pictures with relevant semantics from the original training data, copy the selected pictures, keep the pictures in the original training data unchanged, set the labels of the copied pictures as the target labels, and obtain the backdoor data. In S1, initialize the local model h i is the classifier part is the feature extractor part
[0067] In S2, at the beginning of each global iteration, obtain the global feature extractor from the federated learning server Update
[0068] The steps from S3 to S5 are as follows
[0069] Traverse the local data for training. Each time, select a certain number of data, and replace a part of the data with a proportion of p in this batch of data with backdoor data
[0070] Select a certain number of normal data with target labels, and calculate the average value r of the output representations of their corresponding feature extractors t .
[0071] Add a regularization term regarding the representation similarity to the original loss function. μL r (q i ), where
[0072]
[0073] D b represents the backdoor data in this batch. μ is used to control the proportion of the regularization term and the original loss function. Here, the original loss function varies according to the model task and has no special association with this algorithm. For example, for an image classification task, the original loss function can adopt cross-entropy
[0074] Obtain the gradient of the loss function with respect to the feature extractor part of the model for updating the feature extractor part of the model. Here, various parameter update methods are applicable, and generally, gradient descent can be used
[0075] In S6, repeat S3 - S5 until the local round number reaches l times
[0076] In S7, upload to the server for aggregation, that is, upload after magnifying the update of the local feature extractor by k times
[0077] In S8, repeat S2 - S7 until the iteration number is reached to obtain each local model that has completed training
[0078] In S9, the input with the backdoor feature is an image with specific pixel representations or an image with relevant semantics.
[0079] Based on the above analysis method, the present invention also proposes a defense method, which uses the local dataset to retrain and adjust the model after the end of the joint training stage of federated learning. For the federated learning backdoor attack algorithm based on representation similarity, in order to retain the effective knowledge of the global model, avoid overfitting, and efficiently eliminate the backdoor based on representation similarity.
[0080] The defense method is as follows:
[0081] 1. After the end of the joint training stage of federated learning, that is, after the last step of the backdoor attack described in the previous part of this document. First, re-initialize the parameters of the layer before the classifier of the local model.
[0082] 2. Fix the parameters of other layers, and use the local data to train and adjust the parameters of the layer before the classifier.
[0083] 3. Release the fixation of the local classifier, and use the local data to train and adjust the parameters of the classifier and the layer before the classifier.
[0084] The above method only retrains a small part of the parameters of the model (the classifier and the parameters of the layer before the classifier), and the training efficiency is relatively high. The parameters closer to the output layer are more significantly affected by the backdoor attack, which is also the reason why only the extractor part of the model is backdoor-injected in the attack method proposed in this patent. When only the extractor part is backdoor-injected, it is the last layer of the extractor, that is, the layer before the classifier, that is most deeply affected by the backdoor injection. Therefore, the layer before the classifier can be retrained, and the parameters of other layers are fixed during the retraining to ensure that the positive global knowledge can be retained as much as possible. After the training of the layer before the classifier is completed, the parameters of the classifier and the layer before the classifier also need to be fine-tuned to ensure the performance of the model on normal tasks.
[0085] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative labor. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should fall within the protection scope determined by the claims.
Claims
1. An analysis method for backdoor attacks in federated learning, characterized in that, The method includes: S1. In the client, according to the backdoor type, select part of the data from the original training data for processing to generate backdoor data, and then initialize the local model. The label of the backdoor data is the target label; S2. Initialize the feature extractor of the local model as the feature extractor of the global model sent by the server; S3. Select a portion of the data from the original training data as the current batch of data, replace a certain proportion of the data in the current batch with backdoor data, select a certain number of normal data with target labels, input them into the feature extractor of the local model, and output the average value r of the representations t , where the normal data is non-backdoor data; S4. Input the data replaced with backdoor data in this batch into the feature extractor of the local model, calculate the regularization term regarding the representation similarity based on the output of the feature extractor of the local model and the average value of the representations, and add the regularization term to the original training function of the local model to obtain the loss function; S5. Obtain the gradient of the loss function with respect to the feature extractor of the local model, and update the feature extractor of the local model based on the gradient; S6. Repeat S3 - S5 until the local round number reaches l times; S7. Aggregate the updated feature extractor and the feature extractor of the global model to obtain the feature extractor of the new global model, and distribute the feature extractor of the new global model to each local model; S8. Repeat S2 - S7 until the number of iterations is reached to obtain each trained local model; S9. Input the input with backdoor features into the local model, and analyze the defense ability of the local model based on the output of the local model.
2. The analysis method of a backdoor attack in federated learning according to claim 1, characterized in that, The specific steps of S1 are: If the backdoor type is pixel representation type, select pictures from the original training data, copy the selected pictures, add specific pixel representations to the copied pictures, keep the pictures in the original training data unchanged, and set the label of the copied pictures as the target label to obtain backdoor data; If the backdoor type is semantic type, select pictures with relevant semantics from the original training data, copy the selected pictures, keep the pictures in the original training data unchanged, and set the label of the copied pictures as the target label to obtain backdoor data.
3. The analysis method of a backdoor attack in federated learning according to claim 2, wherein The input with backdoor features is specifically an image with specific pixel representations or an image with relevant semantics.
4. The analysis method of a backdoor attack in federated learning according to claim 1, wherein The regularization term regarding the representation similarity is specifically: μL r (q i ) where D b Indicates the data replaced with backdoor data in this batch of data Indicates belonging to D b The image of a sample belonging to D, y t Indicates belonging to D b The label of a sample belonging to D, and this label is the target label Indicates the feature extractor of the i-th local model Input The output feature after input, r t Indicates the average value of the representation, q i Indicates the i-th local model, and μ represents the proportion parameter of the regularization term and the original training function 5. The analysis method of a backdoor attack in federated learning according to claim 4, wherein The original training function is the cross - entropy loss function.
6. The analysis method of a backdoor attack in federated learning according to claim 1, characterized in that The method for updating the feature extractor of the local model is the gradient descent method.
7. The analysis method of a backdoor attack in federated learning according to claim 1, characterized in that The specific steps of S7 are: Amplify the update of the feature extractor of the local model by k times, then combine it with the feature extractor of the global model, and upload the combined result to the server. The server aggregates the combined results of each local model to obtain the feature extractor of the new global model.
8. The analysis method of a backdoor attack in federated learning according to claim 7, characterized in that, The combined result is: Among them, is the updated feature extractor after executing S6, is the feature extractor of the global model, and k represents the magnification factor.
9. A defense method corresponding to the analysis method of the federated learning backdoor attack described in any one of claims 1 to 8, characterized in that, The method includes: Obtain the feature extractor of the new global model distributed to each local model in S7 of the analysis method for federated learning backdoor attacks; Retrain the parameters of the feature extractor of the new global model received by the local model and the parameters of the classifier of the local model.
10. The defense method according to claim 9, characterized in that, The specific steps for retraining the parameters of the feature extractor of the new global model received by the local model and the parameters of the classifier of the local model are: Re-initialize the parameters of the last layer of the received feature extractor, fix the parameters of the received feature extractor and the classifier of the local model, and train the parameters of the last layer of the feature extractor based on the original training data; Unfix the fixed parameters and retrain the parameters of the last layer of the classifier and the feature extractor based on the original training data.