Self-adaptive simulation backdoor attack method in federated learning scene
Through the adaptive simulation of backdoor attack method, the model's evasion detection ability is enhanced by using adaptive modules and large-scale simulations, and the backdoor impact is amplified by stimulating the model, the problems of accuracy and stability of the existing technology in a variety of defense scenarios are solved, achieving efficient and stable backdoor attack effects.
Patent Information
- Application Number
- CN202510239108.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-20
AI Technical Summary
The existing backdoor attack technology is difficult to maintain high accuracy in a variety of defense scenarios and is easily detected, resulting in unstable model and limited backdoor effects.
An adaptive simulation backdoor attack method is proposed. By invading some federated learning participants in advance, controlling the local model training process, combining adaptive modules and large-scale simulations, the model's ability to evade anomaly detection, and amplifying the impact of backdoors in the global model by stimulating the model.
Maintain high attack stability and success rate under multiple defense mechanisms, improve the influence of backdoors in the global model, and enhance generalization and attack efficiency.
Smart Images

Figure CN120181184A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of federated learning and backdoor attacks, and particularly relates to a backdoor attack method based on federated learning. Background Art
[0002] In the vast field of federated learning, although its distributed training mode provides strong protection for user data privacy, backdoor attacks still pose a security threat that cannot be ignored. Backdoor attacks manipulate the local training process of user systems, submit malicious models to the central server, and then implant backdoors in the global model, which may lead to a decrease in the accuracy of the main task, instability in the model training process, and serious security vulnerabilities.
[0003] To address this challenge, academia and industry have developed various defense strategies aimed at detecting and preventing backdoor attacks. However, existing defense methods often have limitations and are difficult to comprehensively resist backdoor attacks. Therefore, researching new backdoor attack technologies to evaluate and enhance the defense capabilities of federated learning systems has important theoretical and practical significance.
[0004] Among existing backdoor attack technologies, most methods are difficult to maintain high accuracy in multiple defense scenarios. To promote the progress of defense technologies, this paper proposes a new backdoor attack technology - adaptive simulation backdoor attack. This technology aims to enhance the model's ability to evade anomaly detection through more complex methods, combining adaptive modules and large-scale simulations, and maintain high accuracy in multiple defense scenarios.
[0005] Currently, one of the closest prior arts to the present invention is the ModelReplace attack. ModelReplace is a typical backdoor attack method that introduces a backdoor into the global model by replacing local model parameters. However, the ModelReplace attack shows instability under various defense mechanisms, such as Flame, FLDetector, etc., which can detect abnormal changes in model parameters and thus prevent the success of backdoor attacks.
[0006] Analysis of the advantages and disadvantages of ModelReplace:
[0007] Advantages of the ModelReplace attack:
[0008] Simple to implement: The ModelReplace attack introduces a backdoor by directly replacing model parameters, and the operation is relatively simple.
[0009] Disadvantages of the ModelReplace attack:
[0010] 1. Instability: Under various defense mechanisms, the ModelReplace attack is easily detected, making it difficult for the model to converge.
[0011] 2. Inefficiency: Due to the existence of defense mechanisms, the backdoor effect of the ModelReplace attack is often limited. Summary of the Invention
[0012] The present invention provides an adaptive simulated backdoor attack method in the context of federated learning, which not only considers how to maintain the stability of the attack under various defense mechanisms, but also amplifies the influence of the backdoor in the global model by introducing a stimulus model, thereby improving the success rate of the attack.
[0013] An adaptive simulated backdoor attack method in the context of federated learning includes the following steps:
[0014] Step S1: Infiltrate some federated learning participants in advance and control the local model training process of these participants for implementing the backdoor attack.
[0015] Step S2: Prepare benign training data for training normal local models; prepare backdoor training data for implementing backdoor training among the infiltrated participants.
[0016] Step S3: Among the infiltrated participants, use the backdoor training data to train the local model to embed the backdoor function.
[0017] Step S4: By calculating the Euclidean distance between the benign model and the malicious model and combining with the stimulus factor, determine the model scaling injection intensity and injection direction to formulate a backdoor injection strategy.
[0018] Step S5: According to the attacker's understanding of the central server information, i.e., white-box or black-box, implement corresponding defense simulations.
[0019] Step S6: By monitoring the attack effect and model performance metrics, identify inefficient or unstable backdoor injection areas and correct and trim these areas.
[0020] Step S7: Submit the malicious model trained with the backdoor to the central server for aggregation. The aggregated global model will be used as the starting point for the next round of iteration, and repeat the above steps S3 - S7 to continuously optimize the backdoor attack effect.
[0021] The present invention aims to deeply explore and reveal the potential weaknesses of defense strategies in the current federated learning environment, skillfully leveraging the inherent characteristic that the central server of the federated learning framework is invisible to the data of each participating party. Through infiltration means, this technology maliciously manipulates the participating clients during training, enabling them to carry out a specifically designed training process without raising suspicion. Subsequently, these manipulated participants will submit carefully constructed malicious models, aiming to secretly influence the decision-making logic of the global model while maintaining the appearance of normal system operation, thereby successfully implanting a backdoor function. Figure 1 It details the overall architecture of this complex attack strategy. The adaptive simulated backdoor attack method we proposed includes three core modules: the adaptive module, the large-scale simulation module, and the pruning module. These modules integrate two strategies of data poisoning and model poisoning to achieve an efficient attack on the federated learning system.
[0022] The present invention discloses an adaptive simulated backdoor attack method in the context of federated learning. This technology can maintain high attack performance while possessing stronger generalization and stability in the federated learning scenario.
[0023] Compared with the prior art, the advantages of the present invention are as follows:
[0024] 1. High stability: Through the adaptive algorithm and large-scale simulation, the present invention can maintain the stability of the attack under various defense mechanisms.
[0025] 2. High efficiency: By introducing a stimulus model, the present invention can amplify the influence of the backdoor in the global model and improve the success rate of the attack.
[0026] 3. Generalization: The present invention is not only applicable to complex data sets such as CIFAR-10, but also has good generalization and can perform equally well on other data sets. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] To more clearly illustrate the technical solution of the present invention, the accompanying drawings required for the implementation will be briefly introduced below.
[0028] Figure 1 It is a schematic diagram of the overall architecture of the backdoor attack based on federated learning of the present invention;
[0029] Figure 2 It is a schematic diagram of the comparison between the backdoor attack of the present invention and the traditional backdoor attack in the FedAvg defense-free scenario;
[0030] Figure 3 It is a schematic diagram of the comparison between the backdoor attack of the present invention and the traditional backdoor attack in the Deepsight defense scenario;
[0031] Figure 4It is a comparison schematic diagram of the backdoor attack of the present invention and the traditional backdoor attack in the Flame defense scenario;
[0032] Figure 5 It is a comparison schematic diagram of the backdoor attack of the present invention and the traditional backdoor attack in the Fldetector defense scenario;
[0033] Figure 6 It is a comparison schematic diagram of the backdoor attack of the present invention and the traditional backdoor attack in the Foolsgold defense scenario;
[0034] Figure 7 It is a comparison schematic diagram of the backdoor attack of the present invention and the traditional backdoor attack in the Rflbat defense scenario. Detailed implementation manners
[0035] Next, the technical solutions of the present invention will be further elaborated in detail in conjunction with the attached Figure 1 In a preferred embodiment of the present invention, an adaptive simulated backdoor attack method in the federated learning scenario is provided.
[0036] As Figure 1 shown in the overall architecture schematic diagram, an adaptive simulated backdoor attack method in the federated learning scenario includes the following steps:
[0037] Step S1: Invade some federated learning participants in advance and control the model training process of these participants to implement the backdoor attack. This entire process is called malicious training (thread).
[0038] Step S2: Prepare benign training data for training a normal model; prepare backdoor training data for implementing backdoor training in the invaded participants. Two widely used datasets - MNIST and CIFAR-10 are adopted. The scales of these datasets gradually increase, representing low and high data volumes respectively. For MNIST, a simple 2-conv-2 fully connected model is adopted, while for CIFAR-10, the ResNet18 model is selected. To simulate the backdoor attacks that may be encountered in the real environment, we set 20% of the participants to perform malicious behaviors, and the remaining 80% of the participants conduct normal benign training.
[0039] Step S3: Among the invaded participants, use the backdoor training data to perform backdoor training on the initial model to obtain a malicious model with backdoor functionality. Among the 20 participants in MNIST, 4 were set as attackers, and among the 100 participants in CIFAR-10, 20 were attackers each. To be closer to the actual situation, in each training round, we randomly selected 20% of the participants from all participants as attackers through mini-batch random sampling. When performing the backdoor attack, we selected the 8th target class as the target dataset for the backdoor attack. During the training process, to improve the experimental efficiency, we performed 200 rounds of pre-training, during which all participants performed benign local model training. Based on the models obtained through pre-training on each dataset, we further conducted 100 rounds of backdoor attack experiments.
[0040] Step S4: By calculating the Euclidean distance between the benign model and the malicious model and combining the stimulation factor, determine the model scaling scale (i.e., the injection intensity) and the injection direction to formulate a backdoor injection strategy.
[0041] Step S41: The definition of calculating the Euclidean distance d(x,y) is as follows, where x=(x1,x2,…,x n ) and y=(y1,y2,…,y n ) represent the feature vectors of the benign model and the malicious model respectively:
[0042]
[0043] Here, x i and y i are the i-th elements of the feature vectors x and y.
[0044] Step S42: The backdoor injection strategy can be described by the model incremental gradient, where α is the stimulation factor and v is the unit vector of the injection direction. Assume x is the feature vector of the benign model, then the feature vector x backdoor after backdoor injection can be written as:
[0045] x backdoor =x + αv
[0046] Here, x backdoor represents the feature vector after backdoor injection.
[0047] Even if the attack is not successful in this way, it can increase the noise of the malicious model to a certain extent, affect the detection ability of the central server for local models, and improve the attack success rate of subsequent malicious models. However, once the attack is successful under the action of the stimulation factor, it can significantly improve the final backdoor accuracy of the global model. The adaptive module can cope with different federated learning environments and attack targets.
[0048] Step S5. Implement corresponding defense simulations according to the attacker's knowledge of the central server information (white box or black box).
[0049] Step S51. Black box scenario: The attacker knows nothing about the internal information of the central server. We simulate various aggregation algorithms to obtain the metrics under each defense method, and then obtain the final metrics through normalization to achieve the backdoor injection effect.
[0050] Assume that I is the normalized metric vector. The training process of the backdoor model can be expressed as:
[0051] I = Normalize(f(m))
[0052] where f is the simulation function of various aggregation algorithms and m is the backdoor attack model.
[0053] After each round of training, adaptively adjust the weight λ according to the results of the previous round of training i , so that the weights close to the central server defense mechanism gradually increase to approach the effect of the white box model.
[0054]
[0055] Here, η is the learning rate and ΔI is the metric change between two adjacent rounds of training.
[0056] Step S52. White box scenario: The attacker understands the aggregation method and defense mechanism of the central server, so as to simulate the synthesis and aggregation process of adversarial samples to optimize the backdoor injection effect to the greatest extent.
[0057] Step S6. Identify inefficient or unstable backdoor injection areas by monitoring the attack effect and model performance metrics. Correct and trim these areas.
[0058] Step S7. Submit the local models trained by the backdoor attack to the central server for aggregation (Average). After this step, the aggregated model is called the global model. Subsequently, this global model will be downloaded and distributed to each participant as the starting model for their next round of iterative training, repeating the above steps (S3 - S7) to continuously optimize the backdoor attack effect. After 100 attack iterations, the training ends. To comprehensively evaluate the effects of different defense strategies, we implemented six advanced defense algorithms in the experiment: FedAvg (no defense), FLAME, RFLBAT, Deepsight, Foolsgold, and FL - Detector. Figures 2 - 7Shows the specific experimental comparison situation, where the x-axis represents the attack epoch, the y-axis represents the backdoor accuracy in the model, modelreplace represents the traditional backdoor attack method, and Adaptive SimulationBackdoorAttack represents the attack method of the present invention. It can be clearly seen that the method of the present invention is superior to the traditional backdoor attack method in the scenario with defense, and its poor performance in the scenario without defense is also within our expectation because we have taken many measures to improve the ability of the model to evade anomaly detection.
Claims
1. An adaptive simulated backdoor attack method in a federated learning scenario, characterized in that The steps include: Step S1: Invade some federated learning participants in advance and control the local model training process of these participants to implement backdoor attacks; Step S2: preparing benign training data for training a normal local model; preparing backdoor training data for implementing backdoor training in intruding participants; Step S3: In the hacked participant, the local model is trained using the backdoor training data to embed the backdoor function; Step S4, by calculating the Euclidean distance between the benign model and the malicious model, combined with the stimulus factor, the model scaling injection strength and injection direction are determined to formulate a backdoor injection strategy; Step S5: Implement corresponding defense simulation according to the attacker's knowledge of the central server information, i.e., white box or black box; Step S6: Identify inefficient or unstable backdoor injection areas by monitoring attack effects and model performance indicators, and correct and trim these areas; Step S7: Submit the malicious model trained with the backdoor to the central server for aggregation. The aggregated global model will be used as the starting point for the next round of iteration. Repeat the above steps S3-S7 to continuously optimize the backdoor attack effect.
2. According to claim 1, the adaptive simulated backdoor attack method in a federated learning scenario is characterized in that The specific process of the above step S3 is: Step S31, using most of the data sets in cifar10 and mnist as benign training data sets, and performing benign local training on the model; Step S32: Use a small portion of the data set as a backdoor training data set, tamper with the data set to implement a specified backdoor, and perform multiple rounds of iterations on the backdoor data set training.
3. According to claim 2, the adaptive simulated backdoor attack method in a federated learning scenario is characterized in that The specific process of the above step S4 is: Step S41, calculate the Euclidean distance d(x,y) as follows, where x = (x1, x2, ..., x n ) and y=(y1,y2,…,y n ) represent the feature vectors of the benign model and the malicious model respectively: where x i and i is the i-th element of the eigenvectors x and y; Step S42: The backdoor injection strategy is described by the model incremental gradient. The feature vector x after backdoor injection backdoor Written as: x backdoor =x+αv where x backdoor represents the feature vector after backdoor injection, α is the stimulus factor, and v is the unit vector of the injection direction; assuming that x is the feature vector of the benign model.
4. According to claim 3, the adaptive simulated backdoor attack method in a federated learning scenario is characterized in that The specific process of the above step S5 is: Step S51, black box scenario: the attacker knows nothing about the internal information of the central server. By simulating multiple aggregation algorithms, the attacker obtains the indicators under each defense method, and then obtains the final indicator by normalization to achieve the backdoor injection effect; Assuming I is the normalized indicator vector, the training process of the backdoor model is expressed as: I = Normalize(f(m)) Where f is the simulation function of multiple aggregation algorithms, and m is the backdoor attack model; After each round of training, the weight λ is adaptively adjusted according to the results of the previous round of training. i , so that the weight of the defense mechanism close to the central server gradually increases to approach the effect of the white box model; Where η is the learning rate, ΔI is the change in the index between two consecutive rounds of training; Step S52, white box scenario: the attacker understands the aggregation method and defense mechanism of the central server, simulates the synthesis and aggregation process of adversarial samples, and optimizes the backdoor injection effect.