Federal learning backdoor attack defense method, system and device, medium and program product thereof

By using cascaded detection of indicators such as gradient cosine similarity, gradient magnitude, and gradient direction angle in a federated learning system, backdoor attacks can be identified and defended against. This solves the defense problem in scenarios involving adaptive adversaries and non-independent identically distributed data, and achieves a highly efficient backdoor attack defense effect.

CN120856444APending Publication Date: 2025-10-28XIDIAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511170025.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-20
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing backdoor attack defense schemes in federated learning systems fail to effectively identify the ability of adaptive adversaries and do not fully consider the characteristics of non-independent and identically distributed data, resulting in a significant reduction in defense effectiveness.

Method used

Five interrelated and non-simultaneously optimizable metrics are used to detect backdoor attacks in a cascaded manner, including gradient cosine similarity, gradient magnitude, gradient direction angle, upper bound of local model gradient update prediction, and lower bound of local model gradient update prediction. Statistical testing and analysis are performed through an aggregation server to identify adversary characteristics, and the degree of user malice is measured using a method without a fixed threshold.

Benefits of technology

In both independent and non-independent identically distributed data scenarios, it effectively defends against both ordinary and adaptive adversaries, maintains high accuracy in the main task, reduces backdoor task accuracy to near 0%, and is unaffected by the number of adversaries or data distribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120856444A_ABST
    Figure CN120856444A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of network space security, and discloses a federated learning backdoor attack defense method, system and device, a medium and a program product thereof, the method comprises the following steps: generating an initialized global model through an aggregation server and issuing the initialized global model to a user, and training by the user by using a local data set and the latest global model issued by the aggregation server, generating an updated local model, uploading the updated local model to an aggregation server, then performing anomaly detection on local model gradient information uploaded by a user by the aggregation server, aggregating model gradients of normal users according to a predefined aggregation rule to obtain an encrypted global model, and then issuing the global model to the users in the system; the system, the equipment and the medium are used for implementing the method. A program product comprising a computer program of the method; according to the method, the negative influence of backdoor attack on the federated learning system is avoided, and the robustness of the system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of cyberspace security technology, specifically relating to a method, system, device, medium, and program product for defending against federated learning backdoor attacks. Background Technology

[0002] With the rapid development of machine learning and deep learning technologies, artificial intelligence (AI) has been widely applied in various fields such as healthcare, financial risk control, autonomous driving, and intelligent manufacturing. The widespread adoption of these technologies has transformed traditional business models and workflows. However, the extensive application of AI has also brought about some problems, particularly regarding data privacy and security. In traditional machine learning, data needs to be centrally processed on servers, which may pose a risk of privacy breaches and thus data security issues. Simultaneously, data is distributed across different institutions and enterprises, highlighting the growing problem of data silos. Due to privacy concerns and other reasons, this data cannot be effectively shared and integrated, leading to wasted data resources and decreased model performance. Therefore, how to break down data silos while ensuring data privacy and security, and promote cross-institutional and cross-domain data sharing and collaboration, has become a current research hotspot in the field of artificial intelligence.

[0003] Federated learning, first proposed by Google in 2016, was initially used to improve its automatic input completion keyboard system. In a federated learning system, data remains locally on the user's machine. Users train their models using local data and send model updates (not the data itself) to an aggregation server. The aggregation server receives these updates from each user, aggregates the local models to generate a global model, and returns it to the individual users. In this process, the user's original data never leaves their local machine, thus avoiding the risk of data leakage and effectively protecting user privacy. Furthermore, federated learning enables cross-institutional data sharing and collaboration, successfully overcoming the data silo problem. Compared to traditional centralized machine learning, federated learning has significant advantages in privacy protection and data collaboration. However, during federated learning training, because the user's local data and training process are invisible to other users and the aggregation server, it is difficult to effectively verify the identity of each user, making it vulnerable to backdoor attacks. Backdoor attacks are typically covert and complex; the adversary's goal is to cause the model to behave abnormally under specific conditions without affecting the overall performance of the global model. For example, by adding backdoor triggers to the training samples, the global model can output a predefined target prediction when it encounters an input containing the trigger. Such attacks can maintain high accuracy in both the main task and the backdoor task, while maximizing attack effectiveness and stealth through precise computation and policy tuning. Therefore, in federated learning systems, effective methods are needed to detect and prevent backdoor attack adversaries to ensure the reliability and security of the federated learning system.

[0004] Most existing backdoor attack defense schemes do not fully consider the capabilities of adaptive adversaries (i.e., the ability to adaptively attack multiple defense metrics). For example, the FoolsGold scheme weights the contribution of local models by analyzing the cross-cosine distance of the last layer model update. However, because it only focuses on the last layer, it is easily bypassed by adaptive adversaries by fixing the parameters of that layer. In scenarios where parameters are fixed and PDR is increased, the backdoor accuracy can reach 63.54% (Clement Fung, Chris JM Yoon, and Ivan Beschastnikh. 2020. The Limitations of Federated Learning in SybilSettings. RAID (2020).). The M-Krum scheme, based on the Euclidean distance between local models, allows attackers to make the Krum score of the poisoned model closer to that of a benign model through constraint methods, thereby bypassing the defense. In relevant scenarios, the backdoor accuracy reached 89.90% and 95.80%, respectively (Peva Blanchard, El Mahdi El Mhamdi, Rachid Guerraoui, and Julien Stainer. 2017. Machine Learning with Adversaries: Byzantine Tolerant Gradient Descent. NIPS (2017). However, these solutions also overlook the fact that user local data is non-independent and identically distributed (Non-IID) in real-world scenarios. Considering different adversary capabilities and data distribution scenarios, the effectiveness of many existing backdoor attack defense solutions is significantly reduced. Summary of the Invention

[0005] To overcome the shortcomings of the prior art, the present invention aims to provide a federated learning backdoor attack defense method, system, device, medium, and program product. This method first detects backdoor attacks through a cascade of five interrelated and non-simultaneously optimizable metrics, including gradient cosine similarity, gradient magnitude, gradient direction angle, upper bound of local model gradient update prediction, and lower bound of local model gradient update prediction. Metrics for each layer of the model are calculated from the local model gradient and the global model gradient. Then, based on the assumption of a majority of normal users, the aggregation server analyzes the values ​​of each metric through statistical testing to identify potential adversary characteristics. Finally, the aggregation server marks the adversaries and performs further cluster analysis to filter out a set of normal users for whom no anomalies are detected in any of the five metrics, thereby avoiding the negative impact of malicious models generated by backdoor attack adversaries on the global model. Simultaneously, the use of a non-fixed threshold method to measure the degree of user malice further improves the robustness of the system.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0007] A method for defending against federated learning backdoor attacks includes the following steps:

[0008] Step 1: After the aggregation server initializes the global model and distributes it to all users within the federated learning system, users train their local models and upload them to the aggregation server.

[0009] Step 2: After receiving the local models, the aggregation server calculates the metric values ​​of each local model in terms of gradient cosine similarity, gradient magnitude, gradient direction angle, upper bound of local model gradient update prediction, and lower bound of local model gradient update prediction. It also calculates the median of each metric value as a reference threshold for normal users. If any one of the four metrics of a user's local model (gradient cosine similarity, gradient magnitude, gradient direction angle, or upper bound of local model gradient update prediction) is greater than the reference threshold, or if the metric value of the lower bound of local model gradient update prediction is less than the reference threshold, the user is identified as an adversary and will not participate in global model aggregation.

[0010] Step 3: After the aggregation server identifies and marks adversaries, it uses the FedAvg algorithm to aggregate normal users to obtain a new global model, which is then distributed to all users within the federated learning system. The new global model is used as the base model for the next round of local training to obtain a new local model aggregation server, which then enters a new round of screening and aggregation process.

[0011] Step 4: Repeat steps 1-3 for iteration until the global model converges or the preset termination condition is met.

[0012] The federated learning system contains m adversaries, and the set of adversaries is denoted as U. mal ={c1,c2,c3…cm}, where \(2m + 1 < n\), \(n\) represents the number of users. Each adversary has the same backdoor attack target and understands the working mechanism of the aggregation server, including the anomaly detection metrics adopted by the defense scheme. The amount of local training data of the adversary is the same as that of normal users. During the local model training process, the adversary uses both backdoor data and normal data. In addition, the training parameters and model structure of the adversary are the same as those of normal users;

[0013] The federated learning system considers two types of adversaries, namely ordinary adversaries and adaptive adversaries, and defines these two types of adversaries as follows:

[0014] Definition 1: The goal of the ordinary adversary is to improve the concealment of the attack while maintaining the effectiveness of the backdoor attack. Its local training loss function is defined as shown in formula (1):

[0015]

[0016] where, is the loss function when the adversary conducts local model training, is the main task loss, is the avoidance loss, and assume that is Lipschitz continuous. The ordinary adversary adjusts \(\alpha\) to balance the attack intensity and concealment degree;

[0017] Definition 2: The adaptive adversary can adapt to multiple detection metrics. Its optimization goal is shown in formula (2). This adversary not only minimizes the main task loss and avoidance loss but also enhances the concealment of the attack by jointly adjusting the local model gradient deviation caused by the backdoor attack data, making it more difficult to be detected:

[0018]

[0019] where \(\mu_1\), \(\mu_2\) are weight parameters that respectively control the influence intensity of the gradient deviation adjustment term and the model consistency term, \(\|\cdot\|\) Fro is the Frobenius norm, \(t\) represents the current training round, \(t\in[1,T]\), \(T\) represents the total number of training rounds, represents the local model obtained by user \(u\) ( at the \(t\)-th round of training, is the benign model obtained by user \(u\) ( at the \(t\)-th round of training, \(i,i'\in[1,n]\), \(u\) (2 represents a malicious user, and assume that is Lipschitz continuous.

[0020] In step 1, the initialization is that before the start of the first round of local model training, the aggregation server initializes the global model gradient information \(M\) <Among them, the global model gradient information M < Includes gradient information of the L-layer model, represented as Then, the aggregation server will aggregate the global model gradient information M < Distribute to all users within the Federated Learning System.

[0021] In step 1, the local model is trained on a normal user u within the federated learning system. ( Using local training data d ( Perform model training to obtain a local model Represented as An adversary within the federated learning system uses poisoned local data to train a malicious local model according to formula (1) or formula (2), thereby generating a malicious local model. After updating their local models, all users upload their local models to the aggregation server.

[0022] In step 2, the specific steps for calculating the metrics of gradient cosine similarity, gradient magnitude, gradient direction angle, upper bound of local model gradient update prediction, and lower bound of local model gradient update prediction are as follows:

[0023] Gradient cosine similarity and gradient magnitude calculation: The aggregation server uses formulas (5) and (6) to calculate the gradient magnitude of the local model in round t and its gradient cosine similarity with the global model in the previous round, where d l and d l@1 These represent the dimensions of the model gradient information in the l-th and (l+1)-th layers, respectively. This represents the parameter matrix of the l-th layer of the local model for the i-th user during the t-th round of training. This represents the gradient parameter matrix of the l-th layer in the (t-1)-th round of the global model. This represents the L2 norm of the local model gradient for the i-th user in the t-th round of training, i.e., the gradient magnitude, used to measure the size of the gradient. This represents the parameter element information in the i-th row and j-th column of the user model gradient matrix;

[0024]

[0025] Gradient orientation angle calculation: The aggregation server further calculates the orientation angle between the local model gradient and the previous round global model according to formula (7);

[0026]

[0027] Calculation of the upper and lower bounds of the local model gradient update prediction: The aggregation server calculates the upper bound of the local model gradient update prediction according to formulas (8) and (9), respectively. Local model gradient update predicts lower bound in, This represents the gradient information of the i-th opponent's model in the l-th layer during the t-th round of model training. For the model gradient information of the l-th layer of the global model in the (t-1)th round, ||·|| * Let represent the spectral norm of the gradient matrix, and ∈ represent the perturbation term related to the backdoor attack. These are the weight parameters of the malicious model, p is the proportion of backdoor data in the local training data, and η is... t It is the learning rate during the training process, ρ l It is the minimum gradient information of the l-th layer model, d l It is the dimension of the l-th layer, Lip is the Lipschitz constant, v l yes The minimum non-zero singular value of the product of all layers of the model gradient;

[0028]

[0029]

[0030] Ultimately, the aggregation server obtains the user's local model. The metric is based on five indicators: gradient cosine similarity. gradient magnitude Gradient direction angle Local model gradient update prediction upper bound Local model gradient update predicts lower bound The above l∈[1,L].

[0031] Step 3 specifically involves, after the aggregation server identifies and marks adversaries, using the FedAvg algorithm to aggregate the filtered normal users; the resulting global model is represented as follows: Then, the aggregation server will use the global model M obtained from the t-th round of training. t The global model M, obtained from the t-th round of training, is distributed to all users within the federated learning system. t Then, it will be used as the base model for the next round. Subsequently, users use local training data to conduct the next round of training, obtain a new local model, and upload it to the aggregation server for a new round of filtering and aggregation.

[0032] A federated learning backdoor attack defense system includes:

[0033] The system initialization and local model training module: After the aggregation server initializes the global model and distributes it to all users in the federated learning system, the users train their local models and upload them to the aggregation server.

[0034] The adversary detection module, after receiving the local models, calculates the metric values ​​of each local model in terms of gradient cosine similarity, gradient magnitude, gradient direction angle, upper bound of local model gradient update prediction, and lower bound of local model gradient update prediction. It also calculates the median of each metric value as a reference threshold for normal users. If any one of the four metrics of a user's local model is greater than the reference threshold, or if the metric value of the lower bound of local model gradient update prediction is less than the reference threshold, then the user is identified as an adversary and will not participate in global model aggregation.

[0035] In the global model update module, after the aggregation server identifies and marks adversaries, it uses the FedAvg algorithm to aggregate normal users to obtain a new global model, which is then distributed to all users within the federated learning system. The new global model is used as the base model for the next round of local training, resulting in a new local model aggregation server. This process then enters a new round of screening and aggregation, iterating until the global model converges or reaches the preset termination condition.

[0036] A federated learning backdoor attack defense device includes:

[0037] Memory: Used to store computer programs that implement methods for defending against federated learning backdoor attacks;

[0038] Processor: Used to implement the federated learning backdoor attack defense method when executing the computer program.

[0039] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of a federated learning backdoor attack defense method.

[0040] A computer program product includes a computer program that, when executed by a processor, implements a federated learning backdoor attack defense method.

[0041] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0042] This invention employs five interrelated metrics that cannot be optimized simultaneously: gradient cosine similarity, gradient magnitude, gradient direction angle, upper bound of local model gradient update prediction, and lower bound of local model gradient update prediction. Through statistical testing and analysis, it determines the adversary's dynamic threshold and performs metric extraction and cascaded detection layer by layer on the local model. This achieves effective defense against both ordinary and adaptive adversaries in backdoor attacks in both independent and identically distributed (IID) and non-independent and identically distributed (Non-IID) data scenarios. It maintains high accuracy in the main task across different datasets and model structures, while reducing backdoor task accuracy to near 0%, and is unaffected by the number of adversaries, the degree of non-independent and identically distributed local data, or the attack scaling factor. Attached Figure Description

[0043] Figure 1 This is a framework diagram of the federated learning system of the present invention.

[0044] Figure 2 This is a flowchart of the defense method of the present invention.

[0045] Figure 3(a) is a comparison of the accuracy trends of the defense method of the present invention and the method of using only federated averaging for global model aggregation on the MNIST dataset to perform backdoor attacks in scenario 1.

[0046] Figure 3(b) is a comparison of the accuracy trends of the defense method of the present invention and the method of using only federated averaging for global model aggregation on the MNIST dataset to perform backdoor attacks in scenario 2.

[0047] Figure 4(a) is a comparison of the accuracy trends of the defense method of the present invention and the backdoor attack in scenario 1 on the CIFAR-10 dataset using only federated averaging for global model aggregation.

[0048] Figure 4(b) is a comparison of the accuracy trends of the defense method of the present invention and the global model aggregation using only federated averaging on the CIFAR-10 dataset for backdoor attacks in scenario 2.

[0049] Figure 4(c) is a comparison of the accuracy trends of the defense method of the present invention and the method of using only federated averaging for global model aggregation on the CIFAR-10 dataset for backdoor attacks in scenario 3.

[0050] Figure 4(d) is a comparison of the accuracy trends of the defense method of the present invention and the method of using only federated averaging for global model aggregation to perform backdoor attacks in scenario 4 on the CIFAR-10 dataset.

[0051] Figure 5(a) is a comparison of the accuracy trends of the main task and backdoor task under the condition that IID_ratio=1, using the defense method of the present invention and only using federated averaging for global model aggregation.

[0052] Figure 5(b) is a comparison of the accuracy trends of the main task and backdoor task under the condition of IID_ratio = 0.9, using the defense method of the present invention and using only federated averaging for global model aggregation.

[0053] Figure 5(c) is a comparison of the accuracy trends of the main task and backdoor task under the condition of IID_ratio = 0.8, using the defense method of the present invention and using only federated averaging for global model aggregation.

[0054] Figure 5(d) is a comparison of the accuracy trends of the main task and backdoor task under the condition of IID_ratio = 0.7, using the defense method of the present invention and using only federated averaging for global model aggregation. Detailed Implementation

[0055] The present invention will now be described in detail with reference to the accompanying drawings.

[0056] like Figure 1 As shown, the federated learning system framework consists of two parts: local model training and server aggregation. The main process is as follows:

[0057] During the system initialization phase, the aggregation server generates an initial global model and distributes it to users.

[0058] During the model training phase, users train their local models using the latest global models distributed by the aggregation server, generating updated local models.

[0059] Users upload the updated local model to the aggregation server.

[0060] The aggregation server performs anomaly detection on the local model gradient information uploaded by users, aggregates the model gradients of normal users according to predefined aggregation rules to obtain an encrypted global model, and then distributes the global model to users within the system.

[0061] A federated learning backdoor attack defense method has the following system structure and assumptions:

[0062] The federated learning system contains m adversaries, and the set of adversaries is denoted as U. mal ={c1,c2,c3…c m}, where \(2m + 1 < n\), and \(n\) represents the number of users. Each adversary has the same backdoor attack target and understands the working mechanism of the aggregation server, including the anomaly detection metrics adopted by the defense scheme. The amount of local training data of the adversary (including normal data and backdoor data) is the same as that of normal users. During the local model training process, both backdoor data and normal data are used to poison the global model. In addition, the training parameters and model structure of the adversary are the same as those of normal users. This invention considers two types of adversaries, namely ordinary adversaries and adaptive adversaries, and defines them as follows.

[0063] Definition 1 Ordinary Adversary: The goal of an ordinary adversary is to improve the concealment of the attack while maintaining the effectiveness of the backdoor attack. Its local training loss function is defined as shown in formula (1), where is the loss function when the adversary conducts local model training, is the main task loss, is the evasion loss (i.e., the loss used to avoid the defense algorithm), and it is assumed that is Lipschitz continuous.

[0064]

[0065] The ordinary adversary balances the attack intensity and concealment degree by adjusting \(\alpha\). However, this adjustment will cause differences between the local model of the adversary and the global model in some detection metrics (such as the magnitude and direction of the model gradient), making it easy to be detected under some defense schemes.

[0066] Definition 2 Adaptive Adversary: This scheme introduces an adaptive adversary that can adapt to multiple detection metrics. Its optimization goal is shown in formula (2). This adversary not only minimizes the main task loss and the evasion loss but also enhances the concealment of the attack by jointly adjusting the local model gradient deviation caused by the backdoor attack data, making it more difficult to be detected. In this way, the adaptive adversary can achieve a stealthy and effective backdoor attack under the existing defense schemes for backdoor attacks in federated learning (LYU X, HAN Y, WANG W, et al. Poisoning with cerberus: Stealthy and colluded backdoor attack against federated learning [C] / / Proceedings of the AAAI Conference on Artificial Intelligence. 2023, 37(7): 90XXXX - 90XXXX). Where \(\mu_1\), \(\mu_2\) are weight parameters that respectively control the influence intensity of the gradient deviation adjustment term and the model consistency term, \(\|\cdot\|\) FroIt is the Frobenius norm, where t represents the current training round number, t∈[1,T], and T represents the total number of training rounds. Indicates user u ( The local model obtained in the t-th round of training User u ( The benign model obtained in the t-th training round, i, i′∈[1,n], u (2 Indicates a malicious user, and assumes It is the Lipschitz continuum.

[0067]

[0068] In practical applications of federated learning, the stealth of backdoor attacks often makes it difficult to effectively identify malicious model updates using a single detection metric. Therefore, this invention proposes a backdoor attack defense scheme based on multi-metric cascading. This scheme comprehensively analyzes local model updates through five model gradient-related metrics to improve the detection capability of malicious updates. Specifically, multi-metric uses five metrics—gradient magnitude, gradient cosine similarity, gradient direction angle, and the upper and lower bounds of the local model's gradient update prediction—to perform multi-dimensional feature cross-analysis of abnormal patterns, thereby effectively distinguishing between malicious and normal models. The following sections will provide a detailed introduction to these five metrics.

[0069] Gradient magnitude and gradient cosine similarity: These two metrics are used to detect deviations in the gradient magnitude and direction of model updates, respectively. Gradient magnitude is used to measure the Euclidean distance between the local model update and the global model. The application of this metric can be referenced in the distance metric approach of the FABA algorithm proposed by XIA Q et al. at the 2019 IJCAI conference (XIA Q, TAO Z, HAO Z, et al. FABA: an algorithm for fast aggregation against byzantine attacks in distributed neural networks [C] / / IJCAI.2019.). Gradient cosine similarity is used to measure the directional consistency between the local model update and the global model, similar to the Fldetector scheme proposed by ZHANG Z et al. at the 2022 ACM SIGKDD conference (ZHANG Z, CAO X, JIA J, et al. Fldetector: Defending federated learning against model poisoning attacks via detecting malicious clients [C] / / Proceedings of the 28th ACM SIGKDD conference on knowledge discovery and data mining.2022:2545-2555.) and MA Z et al.'s work in IEEE Transactions on Information Forensics and The ShieldFL scheme, as described in the Security journal (2022), measures directional similarity in a specific way (MA Z, MA J, MIAO Y, et al. ShieldFL: Mitigating model poisoning attacks in privacy-preserving federated learning[J]. IEEE Transactions on InformationForensics and Security, 2022, 17: 1639-1654.). For normal users, the gradient magnitude and direction maintain high consistency with the global model, while backdoor attacks by adversaries cause abnormal changes in the magnitude and direction of the local model.

[0070] Gradient direction angle: gradient cosine similarity M is used to measure the similarity between two gradient directions. tThis represents the global model obtained in the t-th training iteration, but it is essentially a scalar and cannot fully distinguish the gradient direction. When multiple local updates have similar gradient cosine similarity... Even then, significant directional deviations may still exist. Therefore, this invention further calculates the gradient direction angle. As an auxiliary indicator, to enhance the characterization of the model's gradient update direction, This represents the parameter matrix of the l-th layer of the local model for the i-th user during the t-th round of training. Let represent the gradient parameter matrix of layer l in the (t-1)th round of the global model. This is compared to using gradient cosine similarity alone. Gradient direction angle is used to measure the degree of user malice. It provides more refined differentiation capabilities, making malicious updates easier to detect.

[0071] Upper and lower bounds of local model gradient update prediction: In the backdoor attack scenario of federated learning systems, some adversaries have strong adaptive capabilities. As shown in formula (2), the adversary can adjust the local model update while optimizing the main task loss and the backdoor task loss, so that it produces the smallest possible deviation in gradient magnitude and gradient cosine similarity index, in order to circumvent existing defense methods. In the study of LYU X et al. (LYU X, HAN Y, WANG W, et al. Poisoning with cerberus: Stealthy and colluded backdoor attack against federated learning[C] / / Proceedings of the AAAI Conference on Artificial Intelligence.2023,37(7):9020-9028.), when the backdoor attack is successfully embedded, the upper bound of the gradient prediction of its model update is expressed as formula (3), and the lower bound of the gradient prediction is expressed as formula (4). M represents the parameter update matrix of the adversary's local model that initiated the backdoor attack during the t-th round of training for the i-th user. nor,t-1 This represents the benign parameter matrix of the global model in round t-1. Let C represent the spectral norm of the adversary's local model update, and C represent all factors unrelated to the backdoor attack. According to formula (2), when the adaptive adversary minimizes gradient magnitude or gradient similarity bias, it reduces the singular values ​​of the model update, thereby lowering the spectral norm. This means that when the adversary simultaneously evades gradient magnitude and similarity detection, it will suppress... The gradient prediction upper and lower bounds are used as indicators for adaptive adversary detection.

[0072]

[0073] Relationship between indicators: These five gradient-related indicators influence each other in the game of attack and defense, forming a cascade relationship, making it difficult for adversaries to optimize all indicators simultaneously, thereby improving the effectiveness of defense. Specifically, if the adversary's capabilities meet the definition of an ordinary adversary in Definition 1, its optimization goal is to influence the behavior of the global model, causing it to perform abnormally in the backdoor task. However, any attack that can effectively change the behavior of the global model will inevitably cause a large deviation in the local model update. Otherwise, under the action of aggregation strategies such as federated averaging, the attack's impact will be weakened by the updates of most normal users, making it difficult to take effect. This view is consistent with the research conclusions of KRAUβT and DMITRIENKO A in "Mesas: Poisoning defense for federated learning resilient against adaptive attackers" (KRAUβT, DMITRIENKO A.Mesas:Poisoning defense for federated learning resilient against adaptive attackers[C] / / Proceedings of the 2023ACM SIGSAC Conferenceon Computer and Communications Security.2023:1526-1540.). However, local model update bias is directly reflected in three metrics: gradient magnitude, gradient cosine similarity, and gradient direction angle, causing the adversary's model updates to deviate from the distribution of normal user model updates in the gradient space. Therefore, these metrics can effectively detect and identify anomalous updates. However, to evade detection based on gradient bias similarity, the adaptive adversary attempts to minimize local model update bias, but this reduces the spectral norm of the model update, thus suppressing the upper and lower bounds of the local model update's predictions. In other words, when the adversary adjusts its local model update to evade the aforementioned three metrics, it inevitably affects the boundaries of the model update. Therefore, backdoor attack adversaries cannot simultaneously optimize all metrics optimally, and their attempts to evade detection leave exploitable attack metrics at different levels.

[0074] The symbols used in the federated learning backdoor attack defense method described in this invention are illustrated in Table 1.

[0075] Table 1. Symbol Explanation

[0076]

[0077] like Figure 2As shown, the specific steps of this method are as follows:

[0078] Step 1, System Initialization and Local Model Training

[0079] System initialization: Before the first round of local model training begins, the aggregation server initializes the global model gradient information M. < Among them, the global model gradient information M < Includes gradient information of the L-layer model, represented as Then, the aggregation server will aggregate the global model gradient information M < It is distributed to all users within the federated learning system to ensure that the model structure is identical among users.

[0080] Local model training: Normal user u within a federated learning system ( Using local training data d ( Perform model training to obtain a local model Represented as Adversaries within the federated learning system use poisoned local data and train a malicious local model according to formula (1) or formula (2) to generate a malicious local model. After updating their local models, all users upload their local models to the aggregation server.

[0081] Step 2, Adversary Detection

[0082] The aggregation server receives local models from various users. (including malicious local models) Then, the local model's metrics on the above five indicators are calculated respectively, as detailed below:

[0083] Gradient cosine similarity and gradient magnitude calculation: The aggregation server uses formulas (5) and (6) to calculate the gradient magnitude of the local model in round t and its gradient cosine similarity with the global model in the previous round, in order to detect the deviation values ​​of normal users and adversaries in gradient magnitude and gradient direction. Wherein, d l and d l@1 These represent the dimensions of the model gradient information in the l-th and (l+1)-th layers, respectively. This represents the parameter matrix of the i-th user's local model layer i during the t-th round of training. This represents the gradient parameter matrix of the l-th layer in the (t-1)-th round of the global model. This represents the L2 norm (i.e., gradient magnitude) of the local model gradient for the i-th user in the t-th round of training, used to measure the magnitude of the gradient. This represents the parameter element information in the i-th row and j-th column of the user model gradient matrix.

[0084]

[0085] Gradient direction angle calculation: Even if the gradients of two models have the same cosine similarity, there may be significant differences. The aggregation server further calculates the direction angle between the local model gradient and the previous global model according to formula (7).

[0086]

[0087] Calculation of the upper and lower bounds of the local model gradient update prediction: When the backdoor task is successfully embedded during model training, it will cause a deviation between the upper and lower bounds of the local model gradient update prediction. This result has been studied in the article "Poisoning with cerberus: Stealthy and coluded backdoor attack against federated learning" by LYU X et al. (LYU X, HAN Y, WANG W, et al. Poisoning with cerberus: Stealthy and coluded backdoor attack against federated learning[C] / / Proceedings of the AAAI Conference on Artificial Intelligence.2023,37(7):9020-9028.). The aggregation server calculates the upper bound of the local model gradient update prediction according to formulas (8) and (9). Local model gradient update predicts lower bound in, This represents the gradient information of the i-th opponent's model in the l-th layer during the t-th round of model training. For the model gradient information of the l-th layer of the global model in the (t-1)th round, ||·|| * Let represent the spectral norm of the gradient matrix, and ∈ represent the perturbation term related to the backdoor attack. These are the weight parameters of the malicious model, p is the proportion of backdoor data in the local training data, and η is... t It is the learning rate during the training process, ρ l It is the minimum gradient information of the l-th layer model, d l It is the dimension of the l-th layer, Lip is the Lipschitz constant, v l yes The minimum non-zero singular value of the product of all layers of the model gradient;

[0088]

[0089] Ultimately, the aggregation server obtains the user's local model. The metric is based on five indicators: gradient cosine similarity. gradient magnitude Gradient direction angle Local model gradient update prediction upper bound Local model gradient update predicts lower bound The above l∈[1,L].

[0090] Next, the aggregation server first calculates the median of each metric and, based on the assumption of a majority of normal users, uses this median as a reference threshold for normal users. If any one of the four metrics—gradient cosine similarity, gradient magnitude, gradient direction angle, or upper bound of the local model's gradient update prediction—is greater than the reference threshold, or if the metric of the lower bound of the local model's gradient update prediction is less than the reference threshold, then the user is considered an adversary. For example, assuming the reference threshold for model gradient cosine similarity is th_cos, then when... At that time, user u ( If it is identified as an adversary, it will be unable to participate in global model aggregation. The adversary detection algorithm is shown in Algorithm 1.

[0091] Step 3, Global Model Update

[0092] After the aggregation server detects and marks adversaries, it uses the FedAvg algorithm to aggregate the filtered normal users. The resulting global model is represented as follows: Then, the aggregation server will use the global model M obtained from the t-th round of training. t The global model M, obtained from the t-th round of training, is distributed to all users within the federated learning system. Each user receives the global model M after the t-th round of training. t Then, it will be used as the base model for the next round. Subsequently, users use local training data to conduct the next round of training, obtain a new local model, and upload it to the aggregation server to enter a new round of screening and aggregation process.

[0093] Step 4: The entire federated learning system iterates according to the above process until the global model converges or reaches the preset termination condition.

[0094]

[0095]

[0096] A federated learning backdoor attack defense system includes:

[0097] The system initialization and local model training module involves the aggregation server initializing the global model and distributing it to all users within the federated learning system. After the users train their local models and upload them to the aggregation server, this module is used to implement steps 1 and 4 of the method described in this invention.

[0098] The adversary detection module, after receiving the local models, calculates the metric values ​​of each local model in terms of gradient cosine similarity, gradient magnitude, gradient direction angle, upper bound of local model gradient update prediction, and lower bound of local model gradient update prediction. It also calculates the median of each metric value as a reference threshold for normal users. If any one of the four metrics of a user's local model (gradient cosine similarity, gradient magnitude, gradient direction angle, and upper bound of local model gradient update prediction) is greater than the reference threshold, or if the metric value of the lower bound of local model gradient update prediction is less than the reference threshold, then the user is identified as an adversary and will not participate in global model aggregation. This is used to implement steps 2 and 4 of the method described in this invention.

[0099] The global model update module, after the aggregation server identifies and marks adversaries, uses the FedAvg algorithm to aggregate normal users to obtain a new global model, which is then distributed to all users within the federated learning system. The new global model is used as the base model for the next round of local training to obtain a new local model aggregation server, which then enters a new round of screening and aggregation process, iterating until the global model converges or reaches the preset termination condition, in order to implement steps 3 and 4 of the method described in this invention.

[0100] A federated learning backdoor attack defense device includes:

[0101] Memory: Used to store computer programs that implement methods for defending against federated learning backdoor attacks;

[0102] Processor: Used to implement the federated learning backdoor attack defense method when executing the computer program.

[0103] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of a federated learning backdoor attack defense method.

[0104] A computer program product includes a computer program that, when executed by a processor, implements a federated learning backdoor attack defense method.

[0105] Experimental verification

[0106] The effectiveness of this invention is verified through experiments. An independent multi-user distributed learning system was built on a DELL T7920 workstation. The DELL T7920 workstation ran Ubuntu 18.04 with 5*32GB of memory, and used Python 3.8.10 and PyTorch 1.9.0 as simulation software. The training datasets were the MNIST and CIFAR-10 datasets. The MNIST dataset had 60,000 training images and 10,000 test images, each a 28*28 pixel grayscale handwritten digit (0-9). The CIFAR-10 dataset had 50,000 training images and 10,000 test images, each a 32*32 pixel 3-channel color RGB image with 10 categories. Different model structures were used for training on different datasets during the experiments. A Multilayer Perceptron (MLP) was used on the MNIST dataset, with a model structure of [784, 256, 10]. On the CIFAR-10 dataset, a Convolutional Neural Network (CNN) was used, with the CNN model consisting of two convolutional layers and two pooling layers. The relevant system parameters are shown in Table 2, and the specific experimental scenarios are shown in Table 3. Specifically, the backdoor triggers used in this experiment include semantic backdoors and pixel backdoors. Semantic backdoors trigger backdoor attacks based on the semantic content of data samples; for example, a green car is used as a semantic backdoor trigger in the CIFAR-10 dataset. Unlike semantic backdoors, pixel backdoors trigger model attacks by changing the pixel level of data samples. In this experiment, a 2x2 red pixel block is added to a fixed position in the backdoor data samples on both the MNIST and CIFAR-10 datasets. To simulate backdoor attacks in federated learning, this experiment selected two attack methods: scaling attack (BAGDASARYAN E, VEIT A, HUA Y, et al. How to backdoor federated learning[C] / / International conference on artificial intelligence and statistics.PMLR,2020:2938-2948.) and edge attack (WANG H, SREENIVASAN K, RAJPUT S, et al. Attack of the tails: Yes, you really can backdoor federated learning[J]. Advances In Neural Information Processing Systems,2020,33:16070-16084.).

[0107] Table 2 Default parameter settings for multi-metric experiments

[0108]

[0109] Table 3 Multi-metric Experimental Scenario Settings

[0110]

[0111]

[0112] Figures 3(a) and 3(b) show the accuracy trends of the main image classification task and the backdoor task on the MNIST dataset, respectively, under two scenarios: scenario 1 and scenario 2, with the adversary's defense method (multi-metric) deployed and global model aggregation using only federated averaging (FedAvg). (Other default parameters are shown in Table 2.) It can be observed that without the backdoor attack defense mechanism deployed, the adversary's local model participates in global model aggregation; therefore, the accuracy of the global model on the backdoor task continuously improves with the increase in training epochs. After deploying the defense method of this invention, the accuracy of the main task continuously improves with the increase in training epochs, but the accuracy of the backdoor task is almost zero. It can be seen that on the MNIST dataset, this scheme can effectively defend against different backdoor triggers and attack methods used by the adversary.

[0113] Figures 4(a), 4(b), 4(c), and 4(d) show the accuracy changes of the main task and backdoor task on the CIFAR-10 dataset under two scenarios: deploying the multi-metric defense method of this invention to defend against backdoor attacks and using only federated averaging (FedAvg) for global model aggregation, respectively (other default parameters are shown in Table 2). The experimental results show that without any other defense schemes deployed, the accuracy of the global model on the main task continuously increases with the number of training epochs, but the accuracy of the backdoor task also rapidly reaches 100% in a short time, indicating that the adversary successfully embedded the backdoor attack into the global model. After deploying the multi-metric defense method of this invention, the accuracy of the global model on the main task continuously improves with the number of training epochs, while the accuracy of the backdoor task remains almost 0% after model convergence. It can be seen that the performance of the defense method of this invention on the CIFAR-10 dataset is the same as that on the MNIST dataset. Combining Figure 3(a), Figure 3(b) and Figures 4(a)-4(d)The experimental results show that this solution can effectively defend against backdoor attacks regardless of changes in the dataset, model structure, backdoor triggers, and backdoor attack methods.

[0114] Figures 5(a)-5(d) Table 2 presents the accuracy changes of the global model on the main task and backdoor task under four different data distribution conditions (other default parameters are shown in Table 2). It can be seen that after deploying the defense method (multi-metric) of this invention, the accuracy of the main task continuously improves with the increase of training epochs under different IID_ratio values, while the accuracy of the backdoor task remains almost zero. This demonstrates that the effectiveness of this scheme is not affected by the degree of non-independent and identically distributed local data.

[0115] To verify the superiority of the proposed scheme compared with other existing defense schemes, FedAvg, Krum (AHMEDN, NATARAJAN T, RAO K R. Discrete cosine transform[J]. IEEE Transactions on Computers, 2006, 100(1): 90-93.), Flame (GUO H, GREENGARD P, WANG H, et al. Federated learning as variational inference: A scalable expectation propagation approach[J]. arXiv preprint arXiv:2302.04228, 2023.), and FLTrust (WANG H, SREENIVASAN K, RAJPUT S, et al. Attack of the tails: Yes, you really can backdoor federated learning[J]. Advances In Neural Information Processing (Systems, 2020, 33:16070-16084.) These four schemes were used as comparison schemes to compare the accuracy of the global model on the main task and the target task under different system scenarios and different adversary capabilities (other default parameters are shown in Table 2). The experimental results are shown in Tables 4 and 5. First, it can be seen from Tables 4 and 5 that when only the FedAvg algorithm is used, the backdoor task accuracy is 100%. This is because FedAvg, as the most classic federated learning aggregation algorithm, adopts the method of directly averaging the local models of each user, without any defense capability. Therefore, the adversary can successfully attack in all attack scenarios. Next, this experiment evaluated the defense capability of Krum against different adversaries in different scenarios. Krum calculates the Euclidean distance between the user's local models and selects one of the models as the global model. The adversary can manipulate its local model through formula (1) or formula (2) to make the models of the adversaries more similar, thereby easily avoiding Krum defense. As shown in Tables 4 and 5, the accuracy of the backdoor mission in different scenarios mostly reached 100%.

[0116] Next, this experiment evaluated the performance of the Flame defense scheme. Flame effectively defends against common adversary attacks by filtering out malicious models with large angular deviations, dynamically pruning to limit the influence of malicious models, and selecting appropriate noise cancellation backdoors. As shown in Tables 4 and 5, on the CIFAR-10 and MNIST datasets, when the adversary's capability is a common attack, the global model's backdoor task accuracy is only 27.18% at its highest. However, when the adversary has adaptive capabilities, Flame's effectiveness is greatly reduced because it can simultaneously adapt to the angular deviation of the defense scheme and the model's Euclidean distance metric. In scenarios 2 and 4, which include edge attacks, the backdoor task accuracy reaches 100%, and in scenarios 1 and 3, which include scaling attacks, the backdoor task accuracy reaches a maximum of 90.41% and 80.75%, respectively. Furthermore, this experiment evaluated the FLTrust defense scheme, which is based on a root dataset and uses the similarity score between models to update the aggregate weights of the user's local model. If the adversary's anomalous features are not obvious on these two metrics, it will affect its detection success rate. Meanwhile, if the weight assigned to normal users is too small, the accuracy of the main task may be sacrificed. As shown in Tables 4 and 5, when the adversary launches an adaptive attack, the accuracy of the FLTrust global model's main task is lower than other defense schemes under the same conditions. FLTrust is effective in defending against normal attacks. However, when the adversary launches an adaptive attack, the defensive effect of FLTrust decreases; for example, in scenario 4 of the CIFAR-10 dataset, the backdoor task accuracy reaches 96.61%. Finally, this experiment verifies the performance of the defense method (multi-metric) of this invention. It can be seen that multi-metric is effective in defending against both normal and adaptive attacks. In all scenarios, the accuracy of the global model's main task is close to that of FedAvg, and the accuracy of the backdoor task remains 0 in all scenarios.

[0117] Table 4 shows the accuracy comparison between the proposed solution and existing solutions on the CIFAR-10 dataset.

[0118]

[0119] Table 5 shows the accuracy comparison between the proposed solution and existing solutions on the MNIST dataset.

[0120]

Claims

1. A method for defending against federated learning backdoor attacks, characterized in that, The steps include: Step 1: After the aggregation server initializes the global model and distributes it to all users within the federated learning system, users train their local models and upload them to the aggregation server. Step 2: After receiving the local models, the aggregation server calculates the metric values ​​of each local model in terms of gradient cosine similarity, gradient magnitude, gradient direction angle, upper bound of local model gradient update prediction, and lower bound of local model gradient update prediction. It also calculates the median of each metric value as a reference threshold for normal users. If any one of the four metrics of a user's local model (gradient cosine similarity, gradient magnitude, gradient direction angle, or upper bound of local model gradient update prediction) is greater than the reference threshold, or if the metric value of the lower bound of local model gradient update prediction is less than the reference threshold, the user is identified as an adversary and will not participate in global model aggregation. Step 3: After the aggregation server identifies and marks adversaries, it uses the FedAvg algorithm to aggregate normal users to obtain a new global model, which is then distributed to all users within the federated learning system. The new global model is used as the base model for the next round of local training to obtain a new local model aggregation server, which then enters a new round of screening and aggregation process. Step 4: Repeat steps 1-3 for iteration until the global model converges or the preset termination condition is met.

2. The method according to claim 1, characterized in that, There are m adversaries in the federated learning system, and the adversary set is represented as U mal ={c1, c2, c3…c m}, where 2m + 1 < n, n represents the number of users. Each adversary has the same backdoor attack target and understands the working mechanism of the aggregation server, including the anomaly detection metrics adopted by the defense scheme. The amount of local training data of the adversary is the same as that of normal users. During the local model training process, both backdoor data and normal data are used. In addition, the training parameters and model structure of the adversary are the same as those of normal users; The federated learning system considers two types of adversaries: ordinary adversaries and adaptive adversaries, which are defined as follows: Definition 1: The objective of the common adversary is to improve the stealth of the attack while maintaining the effectiveness of the backdoor attack. Its local training loss function is defined as shown in formula (1): in, The loss function used when training a local model for the adversary. Loss to the main task To avoid losses, and assuming It is the Lipschitz continuum, where ordinary adversaries balance attack strength and stealth by adjusting α; Definition 2: The adaptive adversary can adapt to multiple detection metrics, and its optimization objective is shown in Equation (2). This adversary not only minimizes the main task loss and avoidance loss, but also enhances the concealment of the attack by jointly adjusting the local model gradient bias caused by the backdoor attack data, making it more difficult to detect. Where μ1 and μ2 are weight parameters, controlling the influence strength of the gradient bias adjustment term and the model consistency term, respectively. Fro It is the Frobenius norm, where t represents the current training round number, t∈[1,T], and T represents the total number of training rounds. Indicates user u i The local model obtained in the t-th round of training User u i The benign model obtained in the t-th training round, i,i ′ ∈[1,n],u i2 Indicates a malicious user, and assumes It is the Lipschitz continuum.

3. The method according to claim 1, characterized in that, In step 1, the initialization refers to the aggregation server initializing the global model gradient information M before the first round of local model training begins. 0 Among them, the global model gradient information M 0 Includes gradient information of the L-layer model, represented as Then, the aggregation server will aggregate the global model gradient information M 0 Distribute to all users within the Federated Learning System.

4. The method according to claim 1, characterized in that, In step 1, the local model is trained on a normal user u within the federated learning system. i Using local training data d i Perform model training to obtain a local model Represented as An adversary within the federated learning system uses poisoned local data to train a malicious local model according to formula (1) or formula (2), thereby generating a malicious local model. After updating their local models, all users upload their local models to the aggregation server.

5. The method according to claim 1, characterized in that, In step 2, the specific steps for calculating the metrics of gradient cosine similarity, gradient magnitude, gradient direction angle, upper bound of local model gradient update prediction, and lower bound of local model gradient update prediction are as follows: Gradient cosine similarity and gradient magnitude calculation: The aggregation server uses formulas (5) and (6) to calculate the gradient magnitude of the local model in round t and its gradient cosine similarity with the global model in the previous round, where d l and d l+1 These represent the dimensions of the model gradient information in the l-th and (l+1)-th layers, respectively. This represents the parameter matrix of the l-th layer of the local model for the i-th user during the t-th round of training. This represents the gradient parameter matrix of the l-th layer in the (t-1)-th round of the global model. This represents the L2 norm of the local model gradient for the i-th user in the t-th round of training, i.e., the gradient magnitude, used to measure the size of the gradient. This represents the parameter element information in the i-th row and j-th column of the user model gradient matrix; Gradient orientation angle calculation: The aggregation server further calculates the orientation angle between the local model gradient and the previous round global model according to formula (7); Calculation of the upper and lower bounds of the local model gradient update prediction: The aggregation server calculates the upper bound of the local model gradient update prediction according to formulas (8) and (9), respectively. Local model gradient update predicts lower bound in, This represents the gradient information of the i-th opponent's model in the l-th layer during the t-th round of model training. For the model gradient information of the l-th layer of the global model in the (t-1)th round, ||·|| * Let represent the spectral norm of the gradient matrix, and ∈ represent the perturbation term related to the backdoor attack. These are the weight parameters of the malicious model, p is the proportion of backdoor data in the local training data, and η is... t It is the learning rate during the training process, ρ l It is the minimum gradient information of the l-th layer model, d l It is the dimension of the l-th layer, Lip is the Lipschitz constant, v l yes The minimum non-zero singular value of the product of all layers of the model gradient; Ultimately, the aggregation server obtains the user's local model. The metric is based on five indicators: gradient cosine similarity. gradient magnitude Gradient direction angle Local model gradient update prediction upper bound Local model gradient update predicts lower bound The above l∈[1,L].

6. The method according to claim 1, characterized in that, Step 3 specifically involves, after the aggregation server identifies and marks adversaries, using the FedAvg algorithm to aggregate the filtered normal users; the resulting global model is represented as follows: Then, the aggregation server will use the global model M obtained from the t-th round of training. t The global model M, obtained from the t-th round of training, is distributed to all users within the federated learning system. t Then, it will be used as the base model for the next round. Subsequently, users use local training data to conduct the next round of training, obtain a new local model, and upload it to the aggregation server for a new round of filtering and aggregation.

7. A federated learning backdoor attack defense system based on the method of any one of claims 1 to 6, characterized in that, include: The system initialization and local model training module: After the aggregation server initializes the global model and distributes it to all users in the federated learning system, the users train their local models and upload them to the aggregation server. The adversary detection module, after receiving the local models, calculates the metric values ​​of each local model in terms of gradient cosine similarity, gradient magnitude, gradient direction angle, upper bound of local model gradient update prediction, and lower bound of local model gradient update prediction. It also calculates the median of each metric value as a reference threshold for normal users. If any one of the four metrics of a user's local model is greater than the reference threshold, or if the metric value of the lower bound of local model gradient update prediction is less than the reference threshold, then the user is identified as an adversary and will not participate in global model aggregation. In the global model update module, after the aggregation server identifies and marks adversaries, it uses the FedAvg algorithm to aggregate normal users to obtain a new global model, which is then distributed to all users within the federated learning system. The new global model is used as the base model for the next round of local training, resulting in a new local model aggregation server. This process then enters a new round of screening and aggregation, iterating until the global model converges or reaches the preset termination condition.

8. A federated learning backdoor attack defense device, characterized in that, include: Memory: Used to store computer programs that implement the federated learning backdoor attack defense method as described in any one of claims 1 to 6; Processor: Used to implement the federated learning backdoor attack defense method as described in any one of claims 1 to 6 when executing the computer program.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the federated learning backdoor attack defense method as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the federated learning backdoor attack defense method as described in any one of claims 1 to 6.

Citation Information

Cited By

  • Privacy protection federated distillation and backdoor defense method for large model fine tuning

    CN121256789A