Federal learning backdoor defense method based on gradient screening and weight adjustment

By implementing adaptive gradient sparsity and adaptive weight adjustment on the server side, the performance-efficiency contradiction of backdoor defense methods in federated learning is resolved, achieving efficient and robust defense in non-independent and identically distributed data environments, which is applicable to distributed intelligent systems such as finance and healthcare.

CN120979739AActive Publication Date: 2025-11-18NANKAI UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511164019.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-20
Publication Date
2025-11-18
Estimated Expiration
2045-08-20

AI Technical Summary

Technical Problem

Existing federated learning backdoor defense methods struggle to balance defense effectiveness with model performance, suffer from high computational complexity, and exhibit poor adaptability in environments with non-independent and identically distributed data.

Method used

A defense method based on gradient filtering and weight adjustment is adopted. By adaptive gradient sparsity and adaptive weight adjustment on the server side, potential backdoor attacks are identified and weakened.

Benefits of technology

While maintaining high model accuracy, it effectively resists various backdoor attacks, reduces computational overhead, and maintains stability and robustness in complex data environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120979739A_ABST
    Figure CN120979739A_ABST
Patent Text Reader

Abstract

The invention discloses a federated learning backdoor defense method based on gradient screening and weight adjustment, and relates to the field of federated learning security. According to the method, firstly, potential abnormal updating is identified and eliminated through gradient adaptive rarefaction, clustering analysis and abnormal detection, so that redundant information is reduced, and the risk of hostile attacks is reduced; then, the weight of each client is adaptively adjusted according to factors such as gradient similarity in the aggregation process, the model aggregation process is optimized, it is ensured that contribution of a benign client is effectively utilized, and meanwhile the influence of a malicious client is weakened; compared with an existing defense algorithm, on the premise that the accuracy of the model is not remarkably sacrificed, various backdoor attacks can be effectively resisted, excellent defensive performance, robustness and low calculation complexity are shown in different data environments, and the method can be widely applied to various federated learning systems.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of federated learning security, and is a defense method against backdoor attacks. As a new distributed machine learning architecture, federated learning allows multiple devices or institutions (clients) to collaboratively train a shared global model without sharing local sensitive data (such as medical records, financial transaction details), thereby breaking down data silos, improving model performance while protecting data privacy. BACKGROUND

[0002] The distributed nature of federated learning exposes it to serious security threats, especially backdoor attacks. In a backdoor attack, a malicious client intentionally injects hidden malicious behavior into the global model by implanting specific "triggers" in the local training data. When the model is deployed, inputs carrying the trigger will be incorrectly classified into the target class preset by the attacker with high probability, while for normal data, the model behaves as usual, making such attacks difficult to detect. Existing defense strategies are mainly divided into two categories: backdoor mitigation and backdoor detection. They limit the impact of malicious updates through algorithm design, or identify and exclude malicious updates through monitoring systems to ensure model security.

[0003] Existing federated learning backdoor defense technologies generally face a core contradiction in practicality. On the one hand, they face the "trade-off dilemma" between defense effectiveness and model performance. In order to resist attacks, these methods often compromise the model's performance on normal tasks, resulting in a significant drop in classification accuracy. On the other hand, many defense mechanisms involve complex computational processes, leading to a dramatic increase in training time and high computational and communication overhead. This approach of sacrificing performance and efficiency for security greatly limits its application value in real-world scenarios.

[0004] In addition, the universality and stability of existing defense methods also face serious challenges. The most important problem is that most methods assume that client data is independently and identically distributed, but real-world data often exhibits non-independent and identically distributed characteristics. When faced with complex data distribution, the effectiveness of these defense methods will decrease dramatically, making it difficult to achieve the desired effect, severely limiting their universality. At the same time, some defense technologies themselves also have limitations, for example, their effectiveness may be highly dependent on clustering accuracy and other specific conditions, or have limited adaptability when dealing with variable attacks. SUMMARY

[0005] The present application aims to solve the problems of existing federal learning backdoor defense methods, such as difficulty in balancing defense effect and model performance, high computational complexity, and poor adaptability in non-independent and identically distributed data environment, and provides a federal learning backdoor defense method based on gradient screening and weight adjustment, which can effectively resist various backdoor attacks, maintain high precision on normal tasks, and ensure efficient, stable and robust defense performance in complex non-IID data environment.

[0006] The federal learning backdoor defense method based on gradient screening and weight adjustment provided by the present application is characterized by the collaborative processing of the update information uploaded by the clients in the server aggregation stage.

[0007] To achieve the above purpose, the technical solution adopted by the present application is:

[0008] A federal learning backdoor defense method based on gradient screening and weight adjustment, which includes three steps of client local training, server-side collaborative processing and defense, and secure aggregation and model updating, wherein:

[0009] Step 1: Client local training

[0010] At the beginning of each round of communication, the selected clients first download the latest global model from the server, then each client trains the model using its local data set, in this stage, the malicious client will implant a backdoor trigger in its local data set to train a local model containing malicious intent, after training, all selected clients upload the updated model parameters or gradients to the server and wait for aggregation;

[0011] Step 2: Server-side collaborative processing and defense

[0012] After receiving the updates uploaded by the clients, the server does not directly aggregate, but performs collaborative defense operations on all collected model updates, including:

[0013] 1. Gradient adaptive sparsification

[0014] In order to retain key information while eliminating potential malicious updates, the server performs hierarchical and adaptive sparsification processing on the gradients uploaded by each client;

[0015] 2. Adaptive weight adjustment

[0016] After sparsification, the server dynamically adjusts the aggregation weights of the clients according to the similarity between their updates;

[0017] Step 3: Secure aggregation and model update

[0018] The server uses the gradient processed through the above steps, and combines the respective adaptive weights to perform weighted averaging, and finally generates a new round of global model; finally, the updated global model will be distributed to the next round of clients participating in training, and the process is repeated until the model converges.

[0019] Further, in the second step of server-side collaborative processing and defense of the present application, step 1 gradient adaptive sparsification includes:

[0020] 1) Dynamic threshold calculation: first, according to the training progress et and the total parameter quantity N l of the network layer l l(t) , a dynamically changing sparsification base T l is calculated. In the early stage of training, the value is small, and more strict screening is implemented; as the training progresses, the value gradually increases, allowing more parameter updates to pass;

[0021] 2) Layer importance weighting: considering that different layers in the neural network have different importance (for example, weight layers usually carry more learning ability than bias layers), a layer importance coefficient I s is introduced. The weight layer is given a higher weight value, while the bias layer weight is lower, so as to protect the more important layer during sparsification;

[0022] 3) Outlier detection: in order to filter out abnormal gradient values that deviate too much, the scheme calculates the mean and variance of the gradient of each layer, and uses the Z-score method to identify outliers. Only the gradient with a Z-score value less than the preset threshold ζ is considered normal, and an anomaly detection mask z l is generated;

[0023] 4) Generate final mask: combine dynamic threshold, layer importance and outlier detection, the server determines a final parameter quantity k l for each layer, then use the TopK function to retain only the k s largest absolute values in each layer gradient, and multiply it with the anomaly detection mask z

[0024] Further, in the second step of server-side collaborative processing and defense of the present application, step 2 adaptive weight adjustment includes:

[0025] 1) Client clustering: using the HDBSCAN clustering method, the clients are grouped based on the cosine distance of the sparsified gradient, which can effectively cluster clients with similar behavior (possibly benign or malicious) together, and assign each client i a cluster label b i ;

[0026] 2) clipping rate calculation: calculate the Euclidean distance between the gradients of each client, and derive a clipping rate rho based on this i for controlling the magnitude of gradient clipping;

[0027] 3) calculation of adaptive weight factor: in order to solve the shortcomings of traditional fixed weight strategy, this scheme introduces an adaptive weight adjustment factor lambda i , which is calculated based on the cosine similarity distribution of all client sparse gradients, and can dynamically fine-tune the contribution of each client in aggregation according to the data of the current round.

[0028] Further, after accurately calculating the adaptive tuning parameter lambda, by embedding all obtained tuning parameters into the established architecture of the aggregation function, a new generation of model parameters is generated:

[0029]

[0030] In order to effectively reduce the influence of abnormal samples, noise is added to the weight of the model,

[0031]

[0032] where N(0, sigma 2 ) represents the added Gaussian noise function, and the purpose of adding noise is to reduce the influence of backdoor attack and enhance the defense effect.

[0033] The present application uses a cooperative defense mechanism to perform double processing on the update information uploaded by the client in the server aggregation stage of federated learning to identify and weaken potential backdoor attacks. The core principle is first gradient adaptive sparsification, that is, through a dynamically adjusted, hierarchical perception screening strategy combined with outlier detection technology, only the parameters with the largest amplitude and statistically most critical in each client gradient update are retained, thereby filtering out most of the gradient information implanted by attackers, which is usually abnormal or redundant, while retaining the core learning ability of the model. On this basis, the scheme further adopts the principle of adaptive weight adjustment, dynamically allocates an aggregation weight to each client by similarity clustering and quantitative analysis of the sparse gradients, ensuring that the behavior of the client is reliable and the contribution is positive, and the influence of the suspicious client is significantly reduced. Finally, this cooperative processing method of filtering (sparsification) and weighting (weight adjustment) can effectively purify and optimize the contributions of each party before aggregation, thereby stably resisting malicious attacks while ensuring the accuracy of the model.

[0034] The technical effect of the present application is:

[0035] The application provides an innovative collaborative idea and implementation path for backdoor defense of federated learning. Compared with the prior art, it has significant advantages in defense performance, model accuracy, computational overhead and environmental adaptability. The core advantage lies in the unique dual collaborative processing mechanism, which fundamentally solves the contradiction between "high security" and "high model performance" in traditional defense schemes. The scheme first uses gradient adaptive sparsification technology to accurately retain the core gradients that contribute most to model learning while eliminating abnormal updates through dynamic threshold and hierarchical weighting. Then, through an adaptive weight adjustment strategy, clustering and dynamic weighting based on the similarity of client updates are performed to ensure that trusted clients dominate in model aggregation. This "precise filtering and then weighting" collaborative operation enables the application to effectively resist various backdoor attacks while maximizing the protection of the original accuracy of the model.

[0036] The application also achieves major breakthroughs in running efficiency and environmental robustness. In terms of efficiency, the core method of the scheme is relatively lightweight in computation, avoiding the high computational overhead caused by the introduction of complex game theory or generated models in many existing technologies, significantly reducing the aggregation time of the server, making it more suitable for deployment in practical applications that prioritize efficiency. In terms of robustness, the defense mechanism of the application is completely data-driven and does not rely on the idealized assumption of independent and identically distributed data. In the face of non-independent and identically distributed data that is prevalent in the real world, its adaptive characteristics can still maintain stable and reliable defense performance, while many traditional methods will experience a sharp decline in performance in such complex environments. This efficient and robust security solution makes it widely applicable to distributed intelligent systems in finance, healthcare, and other fields that require high security and performance. BRIEF DESCRIPTION OF DRAWINGS

[0037] Figure 1 is a federated learning aggregation framework diagram.

[0038] Figure 2 is a workflow diagram of the federated learning backdoor defense method of the application.

[0039] Figure 3 (a-d) are defense performance diagrams of the application and its two variants on MNIST, FMNIST, EMNIST, and CIFAR data sets.

[0040] Figure 4 (a-d) are experimental diagrams of the influence of different sparsity rates a on the defense performance (accuracy and attack success rate) of the FLSDA method on four data sets.

[0041] Figure 5 is an experimental diagram for evaluating the performance stability of the application when the proportion of malicious clients changes from 0.1 to 0.5. DETAILED DESCRIPTION

[0042] Embodiment 1:

[0043] The specific embodiments of the present application are described in detail below with reference to the accompanying drawings and embodiments. The embodiments are implemented on the premise of the technical solutions of the present application, so that the technical characteristics and functional characteristics of the present application can be better understood by those skilled in the related art.

[0044] The present application provides a federal learning backdoor defense method based on gradient screening and weight adjustment, aiming to solve the problem of backdoor attack caused by malicious clients in federal learning. Existing defense methods often face the challenges of model performance decline, high computational complexity or poor effect on non-independent and identically distributed data. The present application introduces two core strategies of gradient adaptive sparsification and client weight adaptive adjustment in the server aggregation stage, efficiently defends various backdoor attacks, and maintains the performance and computational efficiency of the model.

[0045] In the present embodiment, the present application is implemented based on Python language and Pytorch deep learning framework, and the running environment is a server with 12th Gen Intel(R)Core(TM)i5-12600K, 3.70GHz GPU 4070tisuper and RAM 64G. The present application selects MNIST, CIFAR-10, F-MNIST and E-MNIST as benchmark datasets, and simulates multiple attack schemes including distributed backdoor attack (DBA), target label flipping attack (TLF) and edge case backdoor attack (ECBA).

[0046] First, referring to the accompanying drawings Figure 1 , a standard federal learning aggregation framework is shown. In this framework, the server distributes the global model to N clients, each client trains the local data D i , and then uploads the updated model parameters ω i to the server. The server aggregates all uploaded parameters to update the global model. However, the openness of this process provides an opportunity for malicious clients to inject backdoor attacks.

[0047] The core process of the present application is shown in the accompanying Figure 2 , which mainly cooperatively defends the uploaded model updates in the server aggregation stage. The specific process is as follows:

[0048] Step one, client training and uploading: In each round of communication, part of the selected clients download the current global model and train it using local data. Among them, the normal client performs standard training, while the malicious client adds a specific trigger in the training data, so that the model update carries the backdoor intention, to train a local model containing malicious intent. After training, all selected clients upload the local model update to the server and wait for aggregation.

[0049] Step two, server-side collaborative defense: After receiving the updates uploaded by the clients, the server does not directly aggregate. Before aggregation, the server performs the following collaborative defense operations on all collected model updates:

[0050] Step 1: Gradient adaptive sparsification. This step aims to retain key benign update information while filtering out a large number of redundant or malicious gradients. This method first dynamically calculates the sparsification threshold k l ,

[0051] k l =max(k min ,min(T l (t),k max ))

[0052]

[0053] This threshold function takes into account multiple factors. et represents the training progress (0≤et≤1), allowing dynamic adjustment of the threshold during training. N l represents the number of parameters in layer l. The method adopts a more stringent filtering at the beginning of training and gradually relaxes it later to balance defense strength and model convergence speed. According to the type, position and importance of each layer, different sparsification strategies are tailored for each layer's update.

[0054]

[0055] At the same time, by calculating the mean and variance of each layer's gradient to identify and mask outliers, further purify the gradient data.

[0056]

[0057] where z s represents the obtained anomaly detection mask, d l represents the data of the lth layer, ζ is the anomaly threshold, ∈ is a small constant to ensure numerical stability. After processing, only the most important and non-anomalous gradient values are retained, effectively weakening the impact of malicious updates. Finally, the anomaly detection mask is combined with the sparsification mask to generate a complete processing mask:

[0058]

[0059] where, denotes the processed gradient value for subsequent aggregation, TopK(X, k) denotes the values corresponding to the top k absolute values in dataset X, and all other values are set to 0.

[0060] Step 2: Adaptive weight adjustment. After the sparse processing, the clients are then clustered and the weights are adjusted. First, the HDBSCAN clustering algorithm is used to group the clients based on the cosine distance between their updates, identifying potential clusters of malicious clients. Then, an adaptive aggregation weight λ i is calculated for each client based on factors such as the similarity of their gradients. This weight dynamically adjusts the contribution of each client in the global model update, reducing the influence of clients identified as abnormal. The specific steps are as follows:

[0061] 1) Client clustering: Use the HDBSCAN clustering method to group clients based on the cosine distance of the sparse gradients. This effectively clusters similar behaving clients together and assigns each client i a cluster label b i .

[0062] 2) Clipping rate calculation: Calculate the Euclidean distance between the gradients of each client and derive a clipping rate ρ i based on it, which controls the magnitude of gradient clipping.

[0063] 3) Calculate adaptive weight factor: To address the shortcomings of traditional fixed weight strategies, this solution introduces an adaptive weight adjustment factor λ i .

[0064]

[0065] where This factor is calculated based on the cosine similarity distribution of all clients' sparse gradients, and can dynamically fine-tune the contribution of each client in aggregation according to the data of the current round. After accurately calculating the adaptive tuning parameter λ, the key integration step is performed. By embedding all obtained tuning parameters into the established architecture of the aggregation function, a new generation of model parameters is generated:

[0066]

[0067] To effectively reduce the impact of abnormal samples, noise is added to the model's weights.

[0068]

[0069] where N(0, σ2 ) represents the added Gaussian noise function. The purpose of adding noise is to reduce the impact of backdoor attacks and enhance defense effects, but the amount of added noise is relatively small to avoid affecting model performance.

[0070] Step three, secure aggregation and model update: the server uses the gradient processed by the above steps, and combines the respective adaptive weights for weighted average, and finally generates a new round of global model. To further enhance robustness, the global model after aggregation also adds a small amount of Gaussian noise. Finally, the updated global model will be distributed to the next round of participating clients for training, and this process will be repeated until the model converges.

[0071] The above steps (local training, collaborative processing, secure aggregation) constitute a complete communication round. By repeating T rounds, the global model is continuously optimized, and finally converges to an ultimate model G that can resist backdoor attacks and maintain high performance T .

[0072] The implementation effect of the present application is explained as follows:

[0073] The beneficial effects of the present application are verified by a series of comparative experiments, which are specifically explained as follows:

[0074] The data in Tables 1, 2 and 3 are the core basis for evaluating the performance of the FLSDA of the present application. They compare FLSDA with FedAvg, FLAME, Krum and other seven mainstream defense methods in various attack scenarios, data sets and data distributions. The evaluation criteria take into account the main task accuracy and backdoor attack success rate, and the bold data shows the optimal result in each test. These tables jointly prove the comprehensive advantages of FLSDA in defense effectiveness, robustness and universality.

[0075] Table 1 Performance comparison of FLSDA and other defense methods against DBA

[0076]

[0077] In terms of defending against distributed backdoor attacks (DBA),

[0078] The data in Table 1 highlights the superior stability of FLSDA. The performance of many defense methods drops sharply when moving from independent and identically distributed (IID) data to non-independent and identically distributed (Non-IID) data, for example, the attack success rate of Krum method on CIFAR dataset deteriorates from 3.65% to 74.31%. In contrast, FLSDA effectively overcomes the challenges brought by data heterogeneity, maintaining extremely low attack success rates and high accuracy in both environments, demonstrating strong robustness. For example, on the Non-IID CIFAR dataset, the attack success rate of FLSDA is only 9.96%, significantly better than most other methods.

[0079] Table 2: Performance comparison of FLSDA and other defense methods against TLF

[0080]

[0081] In response to the target label flipping attack (TLF),

[0082] Table 2 shows that although the defense of such attacks is relatively low, FLSDA still performs outstandingly. Especially on F-MNIST datasets, when other methods struggle to achieve ideal defense effects, FLSDA consistently provides robust defense capabilities under any data distribution. This proves that the invention has reliable performance in different attack scenarios, ensuring the security bottom line of the system.

[0083] Table 3: Performance comparison of FLSDA and other defense methods against ECBA

[0084]

[0085] In the most challenging edge case backdoor attack (ECBA) scenario,

[0086] The data in Table 3 confirms the strong competitiveness of FLSDA. In this scenario, FLSDA achieved the best defense results in almost all experimental environments. Especially on the CIFAR dataset, which is difficult to defend, the defense effect of FLSDA is far superior to other methods, with an attack success rate of 1.53% in the IID environment, while the attack success rates of several comparative methods are above 10% or even higher. This fully demonstrates the significant superiority of the invention in dealing with complex and hidden attacks.

[0087] In addition to the table data, the technical effects of the invention are also visually demonstrated through the accompanying drawings:

[0088] Appendix Figure 3The results of ablation experiments are shown to verify the synergistic effect of the strategies in the present application. The baseline method (Fedavg), the method using only sparsification (FLSDA1), the method using only weight adjustment (FLSDA2), and the complete FLSDA method are compared in the figure. The experimental results clearly show that the individual strategies have certain improvement, but the effect is limited, while the complete FLSDA method combining gradient sparsification and weight adjustment can reduce the backdoor attack success rate (line graph) to the lowest while maintaining high main task accuracy (bar graph), proving the necessity and superiority of the synergistic effect of the two strategies.

[0089] The Figure 4 The influence of the key process parameter "sparsity α" on the method performance is shown. By setting α to 1, 1.5, 2, 2.5, 3, 4 and other different values for testing, the results show that when α is 3, the method can achieve the best balance on all test data sets, that is, a higher classification accuracy (solid line) and a lower attack success rate (dashed line). This provides experimental basis for the selection of this process parameter, and is set as the preferred value in the present embodiment.

[0090] The Figure 5 The robustness of the present application under different "proportion of malicious clients" is shown. The proportion of malicious clients is adjusted from 0.1 to 0.5, and the change in method performance is observed. The results show that in the first three data sets such as MNIST, even if the proportion of malicious clients is as high as 50%, the performance of FLSDA remains relatively stable; on the more complex CIFAR data set, the performance has decreased, but still remains within an acceptable range. This shows that the present application has good defense stability and adaptability to attacks of different intensities.

[0091] Based on the above embodiments, the federated learning backdoor defense method based on gradient screening and weight adjustment proposed by the present application is key to a set of synergistic, data-driven adaptive processing mechanisms used in the server aggregation stage. This mechanism is not a single defense means, but is composed of two innovative steps of gradient adaptive sparsification and adaptive weight adjustment, which achieves the purpose of the present application through the overall synergistic effect and achieves the ideal technical effect. The "gradient adaptive sparsification" has the innovation of multi-dimensional adaptability, which is not simply discarding small magnitude gradients, but considering three major elements of training dynamic process, network level importance and statistical anomaly. It dynamically adjusts the screening intensity according to different stages of training, and provides protection for network layers that carry more learning ability, while using statistical methods to accurately identify and shield abnormal gradient values that deviate significantly from the overall distribution.

[0092] The adaptive weight adjustment is key to the self-adapting weight adjustment based on the sparsification. After the preliminary gradient purification, the application introduces a dynamic weighting system based on the similarity of client updates. It uses the HDBSCAN clustering algorithm to group clients with similar update behaviors. At the same time, it calculates an adaptive weight adjustment factor for each client. The factor acts as a real-time regulating valve, which can fine-tune the contribution proportion of each client in the final aggregation according to the similarity of the client update and the group update in the current round. This design completely abandons the static or simple weighting method in traditional defense, and realizes the intelligent and dynamic management of the contribution of clients. Therefore, the core protection point of the application lies in the whole "pre-filtering and post-weighting" collaborative process, which enables the defense system to intelligently identify and suppress malicious behavior according to the actual data feedback of each round, so as to achieve stable and efficient defense effect in complex and variable environment.

[0093] It should be further pointed out that the above embodiments are only used to understand the technical solutions of the application, and are not used to limit the protection scope of the application. Any obvious adjustment and modification of the technical solutions of the application which belongs to the technical concept of the application should also belong to the protection scope of the application.

Claims

1. A federated learning backdoor defense method based on gradient selection and weight adjustment, characterized by: The solution comprises three steps: local client-side training, server-side collaborative processing and defense, and security aggregation and model update. Step 1: Local training on the client side At the start of each round of communication, the selected client first downloads the latest global model from the server. Then, each client trains the model using its local dataset. During this stage, the malicious client will implant a backdoor trigger in its local dataset to train a local model containing malicious intent. After training is completed, all selected clients will upload the updated model parameters or gradients to the server for aggregation. Step 2: Server-side collaborative processing and defense After receiving updates uploaded by the client, the server does not directly aggregate them. Instead, it performs collaborative defense operations on all collected model updates, including:

1. Gradient Adaptive Sparsity In order to retain key information while eliminating potential malicious updates, the server performs layered and adaptive sparsification processing on the gradients uploaded by each client.

2. Adaptive weight adjustment After sparsification, the server dynamically adjusts the aggregation weights of client updates based on their similarity. Step 3: Secure Aggregation and Model Update The server uses the gradients processed by the above steps and performs a weighted average with their respective adaptive weights to generate a new global model. Finally, the updated global model is distributed to the next round of clients participating in training, and this process is repeated until the model converges.

2. The federated learning backdoor defense method based on gradient filtering and weight adjustment according to claim 1, characterized in that: In the second step of server-side collaborative processing and defense, step 1, gradient adaptive sparsity, includes: 1) Dynamic threshold calculation: First, based on the training progress et and the total number of parameters N of network layer l, l Calculate a dynamically changing sparsification cardinality T l(t) In the early stages of training, this value is small, and stricter screening is implemented; as training progresses, this value gradually increases, allowing more parameter updates to pass through. 2) Layer Importance Weighting: Considering the different importance of different layers in a neural network (e.g., weight layers typically carry more learning capacity than bias layers), a layer importance coefficient I is introduced. l The weighted layers are assigned higher weight values, while the biased layers are assigned lower weights, thus protecting the more important layers during sparsification. 3) Outlier Detection: To filter out outlier gradient values ​​that deviate too much, this scheme calculates the mean and variance of the gradient at each layer and uses the Z-score method to identify outliers. Only gradients with a Z-score value less than a preset threshold ζ are considered normal, and an outlier detection mask z is generated. s ; 4) Generate the final mask: Combining dynamic thresholding, layer importance, and outlier detection, the server determines the final number of parameters k to be retained for each layer. l Then, the TopK function is used to retain only the k gradients with the largest absolute values ​​in each layer. l Each value, and the anomaly detection mask z. s Multiplying these yields the final sparsified gradient used for aggregation.

3. The federated learning backdoor defense method based on gradient filtering and weight adjustment according to claim 1, characterized in that: In the second step of server-side collaborative processing and defense, step 2, adaptive weight adjustment, includes: 1) Client Clustering: The HDBSCAN clustering method is used to group clients based on the cosine distance of the sparsified gradient. This effectively groups clients with similar behaviors (which may be benign or malignant) together, and assigns a cluster label b to each client i. i ; 2) Clipping rate calculation: Calculate the Euclidean distance between each client gradient, and derive a clipping rate ρ based on this distance. i This is used to control the magnitude of gradient trimming; 3) Calculate the adaptive weighting factor: To address the shortcomings of the traditional fixed-weight strategy, an adaptive weighting adjustment factor λ is introduced. i This factor is calculated based on the cosine similarity distribution of the gradients after sparsification of all clients, and can dynamically fine-tune the contribution of each client in the aggregation according to the data of the current round.

4. The federated learning backdoor defense method based on gradient selection and weight adjustment according to claim 3 is characterized in that: after accurately calculating the adaptive tuning parameter λ, a new generation of model parameters is generated by embedding all the obtained tuning parameters into the predetermined architecture of the aggregation function. To effectively mitigate the impact of outliers, noise is added to the model's weights. Where N(0,σ) 2 ) represents the Gaussian noise function added. The purpose of adding noise is to reduce the impact of backdoor attacks and enhance the defense effect.

Citation Information

Patent Citations

  • Backdoor attack defense method in federated learning based on multi-dimensional index dynamic identification

    CN116150745A

  • Poison defense method, device and equipment for federal learning, medium and product

    CN119520096A

  • Minkowski distance-based federal learning backdoor defense method

    CN120296727A