A federated learning backdoor attack defense method based on multi-layer collaborative defense strategy

By building a federated learning framework with multi-layer collaborative defense strategies, the problem of existing defense methods lacking systematicity and dynamicity in federated learning systems is solved, efficient defense throughout the life cycle is achieved, and the security and robustness of the system are improved.

CN120434054BActive Publication Date: 2025-08-26CHENGDU UNIV OF INFORMATION TECH

Patent Information

Application Number
CN202510933457.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-08-26
Estimated Expiration
2045-07-08

AI Technical Summary

Technical Problem

The existing federated learning backdoor attack defense methods lack systematic defense strategies throughout the entire life cycle, resulting in unsatisfactory defense effects and excessive computing and communication overhead, which makes it impossible to effectively deal with complex and changeable attacks.

Method used

Build a multi-layer defense framework throughout the entire life cycle of federated learning, including the first defense layer (local defense), the second defense layer (security aggregation of central servers) and the third defense layer (recovery after model deployment), and implement dynamic allocation of defense resources and adaptive adjustments through inter-layer information flow and global security state evaluation.

Benefits of technology

It improves the overall security and robustness of the federated learning system, effectively identify and filter updates of malicious nodes, quickly restore model performance, avoid high-cost retraining, and maximize the synergy of defense technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_1
    Figure QLYQS_1
  • Figure QLYQS_7
    Figure QLYQS_7
  • Figure QLYQS_12
    Figure QLYQS_12
Patent Text Reader

Abstract

This paper proposes a federated learning backdoor attack defense method based on a multi-layered collaborative defense strategy, belonging to the field of network data security technology. Addressing the shortcomings of existing federated learning technologies, which often only cover single-point protection at a specific stage, lack systematic and robust defense across the entire process, suffer from weak backdoor attack detection capabilities, and inefficient model recovery, this paper proposes a multi-layered defense framework that spans the entire federated learning cycle (training, aggregation, and deployment). By establishing specialized information flow channels, a global security assessment system, and an adaptive resource allocation strategy between each layer, this framework organically integrates the three defense layers into a coordinated whole. This collaborative mechanism not only improves the effectiveness of individual defense layers but also achieves a synergistic effect of "1+1+1 > 3," providing a comprehensive, intelligent backdoor attack defense solution for federated learning systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network data security technology, and in particular to a federated learning backdoor attack defense method based on a multi-layer collaborative defense strategy. Background Art

[0002] Federated learning is a distributed machine learning framework that allows multiple participants to jointly train machine learning models without sharing the original data. This technology enables different organizations or end devices to leverage their respective data strengths to collaboratively build high-quality machine learning models while protecting data privacy.

[0003] Backdoor attacks are a key security threat facing federated learning systems. Attackers tamper with training data or model updates, causing the model to behave normally under normal inputs but output erroneous results when encountering specific triggering conditions. The high degree of autonomy among participants in federated learning and the system's distributed architecture make backdoor attacks more covert and significantly more difficult to detect. In recent years, with the widespread application of federated learning, the issue of backdoor attacks has gradually garnered attention and become a hot topic of research.

[0004] In a federated learning system, backdoor attacks mainly occur in the following three stages: 1. Local training stage: Malicious participants may inject samples with triggers into local training data (data poisoning), causing the model to learn these abnormal patterns; 2. Model aggregation stage: Malicious participants may upload model parameters containing backdoors (model poisoning), affecting the global model; 3. Model deployment stage: Attackers may design adversarial samples to trigger the implanted backdoor, causing the model to produce incorrect output. At present, people have proposed a variety of defense methods against these attacks. For example, Sun et al. proposed a federated learning defense framework based on differential privacy in "Provable Defense against Privacy Leakage in Federated Learning from Representation Perspective" (ICML2021). This method adds noise when the client updates the gradient to prevent privacy leakage and backdoor attacks. Specifically, after calculating the gradient, the client adds random noise to the gradient according to a preset noise distribution (such as Gaussian distribution), reducing the possibility of attackers recovering the original data or implanting backdoors through gradient analysis; for example, Wang et al. proposed a defense method based on anomaly detection in "Detecting and Mitigating Backdoor Attacks in Federated Learning" (AAAI 2022). This method identifies abnormal updates that may contain backdoors by analyzing the statistical characteristics of model updates; for example, Li et al. proposed a defense method based on anomaly detection in "Defense Against Backdoor Attacks in Federated Learning via Knowledge Distillation" (IEEE Transactions on Information Forensics and Security, 2023) proposed a method to defend against backdoor attacks using knowledge distillation technology. This method introduces an auxiliary dataset and uses knowledge distillation technology to filter out the backdoor influence in the model. Specifically, its knowledge distillation-based method assumes that backdoor attacks mainly affect the model's behavior on specific trigger inputs, while having little effect on the processing of normal inputs. By performing knowledge distillation on a clean public dataset, the model's behavior on normal inputs is retained while mitigating the impact of backdoors.

[0005] While the aforementioned defense methods have improved the security of federated learning systems to a certain extent, they often focus only on specific stages of federated learning (such as client training, model aggregation, or model deployment), lacking a systematic defense strategy that spans the entire federated learning lifecycle. This fragmented defense approach is inadequate to address complex and diverse backdoor attacks, resulting in significant loopholes and blind spots. For example, the differentially private federated learning defense framework proposed by Sun et al. significantly degrades model performance when excessive noise is added, while insufficient defense effectiveness is insufficient when too little noise is added. Furthermore, it lacks adaptive mechanisms for different types of backdoor attacks, making it difficult to balance privacy protection and model utility, and fails to consider other defenses during the client training process. Another example is the anomaly detection defense method proposed by Wang et al., which uses a static anomaly detection threshold and struggles to adapt to dynamically changing attack strategies. It also struggles to distinguish between malicious updates and legitimate but anomalous updates caused by non-independent and identically distributed (Non-IID) data distributions. It also lacks a long-term evaluation mechanism for historical client behavior, making it impossible to establish an effective trust model. Consequently, it suffers from low detection accuracy for carefully crafted, relatively small backdoor attacks. Defense methods based on knowledge distillation also have obvious shortcomings, mainly reflected in: heavy reliance on the quality and quantity of representative public datasets, which are often difficult to obtain in practical applications; the distillation process itself may lead to degradation of model performance, especially in complex tasks; it fails to solve the problem of how to efficiently restore the attacked model after model deployment, and usually requires complete retraining; and there is a lack of accurate identification and repair mechanisms for specific attacked parts of the model.

[0006] In general, existing defense methods are mostly designed and applied independently, lacking a framework to organically integrate different defense technologies. This prevents them from fully leveraging their synergy. This results in suboptimal overall defense effectiveness even when multiple defense measures are employed, and incurs unnecessary computational and communication overhead. This is why the present invention was proposed. Summary of the Invention

[0007] In response to the shortcomings of the existing technology, the present invention proposes a federated learning backdoor attack defense method based on a multi-layer collaborative defense strategy, the purpose of which is to solve at least one of the above problems.

[0008] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0009] A federated learning backdoor attack defense method based on a multi-layer collaborative defense strategy, the method comprising:

[0010] Step S1: Construct the first, second, and third defense layers throughout the entire life cycle of federated learning. The first defense layer is deployed locally on each client to prevent backdoor implantation during the client training phase. The second and third defense layers are both deployed on the central server. The second defense layer is used to filter malicious updates during the model aggregation phase, and the third defense layer is used to achieve efficient recovery of attacked models during the model deployment phase.

[0011] Step S2: establishing inter-layer information flow between the first defense layer, the second defense layer, and the third defense layer, wherein the inter-layer information flow includes reputation information flow, anomaly detection information flow, and model recovery score information flow;

[0012] Step S3: Perform global security status assessment based on inter-layer information flow;

[0013] Step S4: Implement dynamic allocation of defense resources and adaptive adjustment of defense strategies based on global security status assessment.

[0014] Compared with the prior art, the present invention has at least the following beneficial effects:

[0015] In federated learning systems, existing defense methods are mostly designed and applied independently, lacking a framework that organically integrates different defense technologies. This prevents them from fully leveraging their synergistic effects. This results in suboptimal overall defense effectiveness, even with the use of multiple defenses, and introduces unnecessary computational and communication overhead. This paper proposes a multi-layered defense framework that spans the entire federated learning cycle (training, aggregation, and deployment). By systematically integrating multiple defense technologies, it constructs a comprehensive backdoor attack defense system. During the client training phase, a local defense mechanism combining gradient random perturbation and dynamic dropout is designed to enhance the model's resistance to data poisoning attacks. During the model aggregation phase, a secure aggregation strategy based on reputation assessment and outlier detection is developed to effectively identify and filter updates from malicious nodes. During the model deployment phase, an efficient recovery mechanism based on knowledge distillation and model forgetting is implemented to quickly restore model performance after an attack, avoiding the high cost of retraining. Through information exchange and collaborative optimization between defense layers, the overall security and robustness of the federated learning system are maximized while maintaining high model performance. DETAILED DESCRIPTION

[0016] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with examples. The illustrative embodiments of the present invention and their descriptions are only used to explain the present invention and are not intended to limit the present invention.

[0017] Specifically, the present invention provides a federated learning backdoor attack defense method based on a multi-layer collaborative defense strategy, the method comprising:

[0018] Step S1: Construct the first, second, and third defense layers throughout the entire life cycle of federated learning. The first defense layer is deployed locally on each client to prevent backdoor implantation during the client training phase. The second and third defense layers are both deployed on the central server. The second defense layer is used to filter malicious updates during the model aggregation phase, and the third defense layer is used to achieve efficient recovery of attacked models during the model deployment phase.

[0019] Step S2: establishing inter-layer information flow between the first defense layer, the second defense layer, and the third defense layer, wherein the inter-layer information flow includes reputation information flow, anomaly detection information flow, and model recovery score information flow;

[0020] Step S3: Perform global security status assessment based on inter-layer information flow;

[0021] Step S4: Implement dynamic allocation of defense resources and adaptive adjustment of defense strategies based on global security status assessment.

[0022] It should be noted that existing federated learning systems are vulnerable to backdoor attacks during the training, aggregation, and deployment stages. Since existing defenses often only cover single-point protection at a certain stage, they lack systematic and robust defense for the entire process, allowing attackers to penetrate even during uncovered stages. In addition, traditional defense methods mostly focus on a single stage and are unable to interact with information between layers. Consequently, their credibility assessments rely heavily on static historical data and are unable to adapt to behavioral changes at malicious nodes in real time, leading to lagging defenses. Furthermore, traditional defense methods often require retraining after the model is attacked, resulting in high computational costs and a lack of targeted repair methods. The present invention establishes a defense layer for each stage of the federated learning system, establishes inter-layer information flow between them, performs a global security status assessment through a specific information flow method, and implements dynamic allocation of defense resources and adaptive adjustment of defense strategies based on the global security status assessment, thereby achieving a comprehensive dynamic closed-loop defense and greatly improving the defense effect.

[0023] Specifically, to effectively prevent the risk of backdoor implantation during the client training phase, the first defense layer uses a comprehensive defense mechanism that combines gradient random perturbation with dynamic dropout rate adjustment (the probability of dropping neurons in each training iteration). It also dynamically adjusts the local defense strength based on the client's reputation information obtained in real time from the second defense layer. Gradient random perturbation uses an adaptive perturbation strength calculation method. Its core idea is to add random noise to the calculated gradient during model training, interfering with the attacker's ability to precisely control the gradient direction, thereby reducing the effectiveness of data poisoning attacks. The specific implementation steps are as follows:

[0024] Step S1.11 calculates the adaptive perturbation strength μ according to the current training round, model convergence status and global model update status t , the calculation formula is as follows:

[0025] μ t =β*μ base *(1-exp(-α*L t / L0)) * (1 - t / T) γ ,

[0026] Among them, μ base is the basic perturbation intensity, usually set to a decimal between 0.01 and 0.05; L t is the loss value of the current round, L0 is the initial loss value; t is the current training round, T is the total training round; α, β, γ are hyperparameters that control the change of perturbation intensity, usually α∈[0.5, 2], β∈[0.8, 1.2], γ∈[0.5,1]; where, the expression (1 - exp(-α*L t / L0)) reflects the state of model convergence, which is used to ensure that the perturbation is small in the early stage of training (when the loss is large) and gradually increases as the training progresses; the expression (1 - t / T) γ It reflects the global model update and is used to ensure that the disturbance gradually decreases to ensure convergence as the training nears the end;

[0027] Step S1.12 Based on the calculated adaptive perturbation strength μ t , select and generate noise distribution: dynamically select appropriate noise distribution and generate noise according to the defense requirements of different clients (N o ); where noise (N o ) is selected from Gaussian noise, Laplace noise or uniform noise. Gaussian noise is suitable for scenarios with small perturbations to the gradient, which satisfies: N o ~N(0,σ t ), where σ t is the noise standard deviation, which is related to the gradient (g r ) norm, σ t =μ t * ||g r || 2 Laplace noise uses local differential privacy technology, which is suitable for scenarios that require stronger privacy protection and meets the following requirements: N o ~Laplace(0, b t ), where b t is the noise scale parameter, , ε is the privacy budget; uniform noise is computationally efficient and suitable for resource-constrained devices, which satisfies: N o ~U(-at , a t ), where a t is the range parameter of the uniform distribution, .

[0028] Step S1.13 adds the generated noise to the calculated gradient, performs gradient update, obtains the perturbation-updated gradient, and updates the model parameters; its expression is as follows:

[0029] g rp = g r + N o ,

[0030] m g = m o -L e * g rp ,

[0031] Among them, g rp is the updated gradient, g r is the calculated gradient, N o is noise; m g is the updated model parameter, m o is the model parameter before updating, L e is the learning rate, which is used to control the step size of the model parameter update and determines the amplitude of the model parameter adjustment according to the gradient in each iteration.

[0032] Preferably, in order to balance the defense effect and model performance, during the implementation of the gradient random perturbation, the following steps are used to perform adaptive adjustment of the perturbation intensity:

[0033] 1) Monitor model performance, for example, calculating performance metrics (such as accuracy) on the validation set every k rounds;

[0034] 2) Determine the degree of performance degradation of the model based on the calculation of performance indicators; taking accuracy as an example, the following formula is used for calculation:

[0035] Δa cc =a ccp -a ccv ,

[0036] Where Δa cc is the change in accuracy, a ccp is the accuracy of the validation set in the previous round, a ccv is the accuracy of the validation set of the current round;

[0037] 3) Adaptive adjustment of disturbance intensity: If Δa cc >a cct (a cctis the accuracy change threshold, which can be used to express that if the performance drops by more than the threshold, the perturbation intensity is reduced, for example, μ baseg =μ base *0.8; if Δa cc <0 (performance improvement), then appropriately increase the perturbation intensity: μ baseg =min(μ base *1.05,μmax), μmax is the maximum permissible disturbance intensity, μ baseg is the adjusted disturbance intensity.

[0038] Furthermore, dynamic dropout rate adjustment uses an adjustment algorithm based on sensitivity analysis, which intelligently adjusts the dropout ratio of each layer in the neural network according to the training stage and model state, increasing the randomness of model training and improving the model's robustness to data poisoning. The specific implementation steps are as follows:

[0039] Step S1.21 Initialize layer dropout rate

[0040] According to the network structure, set the initial Dropout rate for each layer:

[0041] d r = [p1, p2, ..., p L ],

[0042] Where L is the number of network layers, p L is the initial dropout rate for layer L, indicating the probability that neurons in layer L will be randomly dropped during training. It is typically set between 0.1 and 0.5, depending on the depth and complexity of the network. Deep networks typically use higher dropout rates to enhance the model's generalization and resistance to overfitting.

[0043] Step S1.22: Periodically perform sensitivity analysis to determine the sensitivity of each layer to backdoor attacks, which includes the following steps:

[0044] S1.221. For each layer i, temporarily increase its dropout rate, calculated as follows:

[0045] pii'=pii+Δp,

[0046] Where Δp is the temporarily increased dropout rate, usually set to 0.1-0.2, pii' is the dropout rate of the i-th layer after the temporary adjustment, and pii is the dropout rate of the i-th layer before the temporary adjustment. The purpose of this step is to observe the changes in model performance by increasing the dropout rate, thereby evaluating the sensitivity of the layer to backdoor attacks.

[0047] S1.222. Use a small number of validation samples containing potential triggers (such as anomalous or synthetic samples) for model testing;

[0048] S1.223. Calculate performance changes using the following formula:

[0049] Δperf=perfo - perfi,

[0050] Among them, Δperf is the performance change, perfo represents the performance of the model under the original Dropout rate, and perfi represents the performance of the model after temporarily increasing the Dropout rate. The performance indicators that characterize the performance change can be accuracy, recall rate, F1 score, etc.

[0051] S1.224. Calculate the sensitivity score of layer i using the following formula:

[0052] Si = Δperf / Δp,

[0053] Among them, the sensitivity score Si indicates the sensitivity of the layer to changes in the Dropout rate. The higher the sensitivity, the more sensitive the layer is to backdoor attacks and the more likely it is to contain backdoor features.

[0054] Step S1.23: Dynamically adjust the Dropout rate

[0055] According to the sensitivity analysis results, the Dropout rate of each layer is dynamically adjusted:

[0056] pi new = pi +η* Si * Snormalize,

[0057] Among them, η is the adjustment step size, usually set to 0.01~0.05; Snormalize is the normalized sensitivity: Snormalize = Si / max(S1, S2, ..., SL); at the same time, set the upper and lower limit constraints: p min ≤ pi new ≤ p max , usually p min =0.05, p max =0.7, pi is the Dropout rate before dynamic adjustment of the i-th layer, pi new is the dynamically adjusted Dropout rate of the i-th layer.

[0058] Preferably, in the implementation process of dynamic Dropout rate adjustment, in order to prevent the Dropout rate from continuously increasing and causing underfitting, a periodic reset mechanism is designed, including: checking the performance of the model on the clean validation set every N training rounds (such as N=10); if the performance drops for K consecutive rounds (such as K=3), resetting the Dropout rate of all layers to 80% of the initial value; and then restarting the dynamic adjustment process.

[0059] Furthermore, to dynamically adjust the local defense strength, the following steps are used to establish a comprehensive defense mechanism that combines gradient random perturbation with dynamic Dropout rate adjustment:

[0060] First, in actual deployment, a dynamic balance defense strength equation is established based on the security requirements of the model training phase:

[0061] d s = ww1*p ti + ww2*p avg ,

[0062] Among them, d s is the defense strength score, ww1 and ww2 are weight coefficients, p ti is the current gradient perturbation strength, p avg is the average Dropout rate;

[0063] Next, using the client reputation information obtained from the model aggregation layer, the local defense strength is adjusted in the following way: if the current client reputation is high, the defense strength is appropriately reduced to improve model performance; if the current client reputation is low or fluctuates greatly, the defense strength is increased to strengthen security. The adjustment formula is:

[0064] μ t_adjusted = μ t * (2 - r score ),

[0065] d rs = 1 + λ * (1 - r score ),

[0066] Among them, μ t_adjusted is the adjusted gradient perturbation strength, r score ∈[0,1], which is the reputation score fed back from the second defense layer, λ is the adjustment coefficient, d rs is the adjusted Dropout rate;

[0067] Finally, the adjusted μ t_adjusted and d rs Feedback to the dynamic equilibrium defense strength equation, update p ti and p avg, thereby dynamically maintaining d s When the target defense strength is near, ensure that the local defense layer optimizes the defense strategy in real time based on security needs and reputation status.

[0068] Next, the second layer of defense is introduced in detail:

[0069] The second defense layer primarily addresses model poisoning attacks during the federated learning aggregation phase. It uses reputation assessment and outlier detection methods to identify and filter out model updates submitted by malicious clients. Reputation assessment analyzes the client's historical behavior and model update quality to establish a dynamically updated reputation score for each client, which is used to guide the model aggregation process. Specifically, it includes the following steps:

[0070] Step S2.11: Multi-dimensional reputation feature extraction

[0071] A multi-dimensional reputation feature vector fc is extracted for each client c, wherein the multi-dimensional reputation feature vector fc includes a consistency feature fconsist, a stability feature fstable, a contribution feature fcontrib, and an abnormality feature fanomaly, wherein:

[0072] Consistency feature f consist It is used to characterize the consistency between the client model update and the global update direction, which is calculated using the following formula:

[0073] f consist =cosine_similarity(Δw c , Δw global ),

[0074] Among them, cosine_similarity refers to cosine similarity, Δw c is the model update of client c, Δw global is the global model update of the previous round;

[0075] Stability characteristic f stable The stability of the client model is analyzed by calculating the cosine similarity between consecutive rounds of updates, which is calculated using the following formula:

[0076] f stable =1-std([cosine_similarity(Δw ct , Δw ct-1 ), ..., cosine_similarity(Δw ct-k+1 , Δw ct-k )]), where std represents the standard deviation of the given cosine similarity list, Δw ctrepresents the model update of client c in round t, and k is the number of historical rounds considered.

[0077] Contribution feature f contrib It is used to characterize the contribution of client model updates to the global model performance and is calculated using the following formula:

[0078] f contrib =perf(Δw global + Δw c ) - perf(Δw global ),

[0079] Among them, perf() represents the performance evaluation function on the validation set;

[0080] Abnormality feature f anomaly It is used to characterize the abnormality of the client model update relative to other clients. It is calculated using the following formula:

[0081] f anomaly =-z_score(dist(Δwc, centroid({Δw i | i∈clients}))),

[0082] Among them, dist() is the distance function, centroid() calculates the center point of all client updates, and z_score() is the normalization function;

[0083] Step S2.12: Credit score calculation

[0084] According to the extracted multi-dimensional reputation feature vector fc, the comprehensive reputation score Rc of client c is calculated:

[0085] ,

[0086] Where: f i is the i-th reputation feature, corresponding to the above consistency feature f consist , stability characteristics f stable , contribution feature f contrib and abnormality feature f anomaly There are four features in total; w i is the feature weight, which is initially set to equal weight; the sigmoid function maps the score to the (0,1) interval;

[0087] Step S2.13: Exponential smoothing update

[0088] In order to take into account the client's historical behavior, the exponential smoothing method is used to update the reputation score. The calculation formula is as follows:

[0089] Rct=δ* Rct+(1-δ)*Rct-1,

[0090] Where Rct is the updated reputation score in round t; Rct-1 is the reputation score in round t-1; δ is the smoothing coefficient, which is usually set to 0.2-0.4 to balance current performance and historical behavior.

[0091] Step 2.14: Take appropriate measures based on the updated reputation score:

[0092] For high-reputation clients (Rc>highest reputation threshold, such as 0.8): fully accept their model updates, and preferably increase their weight in the aggregation; for medium-reputation clients (lowest reputation threshold ≤ Rc ≤highest reputation threshold, such as 0.5≤Rc≤0.8): accept their model updates, but use standard weights; for low-reputation clients (Rc<lowest reputation threshold, such as 0.5): reduce their weight in the aggregation, or completely exclude their updates; and use the following formula to adjust the aggregation weight:

[0093] wc = Rc β / ∑i ∈ clients Ri β ,

[0094] Among them, β is an exponent that amplifies the reputation difference and is usually set to 2 to 3.

[0095] Furthermore, the outlier detection method analyzes the model updates submitted by the client to identify abnormal model updates in real time, preventing malicious updates from contaminating the global model. The specific implementation steps are as follows:

[0096] Step S2.21: Model update feature extraction

[0097] Model update feature extraction is a key step in identifying outliers (abnormal updates). By analyzing model weight changes from multiple dimensions, we can effectively identify potential malicious or abnormal clients. Model update feature extraction includes the following steps:

[0098] S2.211. Perform principal component analysis (PCA) dimensionality reduction on the high-dimensional model update data:

[0099] (a) Represent the client’s model update Δwc as a high-dimensional vector;

[0100] (b) Standardize and preprocess the updated data to ensure fair comparison of different parameter magnitudes;

[0101] (c) Calculate the covariance matrix and analyze the correlation between parameters;

[0102] (d) Solve the eigenvalues ​​and eigenvectors and sort them by eigenvalue size;

[0103] (e) Select the minimum dimension d that can explain 95% of the cumulative variance;

[0104] (f) Update the original projection using the following formula:

[0105] Δw c_reduced =PCA(Δw c , n components =d),

[0106] Where Δw c represents the model update submitted by client c, which is a vector containing the changes in model parameters in the current training round; Δw c_reduced is the model update vector of client c after PCA dimensionality reduction. The dimension of this vector is reduced to d, but the main features and change information of the original data are still retained. components =d represents parameter setting, which means setting the dimension reduction dimension to the minimum dimension d;

[0107] S2.212. Perform multi-dimensional feature extraction, including layer-based feature grouping, statistical moment feature extraction, and relative change feature extraction. Model parameters at different layers have unique update characteristics and sensitivities, allowing for more detailed anomaly patterns to be captured through layered analysis. Statistical moments provide a complete description of update distributions and can effectively identify update patterns with abnormal distributions. Relative change feature extraction, by comparing update differences between consecutive rounds, can detect sudden changes in client behavior, which are often important indicators of attacks or anomalies.

[0108] S2.213, feature merging forms a complete feature vector, which is then used to capture abnormal patterns in model updates. The various features extracted in S2.212 are merged, and the calculation formula is as follows:

[0109] F features =[Δw c_reduced ,F layer1 , ..., F layerN , F stat , F rel ],

[0110] Among them, F features Represents the feature vector after combining various features, F layerN Represents the extracted layer-based feature vector, N represents the number of layers, F stat is the statistical moment eigenvector, F rel is the relative change eigenvector;

[0111] The resulting feature vectors, derived from the merging of these various features, are used to identify outliers in input anomaly detection algorithms (such as Isolation Forest and One-Class SVM), serve as the basis for client-side reputation scoring systems, provide a reference for dynamic aggregation weighting mechanisms, and inform decision-making for defense mechanisms. This comprehensive feature extraction framework effectively captures anomalous patterns in model updates, improving the robustness of federated learning systems.

[0112] Step S2.22: Based on the feature vectors features obtained by the model update feature extraction, multi-model outlier detection is performed to obtain an integrated detection result. The multi-model outlier detection includes local outlier factor (LOF) detection, isolation forest detection, and DBSCAN-based detection. The integrated detection result is expressed as follows:

[0113] S en = w lof *S lof + w if *S if + w dbscan * S dbscan ,

[0114] Among them, w lof 、w if and w dbscan The weights corresponding to local outlier factor detection, isolation forest detection, and DBSCAN-based detection can be tuned through the validation set. lof 、S if 、S dbscan is the corresponding test result score, S en is the final score of the integrated detection results;

[0115] Step S2.23, performing adaptive threshold setting and determining an outlier processing strategy based on the final score of the integrated detection result, which includes the following steps:

[0116] (a) Setting an initial threshold factor (factor) base value, usually between 2 and 3, as the starting point for subsequent dynamic adjustments. This base value is set based on experience and the initial requirements of the model;

[0117] (b) In each training round, the actual proportion of detected outliers is calculated based on the current threshold ( ratio ), calculated as follows:

[0118] O ratio =N outliers / C total ,

[0119] Among them, N outliersrepresents the number of clients detected as outliers, C total Indicates the total number of clients participating in the current training round;

[0120] (c) Dynamically adjust the threshold factor (S factor ): Calculate the actual outlier ratio (O ratio ) and the target outlier ratio (T ratio ) deviation, and dynamically adjust the threshold factor (S factor ):

[0121] If O ratio >T ratio +Y margin , increase the threshold value of the threshold factor according to the following formula:

[0122] S factor后 =S factor前 *1.05,

[0123] If O ratio <T ratio -Y margin , reduce the threshold value of the threshold factor as follows:

[0124] S factor后 =S factor前 * 0.95,

[0125] Among them, T ratio is the target outlier ratio, usually set to 0.05-0.1, S factor前 is the threshold factor before dynamic adjustment, S factor后 is the threshold factor after dynamic adjustment, Y margin is the allowable deviation range;

[0126] (d) Using the dynamically adjusted threshold factor (S factor后 ), combined with the mean and standard deviation of the current score, calculate a new dynamic threshold (threshold, abbreviated as Ts). The dynamic threshold is a dynamic threshold based on statistics or a quantile threshold based on historical data. The dynamic threshold based on statistics is calculated using the following formula:

[0127] Ts= mean(S en ) + S factor后 * std(S en ),

[0128] Among them, mean(S en ) represents the mean of the integrated detection score, that is, the average of all client outlier detection scores, reflecting the central trend of the data; std(S en ) represents the standard deviation of the integrated detection score, which measures the degree of dispersion of the data;

[0129] The quantile threshold based on historical data is calculated using the following formula:

[0130] Ts= percentile(h scores , p)

[0131] Among them, percentile is an indicator to measure the data distribution position, h scores is a dataset of historical detection scores, and p is the percentile, usually set to 95 or 97.5, indicating that only a very small number of data points will be considered outliers;

[0132] (e) Based on the calculated new dynamic threshold (Ts), calculate the outlier degree (D outlier ) and combined with the client's reputation, adopt a differentiated processing strategy, where the outlier degree is calculated using the following formula:

[0133] D outlier = (S en - Ts) / Ts

[0134] Where S en is the final score of the integrated detection results, Ts is the calculated new dynamic threshold;

[0135] The differentiated processing strategy includes: for high outlier and low credibility, completely exclude the client from updating; for high outlier but medium credibility, significantly reduce the weight; for medium outlier, linearly adjust the weight according to the credibility; for normal range, use standard weight; the weight adjustment formula is as follows:

[0136] W adjusted =W original * max(0, 1 - D outlier * (1 - r score ))

[0137] Where W adjusted Represents the adjusted weight, W original represents the original weight, that is, the initial weight of the client update before considering the degree of outliers and reputation, r score is the reputation score, obtained from the feedback of the second defense layer; this formula ensures that high-reputation clients will receive a small penalty even if they are detected as slight outliers, while low-reputation clients will be severely punished for outlier behavior.

[0138] Furthermore, to effectively protect the federated learning system from attacks by malicious clients, the second defense layer uses the following steps to establish a collaborative defense mechanism for reputation evaluation and outlier detection:

[0139] Step S2.31: Information sharing and comprehensive decision-making

[0140] In this step, the reputation assessment system and the outlier detection system conduct two-way information exchange to form a complementary security assessment mechanism: 1) Outlier detection feedback to reputation assessment: The outlier detection system calculates the degree of abnormality of the model update submitted by each client, and then feeds this information back to the reputation assessment system: The system first calculates the anomaly score (outlier_score) of the client model update, normalizes the anomaly score (normalized_outlier_score), and converts the normalized anomaly score into the input feature of the reputation assessment. The expression is: fanomaly = -normalized_outlier_score, where the negative sign indicates that the higher the degree of anomaly, the more negative the contribution to the reputation. This design allows the outlier detection result to directly affect the overall reputation assessment of the client; 2) Reputation affects the sensitivity of outlier detection: At the same time, the reputation assessment system also affects the judgment criteria of outlier detection, and adopts differentiated detection strategies for clients with different reputations. The expression is: outlier_threshold_adjusted = outlier_threshold * (2- When a client's reputation score is high (close to 1), the adjusted threshold is closer to the original threshold. When the client's reputation score is low (close to 0), the adjusted threshold is closer to twice the original threshold, adopting a stricter standard. This adaptive mechanism ensures that the system conducts stricter scrutiny of model updates for low-reputation clients, improving defense capabilities.

[0141] Step S2.32: Based on the comprehensive results of reputation evaluation and outlier detection, the system implements a multi-level security aggregation strategy, including weighted average aggregation, adaptive clipping aggregation, and hierarchical aggregation strategy. Specifically,

[0142] For weighted average aggregation, the system no longer simply averages the model updates of all clients, but assigns different weights based on the client's reputation and outlier level:

[0143] Δw global =∑c∈clients (α c *Δw c ) / ∑c∈clients α c ,

[0144] Among them, clients is the set of all clients participating in federated learning, c∈clients represents traversal of all participating clients (devices or data shards), Δw cThe local model update of client c, αc is the aggregation weight of client c, which is determined by the reputation and outlier degree:

[0145] α c = Rc * (1 - D outlierc ) y ,

[0146] Where Rc is the reputation score of client c (between 0 and 1), D outlierc represents the degree of outlier performance of client c's model update (between 0 and 1), and y is an exponent that amplifies the outlier penalty, typically set to 2 to strengthen the penalty for outlier behavior. This design ensures that clients with high credibility and normal model updates receive higher aggregate weights, while clients with low credibility or abnormal model updates have their weights significantly reduced.

[0147] For adaptive clipping aggregation, to further defend against extreme updates, the system adaptively clips the model updates of each client:

[0148] Δw c_clipped = clip(Δw c ,-τ c ,τ c )

[0149] Where Δw c_clipped The model update vector for the clipped client c, clip(Δw c , -τ c , τ c ) means limiting each element of the model update to the interval [-τ c ,τ c ] to ensure that the model update does not exceed a reasonable range and reduce the impact of abnormal updates; the clipping threshold τc is proportional to the credibility, and its calculation formula is τ c = τ base * (0.5+0.5 * Rc), τ base is the basic clipping threshold, which is used to set the basic scale of the clipping range. This means that high-reputation clients (Rc close to 1) get close to τ base The clipping threshold allows a larger range of updates, and low-reputation clients (Rc close to 0) gain about 0.5×τ base The pruning threshold is set to limit the update range. Through this pruning mechanism, the system can effectively control the impact of extreme updates submitted by malicious clients.

[0150] For the hierarchical aggregation strategy, the system divides clients into different levels based on their reputation and outlier degree, and implements differentiated aggregation strategies, including:

[0151] I) High-Reputation, No-Outlier Clients: These clients are considered the most trustworthy participants; their model updates are fully integrated into the global aggregation without any restrictions, and they receive additional weight boosts, strengthening their influence in the global model.

[0152] II) High-reputation slightly outlier clients: These clients are generally trustworthy, but their model updates are slightly abnormal. Their updates are moderately pruned before being included in the aggregation. The pruned updates are relatively light, preserving most of the original update information.

[0153] III) Moderately reputable or moderately outlier clients: These clients have a certain degree of uncertainty; their impact can be limited by significantly reducing their aggregate weights; stricter pruning measures can also be applied.

[0154] IV) Clients with low reputation or serious outliers: These clients are considered potential malicious actors and are completely excluded from the aggregation process. The system may also record the behavior patterns of these clients for subsequent security analysis.

[0155] This hierarchical aggregation strategy enables refined management of different types of clients, minimizing the impact of malicious attacks while ensuring model performance.

[0156] Next, the third defense layer is introduced in detail:

[0157] The third defense layer primarily addresses the issue of recovering from successful attacks after a federated learning model is deployed. By combining knowledge distillation with model forgetting, the performance of the attacked model can be efficiently restored, avoiding the high cost of complete retraining. The knowledge distillation method transfers knowledge from an unattacked model (the teacher model) to a potentially vulnerable model (the student model). This method helps the model retain its original performance while restoring normal functionality. The specific implementation steps are as follows:

[0158] Step S3.1: Teacher model preparation

[0159] The selection and preparation of the teacher model is a key prerequisite for the success of knowledge distillation. In this step, it is necessary to prepare a teacher model with reliable performance and not affected by backdoor attacks. Only by selecting and preparing a teacher model with reliable performance and not affected by backdoor attacks can we ensure that the knowledge it outputs is accurate and secure, and provide effective guidance for the learning and recovery of the student model. Specifically, the selection of the teacher model includes using the early healthy versions saved during the model development process as a source of knowledge, and selecting the best unattacked version from the model version library based on the historical performance records of the model (such as accuracy, F1 score, etc.); and / or, the selection of the teacher model also includes introducing pre-trained domain expert models as a supplementary source of knowledge transfer. These expert models are usually pre-trained models trained on large-scale data and have strong performance in specific fields. These expert models are combined with historical models to form a more powerful set of teacher models, which are used as a source of knowledge for teacher model selection and preparation. The expression is:

[0160] t en = [H model , E model1 , E model2 , …],

[0161] Among them, t en Represents the teacher model set, H model represents the historical model, E model1 Represents the first expert model, E model2 represents the second expert model, and so on.

[0162] Step S3.2: Distillation data preparation

[0163] Construct a distilled dataset that contains sufficient domain knowledge and does not contain backdoor triggers, for example, through proxy dataset construction, data augmentation and screening, adversarial sample generation, etc.

[0164] Step S3.3: Distillation loss design

[0165] This step uses feature distillation loss to capture the teacher model's knowledge, achieving comprehensive knowledge transfer. In this embodiment, feature distillation loss transfers deep-level feature extraction knowledge by minimizing the differences in intermediate-layer feature representations between the teacher and student models. This approach is particularly effective because intermediate-layer features embody the model's understanding of the input data structure and are often richer than the final output. In practice, key layers (such as the output layer of a residual block or an attention layer) are typically selected for feature extraction, and the L2 norm is used to quantify the differences between feature vectors. Given that feature dimensions and scales may vary between layers, the following approaches are often adopted in practice: using adaptation layers (such as 1×1 convolutions or fully connected layers) to map student features to the same dimensions as the teacher; normalizing features to reduce the impact of scale differences; and using an attention mechanism to highlight important feature regions. Through these methods, feature distillation loss effectively enables the student model to acquire feature extraction capabilities similar to those of the teacher model, especially in deep networks.

[0166] Step S3.4: Progressive Distillation Training

[0167] This step uses a phased strategy to gradually optimize different parts of the model, improving the efficiency of knowledge transfer and damaged model recovery. Specifically, it uses a freeze-thaw strategy, learning rate scheduling, and distillation temperature scheduling to achieve gradual optimization of different parts of the model. Among them, learning rate scheduling is an important strategy for controlling the parameter update step size. Reasonable learning rate changes can accelerate convergence and improve final performance. Distillation temperature scheduling is used to dynamically adjust the degree of soft label smoothing as training progresses. The freeze-thaw strategy is implemented in three stages:

[0168] Phase 1: Freeze most layers and train only the last few. This phase keeps the model's underlying and mid-layer parameters unchanged, updating only the last k layers (typically the classification head or task-specific layers). This allows the model to quickly adapt to the teacher model's decision logic while retaining much of its original feature extraction capabilities. Training in this phase is short (epochs = e1), primarily aiming to adjust the output layer to match the teacher model's predicted distribution.

[0169] Phase 2: Progressively unfreeze more layers. This phase unfreezes more layers (the final k+m layers), allowing mid-level features to be trained by the teacher. This phase uses a medium-length training cycle (epochs = e2), allowing the model to undergo deeper adjustments while still keeping the underlying feature extractor relatively stable.

[0170] Phase 3: Fully unfreeze and fine-tune the entire network. The final phase unfreezes all layers and fine-tunes the entire network end-to-end. This phase uses the longest training epoch (epochs = e3), but typically employs a smaller learning rate, carefully adjusting the entire network to achieve an optimal balance.

[0171] Furthermore, the model forgetting method is used to remove malicious behavior from the neural network model attacked by the backdoor while retaining its performance on normal tasks. This is achieved through the following steps:

[0172] Step S4.1: Backdoor feature location

[0173] The goal of this stage is to precisely identify which parts of the model are associated with the backdoor attack. This is the foundation of the entire repair process, and multiple analysis methods are required to ensure accurate positioning. These analysis methods include activation pattern analysis, gradient-based feature attribution, and feature importance ranking. Specifically, activation pattern analysis determines whether the corresponding layer is involved in backdoor behavior processing by analyzing the differences in activation patterns when the model processes normal data and when it carries trigger data. It first inputs a normal sample into the model, collects the activation values ​​of each layer of the model, and records them as clean activation values. Then, it inputs a sample with a trigger into the same model, collects the activation values ​​of each layer, and records them as triggered activation values. Next, it calculates the difference between the normal activation value and the triggered activation value of each layer. This difference is typically quantified by calculating the mean absolute difference of the activation values. Finally, the layers are ranked according to the size of the difference, and the layers with the most significant differences are selected. These layers are likely involved in backdoor behavior processing. For gradient-based feature attribution, gradients are calculated to determine which neurons contribute most to the backdoor behavior. For feature importance ranking, individual neurons are masked to assess their importance to the backdoor behavior.

[0174] Step S4.2: Selective parameter reset

[0175] After identifying the features associated with the backdoor, this step selectively modifies the parameters of these features to eliminate the backdoor's influence. This selective modification of the parameters of these features includes neuron pruning, weight resetting, and selective fine-tuning. Taking neuron pruning as an example, it prunes the identified backdoor neurons, creating a pruning mask for each neuron identified as backdoor-related. These masks are applied to "mute" the backdoor neurons during the forward propagation process, thereby directly cutting off the backdoor trigger path.

[0176] Step S4.3: Forget the features and behaviors related to the backdoor attack. Use algorithms such as gradient interference forgetting, adversarial forgetting, and memory replay and forgetting to gradually forget them. In other words, this forgetting method gradually weakens the impact of the backdoor rather than completely removing it all at once, which helps maintain the stability of the model.

[0177] It should be noted that the third defense layer achieves efficient recovery of the attacked model through knowledge distillation combined with model forgetting. The two methods complement each other to form a more powerful defense capability. When using them, to ensure the synergy between the two methods, the two methods can be applied in a fixed order or alternately through multiple rounds of iteration. Preferably, the objective functions of the two methods are combined into a unified optimization objective for hybrid optimization, and the hybrid loss function is constructed as follows:

[0178] L mixed =γ1 * L distill +γ2 * L unlearn ,

[0179] Among them, L mixed is the total loss function, γ1 and γ2 are the corresponding two weight coefficients used to adjust the distillation loss (L distill ) and forgetting loss (L unlearn ) in the total loss function; the main advantage of hybrid optimization is high training efficiency, and the model considers two objectives simultaneously in the same optimization process.

[0180] To better achieve model recovery, the third defense layer adopts an adaptive recovery strategy. That is, it dynamically adjusts the recovery method based on the specific situation of the attack and the progress of recovery. It includes the following steps:

[0181] 1) Conduct an attack severity assessment, generate a quantitative metric representing the severity of the attack, and output it to guide the selection of subsequent recovery strategies. This assessment includes comparing the model's current performance with expected performance, analyzing the degree of abnormal behavior during backdoor triggering testing, and assessing the degree to which the model's internal representation deviates from the expected model.

[0182] 2) Select an appropriate recovery strategy based on the assessed severity of the attack. The recovery strategies include: for severe attacks, that is, when the attack severity is higher than the high threshold, focus on using knowledge distillation methods, with the teacher model playing a dominant role and a higher teacher ratio (such as 0.8), relying on external knowledge sources to rebuild the model functions. In this case, the model may have been severely damaged and requires substantial reconstruction; for moderate attacks, that is, when the attack severity is between the medium threshold and the high threshold, balance the use of knowledge distillation and model forgetting, assigning similar weights (such as a distillation ratio of 0.5 and a forgetting ratio of 0.5). In this case, the backdoor has a significant impact, but the model still retains a large amount of useful knowledge; for minor attacks, that is, when the attack severity is lower than the medium threshold, focus on using model forgetting technology, using precise positioning methods, focusing on accurately identifying and removing backdoor features, and minimizing interference with other functions. In this case, most of the model functions are intact, and only the backdoor needs to be accurately removed. Choosing different strategy combinations based on the severity of the attack can avoid "one-size-fits-all" repair solutions and improve recovery efficiency and effectiveness.

[0183] 3) During the recovery process, dynamically adjust recovery parameters based on real-time monitoring of recovery progress. Specifically, regularly evaluate model recovery using clean benchmark data, monitor key metrics such as the accuracy of the main task and the success rate of backdoor attacks, and compare recovery progress with expected targets to determine whether the expected results have been achieved. If recovery progress falls short of expectations, increase the learning rate, increase the weight coefficient, extend the training time for specific stages, or adjust the ratio of the technology combination to focus on the more effective technology. If recovery progress meets or exceeds expectations, reduce the learning rate for fine-tuning, balance the weights of the two technologies, prevent overcorrection, and enhance the model's generalization and stability. This dynamic adjustment mechanism makes the recovery process self-correcting, allowing it to adjust its direction and intensity based on actual results, thereby improving the success rate of recovery.

[0184] To better achieve the objectives of the present invention, in step S2, the flow of reputation information refers to the flow of reputation information from the second defense layer to other defense layers. Specifically, reputation information is obtained by the second defense layer through global calculations. After the second defense layer calculates the reputation score of each client, the reputation information is transmitted to the first defense layer. Based on the received reputation information, the first defense layer adjusts the strength of local defenses for different clients. Clients with lower reputations face stricter detection and filtering measures, while clients with good reputations enjoy a relatively relaxed defense policy. This differentiated treatment improves defense efficiency and reduces unnecessary interference with normal clients. In addition, the reputation information calculated by the second defense layer also flows to the third defense layer to guide the selection of teacher models during the model recovery process. Based on the received reputation scores, the third defense layer selects models with reputations above a threshold from the teacher model set as teacher models. This selection of teacher models based on reputation scores ensures the reliability of the teacher models and prevents contaminated models from participating in the recovery process.

[0185] Furthermore, in step S2, anomaly detection information is obtained by the first defense layer by performing anomaly detection locally on the client, which includes the identified possible malicious behavior or data anomalies. After the anomaly detection information is identified locally on the client, it is passed to the second defense layer. After receiving this information, the second defense layer integrates it into the global outlier detection, thereby improving the accuracy and efficiency of outlier client identification. This local to global information flow enables the system to simultaneously utilize the advantages of local fine-grained detection and global macro analysis. Furthermore, the second defense layer can identify potential backdoor areas, that is, parameter areas in the model where backdoors may be implanted, through outlier analysis and aggregation strategies. This key information will be passed to the third defense layer. After receiving it, the third defense layer will mark these areas as priority processing objects and focus on and process these areas in the subsequent model recovery process. This directional information flow greatly improves the efficiency and accuracy of model recovery.

[0186] Furthermore, in step S2, model recovery score information is generated by the third defense layer. This third defense layer evaluates the effectiveness of model recovery, generates a recovery score, and feeds this score back to the first defense layer. Based on this feedback information, the first defense layer dynamically adjusts its local defense strategy, including adjusting anomaly detection thresholds and updating filtering rules. This feedback mechanism enables the first defense layer to continuously optimize its strategy based on overall defense effectiveness. Furthermore, the third defense layer's recovery effectiveness evaluation is also fed back to the second defense layer, which uses this information to adjust aggregation parameters, such as adjusting outlier detection sensitivity and modifying the aggregation weight calculation method. Through this feedback mechanism, the second defense layer can continuously optimize its aggregation strategy and improve its resistance to backdoor attacks.

[0187] In order to better achieve the purpose of the present invention, in step S3, the global security status evaluation is calculated by the following formula:

[0188] ,

[0189] Where, represents the global safety score, Score the client locally. Score the aggregation phase, Score the model recovery, is the corresponding weight coefficient, n represents the total number of clients, where

[0190] ,

[0191] Where, 、 are the mean and variance of all client updates in this round, is the update vector of client i at time step t, is a small constant used for numerical stability to avoid the denominator being zero;

[0192] ,

[0193] Where, represents the maximum value of cosine similarity among all possible client pairs (i, j), is the update vector of client j at time step t;

[0194] ,

[0195] Where Acc(w t ) represents the model weight w at time step t t The accuracy under , indicates the performance of the current round model.

[0196] In order to better achieve the purpose of the present invention, in step S4, the system implements dynamic allocation of defense resources based on the global security score to cope with different security conditions. Specifically,

[0197] 1) High-risk status: When the security score falls below the danger threshold, the system determines that it faces a serious threat and will increase resource investment in all defense layers. Specific measures include increasing local detection frequency, increasing outlier detection sensitivity, and strengthening model recovery efforts, all to enhance defense strength.

[0198] 2) Warning Status: When the security score is between the warning threshold and the danger threshold, the system analyzes the performance of each defense layer, identifies weak links, and strengthens them in a targeted manner. This precise strengthening strategy can achieve the best defense effect under limited resource conditions;

[0199] 3) Security Status: When the security score is higher than the warning threshold, the system is in a relatively safe state. At this time, defense resource allocation will be optimized to balance security and system efficiency. For example, the detection frequency will be appropriately reduced, the complexity of the aggregation algorithm will be adjusted, and so on. This will improve system performance while ensuring safety.

[0200] Furthermore, in step S4, adaptive adjustment of the defense strategy refers to dynamically adjusting the defense focus based on the observed attack pattern. Specifically, when the system detects that the primary attack pattern is data poisoning, it prioritizes strengthening the first defense layer, including measures such as increasing the sensitivity of local anomaly detection, increasing data cleaning efforts, and strengthening feature filtering. When the system detects that the primary attack pattern is model poisoning, it prioritizes strengthening the second defense layer, including measures such as optimizing the outlier detection algorithm, adjusting the aggregation weight strategy, and enhancing model consistency verification. When the system detects that the primary attack pattern is post-deployment attack, it prioritizes strengthening the third defense layer, including measures such as increasing the frequency of model monitoring, optimizing the knowledge distillation process, and strengthening reverse engineering detection. For hybrid attacks or attack patterns that cannot be clearly identified, the system adopts a balancing strategy to rationally allocate resources to each defense layer to ensure the balance and robustness of the overall defense capability. Through this adaptive defense strategy, the multi-layer defense system can flexibly adjust the defense focus based on the real-time attack situation, maximizing defense efficiency and success rate.

[0201] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A federated learning backdoor attack defense method based on a multi-layer collaborative defense strategy, characterized in that: The method includes: Step S1: Construct the first, second, and third defense layers throughout the entire life cycle of federated learning. The first defense layer is deployed locally on each client to prevent backdoor implantation during the client training phase. The second and third defense layers are both deployed on the central server. The second defense layer is used to filter malicious updates during the model aggregation phase, and the third defense layer is used to achieve efficient recovery of attacked models during the model deployment phase. Step S2: establishing inter-layer information flow between the first defense layer, the second defense layer, and the third defense layer, wherein the inter-layer information flow includes reputation information flow, anomaly detection information flow, and model recovery score information flow; Step S3: Perform global security status assessment based on inter-layer information flow; Step S4: Implement dynamic allocation of defense resources and adaptive adjustment of defense strategies based on global security status assessment; In step S3, the global security status evaluation is calculated by the following formula: , Where, represents the global safety score, Score the client locally. Score the aggregation phase, Score the model recovery, is the corresponding weight coefficient, n represents the total number of clients, where , Where, 、 are the mean and variance of all client updates in this round, is the update vector of client i at time step t, is a small constant used for numerical stability to avoid the denominator being zero; , Where, represents the maximum value of cosine similarity among all possible client pairs (i, j), is the update vector of client j at time step t; , Where Acc(w t ) represents the model weight w at time step t t The accuracy under , indicates the performance of the current round model.

2. The method for defending against backdoor attacks using a federated learning framework based on a multi-layer collaborative defense strategy as claimed in claim 1, wherein: The first defense layer adopts a comprehensive defense mechanism that integrates gradient random perturbation and dynamic Dropout rate adjustment, and dynamically adjusts the local defense strength based on the client credibility information obtained in real time from the second defense layer.

3. The method for defending against backdoor attacks by federated learning based on a multi-layer collaborative defense strategy according to claim 1, characterized in that: The second defense layer identifies and filters out model updates submitted by malicious clients through reputation evaluation and outlier detection.

4. The method for defending against backdoor attacks using federated learning based on a multi-layer collaborative defense strategy as claimed in claim 1, wherein: The third defense layer combines knowledge distillation with model forgetting, transfers teacher model knowledge, and deletes backdoor features in a targeted manner to efficiently restore the performance of the attacked model, avoiding the high cost of complete retraining.

5. The method for defending against backdoor attacks by federated learning based on a multi-layer collaborative defense strategy according to claim 1, characterized in that: In step S2, the reputation information flow refers to the reputation information obtained by the second defense layer through global perspective calculation. After the second defense layer calculates the reputation score of each client, the reputation information is passed to the first defense layer. The first defense layer adjusts the strength of local defense for different clients based on the received reputation information.

6. A federated learning backdoor attack defense method based on a multi-layer collaborative defense strategy as claimed in claim 5, characterized in that: The reputation information calculated by the second defense layer also flows to the third defense layer and is used to guide the selection of teacher models during the model recovery process.

7. A federated learning backdoor attack defense method based on a multi-layer collaborative defense strategy as claimed in claim 6, characterized in that: In step S2, anomaly detection information is obtained by the first defense layer by performing anomaly detection locally on the client. After being identified, the anomaly detection information is passed to the second defense layer. After receiving the anomaly detection information, the second defense layer integrates it into the global outlier detection and identifies potential backdoor areas through outlier analysis and aggregation strategies.

Citation Information

Patent Citations

  • Data privacy trusted computing method and system

    CN120030582A

  • Methods and apparatuses for defense against adversarial attacks on federated learning systems

    US20210383280A1

Cited By

  • Federal cooperative security defense method and system for industrial control network

    CN121864447A