Federal learning defense method and system for coping with poisoning attack
Through multimodal feature extraction and HDBSCAN clustering algorithm combined with reputation value weighting strategy, the detection problem of data poisoning attacks in federated learning is solved, and efficient identification and isolation of malicious clients is achieved, ensuring the stability and security of the global model.
Patent Information
- Application Number
- CN202510404100.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-08
AI Technical Summary
When facing data poisoning attacks, the existing robust federated learning algorithms rely on a single feature to be bypassed by malicious clients, and cannot fully capture the diverse behavior of the client, resulting in the failure of the defense mechanism.
Multimodal feature extraction and HDBSCAN clustering algorithm are used to identify potential malicious clients, combine reputation value weighting strategies and sparse mechanisms, and verify client behavior by calculating L2 norm differences and historical data, and dynamically adjust the client's weight in the global model.
It significantly improves the detection accuracy of malicious client, can identify long-term abnormal behavior, ensure the stability and security of the global model, reduces computing complexity, and reduces storage and communication costs.
Smart Images

Figure CN120281527A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of federated learning, and specifically provides a federated learning defense method and system for dealing with poisoning attacks. Background Art
[0002] As an emerging machine learning paradigm, federated learning allows multiple remote clients to collaboratively train a shared global model without revealing their local private data. Different from traditional centralized machine learning, FL distributes the training process to each client. Each client uses local data to train the model and then uploads the local model updates to the central server. The central server then aggregates these local updates, for example, through the common federated averaging FedAvg algorithm, to construct a global model. This entire process avoids the transmission of sensitive data over the network and greatly reduces the risk of privacy leakage. This makes FL increasingly popular in data environments that require privacy protection, such as the sharing and collaboration of medical data, fraud detection in the banking and financial industries, and personalized service areas in smart devices.
[0003] Existing robust federated learning algorithms usually rely on the similarity of client model parameters to evaluate the credibility of each client update. By comparing the parameter distributions between different clients, the server can detect those updates that deviate from the global model expectation, thereby identifying and punishing potential malicious clients. Such methods have a certain effect in defending against model poisoning attacks, especially when the attacker directly tampers with the model parameters. However, in the face of data poisoning attacks, such a defense mechanism based on parameter similarity may fail. Because data poisoning often affects the model performance by modifying the training data, and this kind of modification does not significantly change the client's model parameter distribution. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a federated learning defense method and system for dealing with poisoning attacks, which solves the problem that traditional methods usually rely on a single feature to judge whether a client is abnormal, but this method cannot comprehensively capture the diversity of client behavior. Malicious clients will modify multiple features to avoid detection, so the detection based on a single feature is easily bypassed.
[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A federated learning defense method for dealing with poisoning attacks, including the following steps:
[0006] First, assign an initial reputation value to each client. After the client completes local training, the server extracts multi-modal features from each client and constructs a multi-modal feature vector based on the multi-modal features;
[0007] Next, use the HDBSCAN clustering algorithm to perform clustering analysis on the extracted multi-modal feature vectors, and regard the clients that deviate from the normal behavior pattern in the clustering results as potential malicious clients;
[0008] Reduce the reputation value of the clients marked as potential malicious clients in the clustering analysis, thereby forming a weighted aggregation strategy for the reputation value, and the reduction amplitude is half of the current reputation value. Combine the client's historical update data and the sparsification mechanism to further verify its malicious behavior;
[0009] Make a preliminary judgment by calculating the L2 norm difference between the current model update of the client and the global model and comparing it with the difference in the client's historical updates;
[0010] After the preliminary judgment, make a secondary judgment based on the dynamically evaluated reputation value and behavior performance of the client, and exclude it from subsequent training to prevent malicious client updates from affecting the global model;
[0011] Then, based on the weighted aggregation strategy of the reputation value, adjust the weight of the client in the global model aggregation.
[0012] Preferably, the multi-modal features include model weights, update directions, information entropy, and the client's dynamic behavior change rate, and the initial reputation value is 1.
[0013] Preferably, the HDBSCAN clustering algorithm is used to identify outliers in the data, and the results of the clustering analysis are used to mark the abnormal clients as noise or abnormal behavior clients.
[0014] Preferably, the comparison and determination method for the preliminary judgment includes that if the L2 distance difference of the current update exceeds 1.5 times the median of the historical update difference, then mark this client as a potential malicious client and further reduce its reputation value.
[0015] Preferably, the method for the secondary judgment of the behavior performance includes that when the reputation value of the client is lower than the preset threshold, or it shows abnormal behavior in multiple training rounds, regard this client as a malicious client.
[0016] Preferably, the method for adjusting the client includes that the update weight contributed by the client with a lower reputation is smaller, or even zero, and the adjustment of the client is used to reduce the impact of malicious clients on the global model.
[0017] Preferably, the sparsification mechanism calculates the L2 norm difference between the current model update of the client and the global model, quantifies the deviation degree of the client update, and combines the difference in the client's historical updates for dynamic evaluation, so as to identify and isolate abnormal clients.
[0018] Preferably, the dynamic evaluation step includes comparing the difference between the historical update data of the client and the current update. If the difference exceeds a preset threshold, its reputation value is further reduced and marked as a potentially malicious client, and the abnormal pattern is further confirmed by comparing the current update with the historical update.
[0019] Preferably, the reputation value-based weighted aggregation strategy can adjust the aggregation weight of the client in the global model according to the reputation value of the client, so that clients with low reputation will not have too much impact on the global model, and all clients are dynamically adjusted to ensure the stability of the final aggregation.
[0020] A federated learning defense system against poisoning attacks includes:
[0021] The client module is used to perform model training locally on the client and upload features to the server;
[0022] The server module is used to receive the features of the client and perform anomaly detection using the HDBSCAN clustering algorithm;
[0023] The aggregation and update module is used to perform weighted aggregation on the updates based on the reputation value of the client and update the global model;
[0024] The sparsification mechanism module is used to detect the deviation between the client update and the global model by calculating the L2 norm difference to help identify abnormal clients;
[0025] The behavior dynamic evaluation module is used to dynamically adjust the reputation value according to the historical performance of the client to further confirm malicious behavior;
[0026] The client behavior monitoring and recording module is used to record the training history and behavior data of the client;
[0027] The abnormal client isolation and elimination module is used to immediately eliminate the client when it is confirmed to be malicious,
[0028] The system monitoring and feedback module is used to monitor the running state and provide a feedback report.
[0029] The present invention provides a federated learning defense method and system against poisoning attacks. It has the following
[0030] Beneficial effects:
[0031] 1. By extracting multi-dimensional features including model weight distribution, update direction deviation, information entropy, etc., FedMRA captures abnormal behaviors from multiple angles, significantly improving the detection accuracy of malicious clients. Combining historical data analysis, FedMRA can identify clients with abnormal long-term performance and effectively cope with dynamic attack behaviors.
[0032] 2. The present invention adapts through non-I ID data: FedMRA can handle complex data distributions through the HDBSCAN clustering algorithm, identify malicious clients without relying on an external validation dataset, and through a dynamic reputation scoring and sparsification mechanism, FedMRA can still operate stably when a high proportion of malicious clients participate, ensuring that the performance of the global model is not severely affected.
[0033] 3. The present invention, through FedMRA, is completely based on the local update information and historical behavior data uploaded by clients, without the server having to hold additional validation data, reducing storage and communication costs. HDBSCAN does not require presetting the number of clusters, reducing the workload of manually adjusting parameters and enhancing the convenience of model deployment. The sparsification mechanism of FedMRA significantly reduces unnecessary computational complexity by focusing on abnormal update behaviors, making the system operate more efficiently. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is a flowchart of a federated learning defense method for coping with poisoning attacks according to the present invention;
[0035] Figure 2 is an algorithm framework diagram of FedMRA of a federated learning defense method for coping with poisoning attacks according to the present invention;
[0036] Figure 3 is a system architecture diagram of a federated learning defense system for coping with poisoning attacks according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0038] Please refer to the attached Figure 1 , the embodiments of the present invention provide a federated learning defense method for coping with poisoning attacks, including the following steps:
[0039] First, an initial reputation value is assigned to each client, and after the client completes local training, the server extracts multi-modal features from each client and constructs a multi-modal feature vector based on the multi-modal features;
[0040] Then, the HDBSCAN clustering algorithm is used to perform clustering analysis on the extracted multi-modal feature vectors, and the clients that deviate from the normal behavior pattern in the clustering results are regarded as potential malicious clients;
[0041] Reduce the reputation value of the potential malicious clients marked in the clustering analysis, so as to form a weighted aggregation strategy for the reputation value, and the reduction amplitude is half of the current reputation value. Combine the client's historical update data and the sparsification mechanism to further verify its malicious behavior;
[0042] Make a preliminary judgment by calculating the L2 norm difference between the client's current model update and the global model and comparing it with the difference in the client's historical updates;
[0043] After the preliminary judgment, make a secondary judgment based on the dynamically evaluated reputation value and behavior performance of the client, and exclude it from subsequent training to prevent malicious client updates from affecting the global model;
[0044] Then, based on the weighted aggregation strategy of the reputation value, adjust the weight of the client in the global model aggregation.
[0045] Specifically, in the initialization stage, assign an initial reputation value of 1 to each client. After the client completes local training, the server extracts multi-modal features from multiple dimensions and constructs a feature vector to evaluate the client's behavior performance. Then, use the HDBSCAN clustering algorithm to detect malicious clients that deviate from benign clients, and reduce the reputation value R of the malicious clients. By combining the sparsification mechanism and historical data analysis, continuously screen out clients with abnormal behaviors and further reduce their reputation value R. Finally, the system adopts a reputation-based weighted aggregation strategy. Clients with too low reputation values are regarded as malicious clients, and their aggregation weights are set to 0 to ensure the robustness of the global model.
[0046] After each round of client completes local training, the server extracts various feature information from each client, including model weights, update directions, information entropy, and the change rate of client dynamic behavior parameters.
[0047] FedMRA uses the HDBSCAN hierarchical clustering algorithm to cluster the extracted multi-modal feature vectors to identify potential malicious clients in the federated learning system. Using HDBSCAN for malicious client detection has the following advantages:
[0048] Automatically detect outliers: HDBSCAN can effectively identify outliers in the data. In a model poisoning attack, the model parameters of malicious clients usually deviate significantly from the normal distribution, and HDBSCAN marks these adversarial model updates as noise data, thereby effectively detecting malicious clients.
[0049] Adapt to complex clustering shapes: HDBSCAN can handle clusters of any shape. Therefore, in scenarios of data imbalance and non-independent and identically distributed (non-IID), it can reasonably group client models with similar distributions, overcoming the limitations of traditional clustering algorithms.
[0050] No need to preset the number of clusters: HDBSCAN does not require specifying the number of clusters in advance, offering high flexibility. It is especially suitable for the federated learning environment where the client data distribution is unpredictable, ensuring effectiveness in diverse data scenarios.
[0051] Through this clustering method, the server can identify clients deviating from the normal distribution and regard them as potential malicious clients, while setting their reputation value R to 1 / 2 of the current reputation value. Further combined with the client's historical update data and the sparsification mechanism, screen and verify the clients marked as potentially malicious during the clustering process.
[0052] In the sparsification mechanism, first calculate the L2 norm difference between the current global model weights and the client's most recent local model update to quantify the deviation degree of the model update. Subsequently, calculate the L2 distance between all the client's historical updates and the corresponding global model weights.
[0053] Obtain the difference value of each update, and use the median of these difference values as a reference standard to reflect the normal fluctuation range of the client's model update. If the L2 distance difference value of the current update exceeds 1.5 times the median of the historical update difference values, then adjust the client's reputation value R to 1 / 2 again and mark it as a potentially malicious client.
[0054] For each client, dynamically evaluate whether it is a malicious client by calculating its reputation value and historical update behavior. When the client's reputation value is lower than the preset threshold, or it continuously shows abnormal behavior in multiple training rounds (for example, the L2 distance difference of its model update exceeds 1.5 times the historical update difference), then the client will be marked as a malicious client and completely excluded from participating in the subsequent model aggregation process. This paper proposes a reputation-based weighted aggregation strategy. For clients identified as potentially malicious, the server assigns a lower aggregation weight to their model updates.
[0055] Multimodal features include model weights, update directions, information entropy, and the client's dynamic behavior change rate, with an initial reputation value of 1.
[0056] Specifically, first, the model weights refer to the model parameter updates generated by the client after completing the local training task. This feature directly reflects the local training results of the client and is one of the important inputs for global model aggregation. By comparing the differences between the model weights and the current state of the global model, it is possible to preliminarily analyze whether the goals of the client's training are consistent and whether there are abnormal phenomena such as deliberate deviation of model parameters;
[0057] Secondly, the update direction is used to describe the gradient direction of the change in the client model weights, that is, the vector direction of the adjustment of the model parameters during the local training process. In multiple rounds of federated training, the update directions of normal clients should be highly consistent with the global optimization direction. If the update direction of a certain client deviates significantly from that of most clients, it may be a manifestation of malicious interference with the convergence of the global model. Therefore, as one of the highly sensitive indicators, the update direction occupies a key weight in the clustering analysis.
[0058] The HDBSCAN clustering algorithm is used to identify outliers in the data, and the results of the clustering analysis are used to label the abnormal clients as noise or clients with abnormal behaviors.
[0059] Specifically, first, HDBSCAN performs clustering by measuring the density relationship between data points. This algorithm can effectively identify the high-density regions in the data, that is, the groups of clients with normal behaviors, and at the same time can discover the sparse regions, that is, the clients with fewer data points distributed or significantly different from the behaviors of most clients. Based on this density difference, HDBSCAN will classify the clients into multiple categories, including normal clients, abnormal clients, and noise clients. The server will perform HDBSCAN clustering analysis on the multimodal features uploaded by all clients. The multimodal feature vector of each client includes information such as model weights, update directions, information entropy, and dynamic behavior change rates. Through HDBSCAN clustering, the system can identify which clients' behaviors deviate from the normal mode based on these features, and thus label these clients as potential clients with abnormal behaviors. In the clustering analysis, HDBSCAN can not only identify the core members in the group, but also label the noise clients with behaviors far different from those of most clients. Noise clients usually refer to those clients that do not obviously belong to a certain group, and usually show abnormal behaviors during the training process, such as data injection, model poisoning, etc. For these noise clients, the system will pay special attention because they may be manifestations of malicious clients;
[0060] Through this clustering analysis, the system can effectively screen out the clients with abnormal behaviors from numerous clients, providing an important basis for subsequent reputation value adjustment and the identification of malicious behaviors. The clients labeled as having abnormal behaviors will be given priority for reputation value adjustment and removed if necessary to prevent them from affecting the global model. By identifying the abnormal patterns and noises in the clients' behaviors, potential malicious clients can be automatically identified, improving the security and robustness of the federated learning system. The advantage of this algorithm is that it can perform automatic clustering based on density differences without prior labels, thus reducing human intervention and enhancing the system's adaptability in the face of complex environments.
[0061] The comparison and determination initial judgment method includes that if the L2 distance difference updated currently exceeds 1.5 times the median of the historical update differences, the client will be marked as a potential malicious client, and its reputation value will be further reduced.
[0062] Specifically, the initial judgment method is used to determine whether there are potential malicious clients based on the L2 norm difference between the client model update and the global model. The L2 norm difference measures the deviation degree between the current update of the client and the global model, and it can be obtained by calculating the Euclidean distance between the model weights of the local update of the client and the global model. If there is a large L2 distance difference between the model update of the client and the global model, this usually indicates that the behavior of the client may be abnormal or even malicious.
[0063] The determination mechanism of the initial judgment is as follows:
[0064] Calculate the L2 norm difference: In each round of training, the client will upload the updated model after local training to the server, and the server will calculate the L2 norm difference based on the model weights updated by the client and the current global model. The larger the L2 norm difference, the more the update of the client deviates from the global model, and there may be a risk of malicious modification.
[0065] The median of the historical update differences: When making the initial judgment, the server will also calculate the median of the L2 norm differences of each client in the historical training rounds. This median reflects the fluctuation range of the normal update of the client and can be used as a reference standard for normal behavior.
[0066] Set the judgment threshold: If the L2 norm difference updated currently by a certain client exceeds 1.5 times the median of its historical update differences, then this client will be marked as a potential malicious client. The multiple of 1.5 is selected because it can effectively distinguish between normal updates and abnormal updates. This multiple can also be adjusted according to the requirements of the actual application environment in actual use, with high flexibility.
[0067] Marked as a potential malicious client: If the current L2 norm difference of the client exceeds the above-set threshold, it means that there is a large difference between the model update of this client and the global model, and it may be a malicious client. The system will mark this client as a potential malicious client and perform subsequent processing.
[0068] Reduce the reputation value: Once the client is marked as a potential malicious client, the system will adopt the strategy of downgrading its reputation value. This strategy is similar to the reputation value adjustment method based on clustering analysis mentioned before: the reputation value of this client will be reduced by half, reducing its weight in the aggregation of the global model. This way of downgrading the reputation value aims to punish potential malicious behaviors and reduce their impact on the global model.
[0069] Subsequent verification: After reducing the reputation value, the system will further verify whether the client is still a malicious client by combining the client's historical update data and the sparsification mechanism. If the client continues to exhibit abnormal behavior, the system will take more stringent exclusion or isolation measures against it.
[0070] The secondary judgment method of behavior performance includes regarding the client as a malicious client when the client's reputation value is lower than the preset threshold or it exhibits abnormal behavior in multiple training rounds.
[0071] Specifically, the secondary judgment method of behavior performance is to conduct a more in-depth evaluation and verification of the client's behavior after the initial judgment. The initial judgment usually relies on the L2 norm difference and the clustering analysis of multi-modal features to discover potential abnormal clients, while the secondary judgment further verifies whether the client continues to exhibit abnormal behavior or whether its behavior has an adverse impact on the global model. Through this mechanism, the system can more accurately identify malicious clients and prevent their impact on the global model.
[0072] The mechanism of the secondary judgment of behavior performance is as follows:
[0073] After the initial judgment, if the client is marked as a potentially malicious client and its behavior performance has not been effectively improved, then its reputation value will be further reduced in subsequent training rounds. The system sets a preset reputation value threshold. When the client's reputation value drops below this threshold, the system will mark it as a malicious client. This reputation value threshold is usually set within a reasonable range to ensure that it can effectively distinguish normal clients from malicious clients. The weight of clients below this threshold in the global model aggregation will be further reduced or even completely excluded;
[0074] In addition to the continuous decrease of the reputation value, the behavior performance of the client in multiple training rounds is also a key basis for the secondary judgment. The system will evaluate the behavior stability of the client in multiple rounds of training, such as the L2 norm difference, update direction, information entropy and other multi-modal features in each training round. If the client continues to exhibit abnormal behavior in multiple training rounds and its behavior never returns to the normal track, it can be determined as a malicious client. This multi-round evaluation can effectively reduce the misjudgment caused by accidental errors or temporary behavior abnormalities, ensuring that the system can accurately identify malicious clients with long-term abnormal performance;
[0075] Through a secondary judgment mechanism, the behavior of the client can be analyzed and evaluated more strictly. This process combines multiple dimensions such as reputation values, multi-round training behaviors, and anomaly criteria to ensure that malicious clients can be identified at an early stage and removed from the global model, thus effectively guaranteeing the stability and robustness of the global model. The secondary judgment provides an additional line of defense to prevent occasional misjudgments or external interferences, further enhancing the system's adaptability in complex environments.
[0076] The way to adjust the client includes that clients with lower reputation contribute less update weight, even zero, to adjust the client to reduce the impact of malicious clients on the global model.
[0077] Specifically, first, calculate the reputation value of the client: The system will calculate the reputation value of each client according to the client's performance in multiple rounds of training, including features such as its model update, L2 norm difference, and update direction. The initial reputation value is set to 1, indicating that the client is trusted by default when the system starts. As the training progresses, the reputation value of the client will be dynamically adjusted according to its behavior performance. If the client behaves abnormally, its reputation value will decrease, otherwise it will increase.
[0078] Secondly, adjust the weight of the client according to the reputation value: The reputation value of the client directly affects its weight in the global model aggregation. Clients with higher reputation will occupy a larger weight, while the contribution of clients with lower reputation (especially potentially malicious clients) in the global model will be reduced. Clients with higher reputation values will contribute more update weights, promoting the global model to converge more accurately. Clients with lower reputation values may even be zero, and their model updates will be ignored and they will not participate in the aggregation of the global model.
[0079] Then, if the reputation value of the client continues to be lower than the preset threshold or behaves abnormally, remove the malicious client: If the reputation value of a client continues to be lower than the set threshold in multiple rounds of training, or its behavior continues to be abnormal in multiple training rounds, the system will remove it from subsequent training, that is, no longer consider the impact of its model update on the global model. Removing malicious clients ensures that these clients do not interfere with the training process of the global model. The system makes dynamic adjustments and optimizations: To ensure the system's adaptability in the face of complex attacks, the system will make dynamic adjustments according to the reputation value and behavior changes of the client. At the end of each round of training, the system will recalculate the behavior data of the client, re-evaluate its reputation value, and dynamically adjust the weight of the client during the global model aggregation process. If the behavior of a certain client returns to normal, the reputation value can gradually recover. On the contrary, the weight of malicious clients will continue to decrease or even be removed.
[0080] Finally, the global model is updated based on the weighted aggregation strategy: the model updates of all clients are weighted and aggregated according to their weights. For clients with a higher reputation value, their updates will account for a larger proportion; while for clients marked as malicious, their updates will be ignored or completely excluded. In this way, the negative impact of malicious clients on the global model is reduced, ensuring the stability and credibility of the global model during training. By gradually adjusting the client weights and monitoring the behavior in this way, the system can effectively identify and isolate malicious clients, reducing their interference with the global model, thus guaranteeing the security and stability of the federated learning system.
[0081] The sparsification mechanism quantifies the deviation degree of the client's update by calculating the L2 norm difference between the client's current model update and the global model, and dynamically evaluates it in combination with the differences in the client's historical updates, so as to identify and isolate abnormal clients.
[0082] Specifically, first, calculate the L2 norm difference between the client's current model update and the global model: in each round of training, the client submits the model update after local training. The system calculates the L2 norm difference between the client's current model update and the global model, that is, calculates the Euclidean distance between the two models. The larger the L2 norm difference, the greater the deviation of the client's update from the global model, and there may be abnormal or malicious updates.
[0083] Secondly, quantify the deviation degree of the client's update: by calculating the L2 norm difference, the system can quantify the deviation degree between the client's model update and the global model. This deviation degree reflects the possible abnormal behavior of the client during training. For example, if a client's update deviates far from the global model, it may mean that the client's update has malicious intentions (such as data pollution or model poisoning), or there are errors in the client's training.
[0084] Then, dynamically evaluate in combination with the differences in the client's historical updates: to more accurately judge whether the client's update is abnormal, the system not only considers the current L2 norm difference, but also incorporates the differences in the client's historical updates into the evaluation. The system calculates the average or median of the L2 norm differences of the client in multiple rounds of training, and compares the difference of the current update with the differences in historical updates. If the L2 norm difference of the current update is much larger than the normal fluctuation range of historical updates, it means that the client's update behavior deviates from the normal training mode and may be the manifestation of a malicious client.
[0085] Finally, isolate the abnormal clients and adjust the weights: After identifying the abnormal clients, the system will make dynamic adjustments based on the reputation values of these clients. If the abnormal behavior of a client persists, the system will reduce its weight in the global model aggregation or even completely exclude it to prevent the updates of malicious clients from having an adverse impact on the stability and accuracy of the global model. In this way, the system can effectively isolate the abnormal clients and ensure that the model updates of normal clients can dominate the training process of the global model. Through the sparsification mechanism, the system can, based on the L2 norm difference of the client model updates and the dynamic evaluation of historical behaviors, promptly detect and isolate abnormal clients. This mechanism ensures that only clients with normal performance and consistent with the global model can participate in the training of the global model, preventing interference from malicious clients and thus guaranteeing the security and stability of the federated learning system.
[0086] The dynamic evaluation step includes comparing the difference between the historical update data of the client and the current update. If the difference exceeds the preset threshold, the system will further reduce its reputation value and mark it as a potentially malicious client, and use the comparison between the current update and the historical update to further confirm the abnormal pattern.
[0087] Specifically, first, compare the difference between the historical update data of the client and the current update: The system will record the historical update data of the client in multiple training rounds, and these data include the L2 norm difference between the client and the global model in each round of training. To evaluate whether the current update is abnormal, the system will calculate the difference between the current model update of the client and the historical update. By comparing the difference between the current update and the historical update, the system can quantify the stability and consistency of the client's behavior.
[0088] Secondly, if the difference exceeds the preset threshold, further reduce its reputation value and mark it as a potentially malicious client: If the difference between the current update of the client and its historical updates exceeds the preset threshold (e.g., more than 1.5 times the historical update difference), the system will regard the behavior of this client as abnormal and adjust its reputation value. Specifically, the system will reduce the reputation value of this client and mark it as a potentially malicious client. This process can reduce the impact of malicious clients on the global model by reducing their weight in the global model aggregation or, in severe cases, excluding them from subsequent training. Then, use the comparison between the current update and the historical update to further confirm the abnormal pattern: After marking the potentially malicious client, the system will further analyze to confirm whether it has malicious behavior. The system will further analyze whether this client shows a continuous abnormal pattern by combining the comparison between the current update and the historical update. If the update behavior of the client continuously deviates from the normal pattern (e.g., both the model update direction and the update amplitude show abnormalities), it can be confirmed as a malicious client. This confirmation step can use different anomaly detection algorithms, such as trend analysis based on historical data, model change monitoring, etc., to ensure the detection accuracy of the system.
[0089] Next, according to the confirmed abnormal pattern, further adjust the reputation value and isolate the abnormal client: If it is further confirmed through dynamic evaluation that the client is a malicious client, the system will handle it more severely. This includes: further reducing its reputation value so that its weight in the global model aggregation is close to zero, or completely excluding its impact on the global model update, and excluding this client from subsequent training rounds to prevent its model update from having any impact on the global model.
[0090] Finally, adjust the aggregation strategy of the global model based on the dynamic evaluation results: Through dynamic evaluation and reputation value adjustment, the system can adapt to changes in client behavior in real time, thereby reducing the impact of malicious clients during the global model aggregation process. For normal clients, their weights can be gradually restored, while for malicious clients, their influence is continuously reduced to ensure the stability and credibility of the global model. Through the dynamic evaluation step, the system can analyze and dynamically adjust the behavior of clients in real time, enabling the interference of malicious clients to be detected and isolated in a timely manner. The comparison with historical updates provides a relatively stable benchmark to help the system judge whether there is abnormal behavior, and the dynamic reputation value adjustment can ensure that the impact of malicious clients on the global model is minimized. Through this mechanism, the system can effectively respond to various potential attacks and ensure the security and robustness of federated learning.
[0091] The weighted aggregation strategy based on reputation value can adjust the aggregation weight of the client in the global model according to the client's reputation value, which is used when clients with low reputation will not have too much impact on the global model, and dynamically adjust all clients to ensure the stability of the final aggregation.
[0092] Specifically, first, calculate the weight according to the client's reputation value: The client's reputation value directly affects its weight in the global model aggregation. Clients with higher reputation contribute more, while clients with lower reputation contribute less or even zero, avoiding malicious clients from affecting the global model.
[0093] Secondly, dynamically adjust the weight: The system dynamically adjusts the weight according to the client's performance. If the client's reputation value is low, its weight will be reduced, and the weight of malicious clients can be reduced to zero. The weight of well-behaved clients will gradually recover to ensure the fairness and reasonableness of the aggregation process.
[0094] Then, perform weighted aggregation: In the global model update, the system performs weighted aggregation according to the client weights to ensure that the updates of clients with higher reputation have a greater impact on the global model. The updates of malicious clients are minimized to ensure the stability of the model. Through this weighted aggregation strategy, the system can reduce the impact of malicious clients on the global model and ensure the stability and security of the model.
[0095] Please refer to the appendix Figure 2 , a federated learning defense system against poisoning attacks, including:
[0096] Client module, used to perform model training locally on the client and upload features to the server;
[0097] Server module, used to receive the features of the client and perform anomaly detection using the HDBSCAN clustering algorithm;
[0098] Aggregation and update module, used to perform weighted aggregation on the updates based on the client's reputation value and update the global model;
[0099] Sparsification mechanism module, used to detect the deviation between the client update and the global model by calculating the L2 norm difference to help identify abnormal clients;
[0100] Behavior dynamic evaluation module, used to dynamically adjust the reputation value according to the client's historical performance to further confirm malicious behavior;
[0101] Client behavior monitoring and recording module, used to record the training history and behavior data of the client;
[0102] Abnormal client isolation and elimination module, used to immediately eliminate the client when it is confirmed to be malicious,
[0103] The system monitoring and feedback module is used to monitor the running status and provide feedback reports.
[0104] Specifically, the client module is responsible for local model training on the client side. Each client trains based on local data, updates local model parameters, and uploads relevant feature data to the server. These feature data include model weights, update directions, information entropy, etc., for subsequent anomaly detection and analysis;
[0105] Secondly, the server receives the feature data from each client. The server performs clustering analysis on the multi-modal features uploaded by the clients through the HDBSCAN clustering algorithm to identify the differences between normal clients and potential malicious clients. The clustering algorithm can identify abnormal clients as noise or abnormal behavior clients based on density differences;
[0106] The aggregation and update module is responsible for weighted aggregation based on the reputation values of the clients and updating the global model. Clients with higher reputation values will have a greater impact on the update of the global model, while clients with lower reputation values contribute less. This module ensures that the impact of malicious clients on the global model is minimized;
[0107] Then, the sparsification mechanism module detects the deviation of the client update by calculating the L2 norm difference between the client model update and the global model. Clients with a large L2 norm difference may exhibit abnormal behavior, and the system uses this mechanism to help identify these potential malicious clients;
[0108] Secondly, the behavior dynamic evaluation module dynamically adjusts the reputation value according to the historical behavior data and current performance of the client. When the update behavior of the client deviates from the normal mode, its reputation value will decrease, and the system can further confirm whether it is a malicious client and make corresponding adjustments;
[0109] Next, the client behavior monitoring and recording module continuously records the training history and behavior data of the client. These data are used to evaluate the stability and credibility of the client and support the subsequent dynamic evaluation and anomaly detection processes;
[0110] And through the abnormal client isolation and elimination module, when it is confirmed that the client is malicious, this module immediately isolates and eliminates it to prevent its impact on the global model. Malicious clients will no longer participate in the global model update in subsequent training to ensure the security of the model;
[0111] Finally, the system monitoring and feedback module monitors the running status of the system in real time, collects the feedback reports of each module, and evaluates the system performance. The monitoring results can help administrators discover potential problems in a timely manner and make adjustments.
[0112] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A federated learning defense method against poisoning attacks, characterized in that, It includes the following steps: First, assign an initial reputation value to each client. After the client completes local training, the server extracts multimodal features from each client and constructs a multimodal feature vector based on the multimodal features; Next, use the HDBSCAN clustering algorithm to perform clustering analysis on the extracted multimodal feature vectors, and regard the clients deviating from the normal behavior pattern in the clustering results as potential malicious clients; Reduce the reputation value of the clients marked as potential malicious clients in the clustering analysis, thereby forming a weighted aggregation strategy for the reputation value, and the reduction amplitude is half of the current reputation value. Combine the client's historical update data and the sparsification mechanism to further verify its malicious behavior; Make a preliminary judgment by calculating the L2 norm difference between the client's current model update and the global model and comparing it with the difference in the client's historical updates; After the preliminary judgment, make a secondary judgment based on the dynamic evaluation of the client's reputation value and behavior performance, and exclude it from subsequent training to prevent malicious clients from updating and affecting the global model; Then, based on the weighted aggregation strategy of the reputation value, adjust the weight of the client in the global model aggregation.
2. The federated learning defense method and system for coping with poisoning attacks according to claim 1, characterized in that: The multimodal features include model weights, update directions, information entropy, and the client's dynamic behavior change rate, and the initial reputation value is 1.
3. A federated learning defense method against poisoning attacks according to claim 1, characterized in that: The HDBSCAN clustering algorithm is used to identify outliers in the data, and the results of the clustering analysis are used to mark abnormal clients as noise or abnormal behavior clients.
4. A federated learning defense method against poisoning attacks according to claim 1, characterized in that: The comparison and determination method for the preliminary judgment includes that if the L2 distance difference of the current update exceeds 1.5 times the median of the historical update difference, mark this client as a potential malicious client and further reduce its reputation value.
5. A federated learning defense method against poisoning attacks according to claim 1, characterized in that: The method for the secondary judgment of the behavior performance includes that when the client's reputation value is lower than the preset threshold, or it shows abnormalities in multiple training rounds, regard this client as a malicious client.
6. A federated learning defense method against poisoning attacks according to claim 1, characterized in that: The method for adjusting the client includes that the update weight contributed by the client with a lower reputation is smaller, or even zero. The adjustment of the client is used to reduce the impact of malicious clients on the global model.
7. A federated learning defense method against poisoning attacks according to claim 1, characterized in that: The sparsification mechanism calculates the L2 norm difference between the client's current model update and the global model, quantifies the deviation degree of the client's update, and combines the difference in the client's historical updates for dynamic evaluation, so as to identify and isolate abnormal clients.
8. A federated learning defense method against poisoning attacks according to claim 1, characterized in that: The dynamic evaluation step includes comparing the difference between the client's historical update data and the current update. If the difference exceeds the preset threshold, further reduce its reputation value and mark it as a potential malicious client, and use the comparison between the current update and the historical update to further confirm the abnormal pattern.
9. A federated learning defense method against poisoning attacks according to claim 1, characterized in that: The weighted aggregation strategy based on the reputation value can adjust the aggregation weight of the client in the global model according to the client's reputation value, which is used when clients with lower reputations will not have too much impact on the global model, and dynamically adjust all clients to ensure the stability of the final aggregation.
10. A federated learning defense system against poisoning attacks, characterized in that, Applied to a federated learning defense method for coping with poisoning attacks described in any one of claims 1-9, it includes: A client module, which is used to perform model training locally on the client and upload features to the server; A server module, which is used to receive the characteristics of clients and perform anomaly detection using the HDBSCAN clustering algorithm; An aggregation and update module, which is used to perform weighted aggregation on updates based on the reputation values of clients and update the global model; A sparsification mechanism module, which is used to detect the deviation between the client update and the global model by calculating the L2 norm difference to help identify abnormal clients; A behavior dynamic evaluation module, which is used to dynamically adjust the reputation value according to the historical performance of clients to further confirm malicious behavior; A client behavior monitoring and recording module, which is used to record the training history and behavior data of clients; An abnormal client isolation and elimination module, which is used to immediately eliminate a client when it is confirmed to be malicious; A system monitoring and feedback module, which is used to monitor the running status and provide a feedback report.
Citation Information
Cited By
Federal learning system fusing hierarchical modeling, dynamic weight and reputation supervision
CN120851251A