Federal learning anti-poisoning method for high-proportion malicious clients

Through principal component analysis and hierarchical aggregation clustering combined with multi-dimensional feature analysis, malicious clients in federated learning are identified and eliminated, and the global model damage caused by a high proportion of malicious clients is solved, and the robustness and security of the system are improved.

CN120474810APending Publication Date: 2025-08-12GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510758149.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

Existing federated learning defense methods are difficult to effectively detect and eliminate malicious clients when facing a high proportion of malicious clients, resulting in the global model being corrupted, especially in the absence of a clean validation dataset.

Method used

Principal component analysis and hierarchical aggregation clustering algorithm are used to reduce and cluster client model parameters, combine multi-dimensional feature analysis to identify potential malicious clients, and design malicious evaluation and detection modules for precise distinction, and eliminate malicious updates.

Benefits of technology

In a high proportion of malicious client environment, the robustness and security of the federated learning system are significantly improved, ensuring the accuracy and stability of the global model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120474810A_ABST
    Figure CN120474810A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of federated learning, and particularly relates to a federated learning anti-poisoning attack defense method for high-proportion malicious clients. Aiming at the problem that the existing robust aggregation method is difficult to deal with high-proportion malicious client poisoning attacks in a federal learning training process, the invention provides a novel defense method. Comprising the following steps: firstly, carrying out dimension reduction processing on model parameters uploaded by clients by utilizing principal component analysis to remove redundant information and extract key features; secondly, grouping the parameters subjected to dimension reduction by adopting a hierarchical agglomeration clustering algorithm to realize preliminary identification of potential malicious clients; and finally, designing a malicious evaluation detection module, comprehensively considering gradient contribution and parameter difference of client model updating, and further accurately distinguishing a benign client and a malicious client. According to the method, malicious updating can be effectively detected and eliminated in an environment with a high malicious client occupation ratio, and the robustness and the safety of a federated learning system are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of federated learning technology, and specifically provides a federated learning anti-poisoning method for a high proportion of malicious clients. Background Art

[0002] Federated Learning (FL) is a collaborative learning technique that allows multiple participants to jointly train machine learning models without sharing their own training data. Each participant trains their local model and shares their local updates with a server, which then aggregates the updates to generate a global model and sends it back to the participants. This technique has two important characteristics. First, it can train a model based on massive amounts of data provided by thousands of participants without compromising their data privacy. Second, decentralized training allows parallel computing across multiple devices, significantly increasing the speed of the training process. Due to these attractive properties, FL technology has been used to solve problems in many different fields, such as autonomous vehicles, Industry 4.0, and healthcare.

[0003] However, due to its distributed nature, FL is vulnerable to model poisoning attacks, in which a malicious client controlled by an attacker corrupts the global model by sending manipulated model updates to the server. The malicious client controlled by the attacker can be injected into a fake client or a real client compromised by the attacker. Based on the attack target, model poisoning attacks can generally be divided into untargeted attacks and targeted attacks. In untargeted model poisoning attacks, the corrupted global model indiscriminately makes incorrect predictions for a large number of test inputs. In targeted model poisoning attacks, the corrupted global model makes incorrect predictions for test inputs selected by the attacker, while the accuracy of the global model for other test inputs is unaffected. For example, the test input selected by the attacker may be a test input embedded with a trigger selected by the attacker, which is also called a backdoor attack.

[0004] Existing defenses against model poisoning attacks primarily rely on Byzantine-robust federated learning methods, such as the provably robust Krum and FLTrust methods. These methods aim to learn an accurate global model even if some clients are malicious and send arbitrary model updates to the server. Byzantine-robust federated learning methods theoretically limit the changes that malicious clients can make to global model parameters, while provably robust federated learning methods guarantee lower bounds on test accuracy in the presence of malicious clients. However, these methods can only handle a small number of malicious clients or require the server to have a clean and representative validation dataset. For example, Krum's method can theoretically tolerate up to (n-2) / 2 malicious clients, where n is the total number of clients. While FLTrust can protect against a large number of malicious clients, it does so only if the server has a clean validation dataset whose distribution does not deviate significantly from the overall training data distribution. Therefore, in typical federated learning scenarios, when the server lacks such a validation dataset, the global model can still be corrupted by a large number of malicious clients.

[0005] This paper addresses the challenging scenario of federated learning where malicious clients may constitute the majority. It addresses the failure of existing detection methods when the number of malicious clients approaches or exceeds the number of benign clients. Furthermore, it does not require any clean validation dataset or a pre-defined number of malicious clients. Summary of the Invention

[0006] The present invention provides a federated learning anti-poisoning method for a high proportion of malicious clients, which solves the detection dilemma when malicious clients account for the majority.

[0007] In order to achieve the above-mentioned purpose, the present invention adopts the following technical solutions.

[0008] A federated learning anti-poisoning method for a high proportion of malicious clients includes the following steps:

[0009] 1) The server distributes the initial global model to each client, including benign clients and malicious clients;

[0010] 2) Each client trains the received global model based on local data and uploads the updated results to the server;

[0011] 3) The server performs principal component analysis and dimensionality reduction on the model update data uploaded by the client to remove redundant information and retain the main features;

[0012] 4) Based on the client model parameters after dimensionality reduction in step 3), a hierarchical agglomerative clustering algorithm is used to perform cluster analysis on the client model parameters, dividing the clients into two clusters to achieve a preliminary distinction between normal and potentially malicious clients, and lay the foundation for subsequent identification and elimination of malicious clusters;

[0013] 5) Based on the clustering results in step 4), perform cluster-level numerical feature analysis on the client model update, integrating statistical features with frequency domain features to identify potential malicious behaviors;

[0014] 6) Based on the feature analysis results from step 5), the clients are classified into benign and malicious clients, model updates from benign clients are retained, and updates uploaded by malicious clients are discarded;

[0015] 7) The server aggregates the model update results of the benign clients screened in step 6) to update the global model, thereby enhancing the robustness and security of the federated learning system in the presence of a high proportion of malicious clients.

[0016] Preferably, in step 3), the server performs principal component analysis and dimensionality reduction processing on the model update data uploaded by the client to remove redundant information and retain the main features, including:

[0017] The server first treats the local model parameters uploaded by each client in the current training round as a high-dimensional vector, and stacks the model update data of all clients into a parameter matrix. Next, the matrix is centralized to eliminate the interference of the mean difference of each dimension feature on subsequent calculations, and then its covariance matrix is calculated to reflect the statistical correlation between each parameter dimension. Subsequently, the server extracts the principal component direction through eigenvalue decomposition and constructs a projection space to map the original model parameters to a feature subspace with lower dimension and more concentrated information, thereby obtaining a low-dimensional representation of each client. This dimensionality reduction process not only significantly reduces data redundancy and reduces the computational overhead of the clustering and evaluation modules, but also effectively retains the main differential features in the client model update, providing a key feature basis and judgment basis for subsequent hierarchical clustering and malicious client identification.

[0018] Preferably, step 4) uses a hierarchical agglomerative clustering algorithm to perform cluster analysis on the client model parameters, dividing the clients into two clusters to achieve a preliminary distinction between normal and potentially malicious clients, and lay the foundation for subsequent identification and elimination of malicious clusters, including:

[0019] Each client is initially a separate cluster, and a bottom-up strategy is used to gradually merge the most similar clusters until a complete cluster structure is constructed. The Ward link method is introduced as an aggregation criterion for measuring inter-cluster distances. Its goal is to minimize the increase in intra-cluster squared error during each cluster merging process. Specifically, for any two clusters A and B, the merging cost is defined as:

[0020]

[0021] Among them, |A| and |B| represent the number of samples in the two clusters respectively, μ A and μ B is the centroid of each cluster, and ||·|| represents the Euclidean norm. This metric minimizes the “information loss” introduced by each merging step, which is conducive to obtaining a compact and discriminative clustering structure.

[0022] The clustering process is recorded by constructing a connection matrix, which records the merging relationship between clusters, the merging distance, and the number of samples in the new cluster, and then constructs a complete binary clustering tree. In this tree structure, the leaf nodes represent the original clients, and the internal nodes record each merging event. Let the connection matrix generated in the clustering process be The i-th row represents the information of the merging, which is in the form of:

[0023] Li=[c1,c2,d,s]

[0024] Where c1 and c2 are the indices of the clusters to be merged, d is the distance value at the time of merging, and s is the number of clients in the merged cluster. After the cluster tree is constructed, the subtree is recursively traversed from the root node from top to bottom to identify and eliminate abnormal clusters that are too small. To this end, a minimum cluster size threshold τ is set. If the number of leaf nodes |C| of a subtree satisfies:

[0025] |C|<τ

[0026] This subcluster is then considered an outlier, and the clients it contains are temporarily removed from the clustering structure to prevent extremely heterogeneous data from disrupting the overall structure. It's important to emphasize that clients in such outlier clusters are not necessarily malicious; they may be mistakenly excluded from the main cluster because their local data distribution deviates from the main modality. Therefore, this strategy employs a conservative removal mechanism, and their behavioral characteristics will be further evaluated in subsequent multidimensional feature analysis.

[0027] Preferably, the step 5) performs cluster-level numerical feature analysis on the client model update based on the clustering results, and integrates statistical features with frequency domain features to identify potential malicious behaviors, including:

[0028] After completing the hierarchical clustering of the clients, the server updates the features based on the model of the clients in each cluster, extracts multi-dimensional statistical indicators and performs comprehensive discriminant analysis to identify potential malicious client behaviors. The feature analysis module includes the following steps:

[0029] First, the average gradient norm of each cluster is calculated to measure the overall strength of the client model update within the cluster, which is defined as:

[0030]

[0031] Where N represents the number of clients in the cluster, g i The gradient vector uploaded by the i-th client. This metric can reveal the behavior of malicious clients amplifying the influence of the model through scaling operations in local updates.

[0032] Secondly, we extract the gradient distribution anomaly index to measure the consistency of the distribution structure of client updates within the cluster. This index is calculated based on the variance of skewness and kurtosis and is defined as follows:

[0033]

[0034] in and are the variances of the skewness and kurtosis within the cluster, respectively, and γ is a modulating factor. Benign clients exhibit higher variance in their high-order statistical features due to large differences in data distribution. However, malicious clients controlled by the same strategy tend to update more consistently, resulting in a significant decrease in variance and an increase in the gradient distribution anomaly index.

[0035] Next, calculate the update direction entropy and measure the consistency of the gradient update direction of each client in the cluster based on the information entropy principle. Let the probability distribution of the update direction angle be {p k}, then the updated direction entropy is expressed as:

[0036]

[0037] Where K represents the number of discrete bins of direction angles, {p k} is the probability of falling into the kth interval, and ∈ is a stability term. The lower the entropy value, the more concentrated the update direction is, which usually indicates highly coordinated malicious behavior.

[0038] Finally, the spectrum normalization index is calculated. For each client i in the cluster, its gradient vector g is obtained. i , and calculate its one-dimensional fast Fourier transform:

[0039]

[0040] Then normalize it to get the power spectral density distribution:

[0041]

[0042] where ε is a numerical stability constant.

[0043] Then calculate the cosine similarity between any two client spectra:

[0044]

[0045] Calculate for all client pairs to obtain the spectrum similarity matrix.

[0046] Finally, the mean μ and standard deviation σ of all spectrum similarities are calculated , Constructing spectrum consistency indicators:

[0047] SCI=μ·(1-σ)

[0048] This indicator encourages the coexistence of overall high consistency and structural stability, reflecting the overall synchronization degree of the client cluster in the spectrum structure.

[0049] To improve the accuracy and robustness of malicious client identification after clustering, we designed a cluster-level discrimination mechanism based on multi-feature priority judgment. This mechanism leverages three statistical features—average gradient norm, gradient distribution anomaly index, and calculated update direction entropy—as well as an auxiliary feature spectral normalization index. This mechanism comprehensively models client update behavior at the cluster level and, combined with hierarchical rules, enables accurate identification of malicious clusters. Compared to traditional heuristic methods that rely on cluster size, this mechanism exhibits greater attack adaptability and maintains effective discrimination, especially when the proportion of malicious clients is high.

[0050] Preferably, the step 6) includes distinguishing clients into benign clients and malicious clients based on feature analysis and judgment, retaining model updates from benign clients, and discarding updates uploaded by malicious clients, including:

[0051] First, if the ratio of the average gradient norms of two subclusters is within a set threshold, it is considered that there is significant gradient amplification behavior. In this case, the cluster with the larger norm is directly judged as malicious.

[0052] If the norm ratio is insufficient to provide effective differentiation, the system further checks for extreme values in three characteristic metrics (gradient distribution anomaly index, computational update direction entropy, and spectral normalization metric). The system counts the number of metrics reaching the extreme value threshold within each cluster, and considers the cluster with the most extreme values malicious. If two clusters have the same number of extreme values, the system proceeds to further fine-grained analysis.

[0053] If the number of extreme values is equal or neither cluster has an extreme value, the two clusters are further evaluated for their gradient distribution anomaly index. This metric reflects the consistency of the distribution structure by measuring the skewness and kurtosis variance of the client gradients within the cluster. A higher gradient distribution anomaly index generally indicates highly similar update patterns within the cluster, indicating a potential coordinated attack. When the difference in the gradient distribution anomaly index between two clusters exceeds a set threshold, the cluster with the higher anomaly is considered malicious.

[0054] If none of the above three metrics are significantly discriminative, the computational update direction entropy is introduced as the final basis. This metric measures the uncertainty of the gradient direction distribution within a cluster. Lower computational update direction entropy values indicate more consistent gradient directions, and thus a higher likelihood of organized malicious behavior. In particular, if a cluster's computational update direction entropy is extremely low, it can be directly considered malicious; otherwise, the cluster with the lower computational update direction entropy is selected as the final judgment.

[0055] This feature analysis mechanism integrates three complementary features: update intensity, statistical structure, and directional information. It has good discrimination and robustness, and is particularly suitable for dealing with federated learning security issues in complex environments with a high proportion of malicious clients. It significantly improves the overall stability and attack defense performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. The drawings in the description are only some embodiments of the present invention.

[0057] Figure 1 It is a schematic flow chart of a specific implementation of the present invention. DETAILED DESCRIPTION

[0058] The following is a further description of the present invention with reference to the accompanying drawings and specific embodiments, but the content of the invention is not limited to the scope of the invention. The basic process of the invention is as follows Figure 1 shown.

[0059] The following combination Figure 1 This article introduces a federated learning anti-poisoning method for a high proportion of malicious clients, provided by an embodiment of the present invention. The specific implementation steps include:

[0060] S1 server distributes the initial global model to each client, including benign clients and malicious clients;

[0061] Each S2 client trains the received global model based on local data and uploads the updated results to the server;

[0062] The S3 server performs principal component analysis and dimensionality reduction on the model update data uploaded by the client to remove redundant information and retain the main features;

[0063] S4 uses a hierarchical agglomerative clustering algorithm to perform cluster analysis on client model parameters, dividing the clients into two clusters to achieve a preliminary distinction between normal and potentially malicious clients, and lay the foundation for the subsequent identification and elimination of malicious clusters;

[0064] S5 performs feature analysis based on the clustering results in step S4, and identifies clients with abnormal features in combination with the numerical features updated by the model, thereby determining whether there is malicious behavior;

[0065] S6 classifies the clients into benign clients and malicious clients based on the feature analysis results of step S5, retains the model updates of benign clients, and removes the updates uploaded by malicious clients;

[0066] The S7 server aggregates the model update results of the benign clients screened out in step S6 to update the global model, thereby enhancing the robustness and security of the federated learning system when there is a high proportion of malicious clients.

[0067] Preferably, in step S3, the server performs principal component analysis and dimensionality reduction processing on the model update data uploaded by the client to remove redundant information and retain the main features, including:

[0068] The server first treats the local model parameters uploaded by each client in the current training round as a high-dimensional vector, and stacks the model update data of all clients into a parameter matrix. Next, the matrix is centralized to eliminate the interference of the mean difference of each dimension feature on subsequent calculations, and then its covariance matrix is calculated to reflect the statistical correlation between each parameter dimension. Subsequently, the server extracts the principal component direction through eigenvalue decomposition and constructs a projection space to map the original model parameters to a feature subspace with lower dimension and more concentrated information, thereby obtaining a low-dimensional representation of each client. This dimensionality reduction process not only significantly reduces data redundancy and reduces the computational overhead of the clustering and evaluation modules, but also effectively retains the main differential features in the client model update, providing a key feature basis and judgment basis for subsequent hierarchical clustering and malicious client identification.

[0069] Preferably, step S4 uses a hierarchical agglomerative clustering algorithm to perform cluster analysis on the client model parameters, dividing the clients into two clusters to achieve a preliminary distinction between normal and potentially malicious clients, and lay the foundation for subsequent identification and elimination of malicious clusters, including:

[0070] Each client is initially a separate cluster, and a bottom-up strategy is used to gradually merge the most similar clusters until a complete cluster structure is constructed. The Ward link method is introduced as an aggregation criterion for measuring inter-cluster distances. Its goal is to minimize the increase in intra-cluster squared error during each cluster merging process. Specifically, for any two clusters A and B, the merging cost is defined as:

[0071]

[0072] Among them, |A| and |B| represent the number of samples in the two clusters respectively, μ A and μ B is the centroid of each cluster, and ||·|| represents the Euclidean norm. This metric minimizes the “information loss” introduced by each merging step, which is conducive to obtaining a compact and discriminative clustering structure.

[0073] The clustering process is recorded by constructing a connection matrix, which records the merging relationship between clusters, the merging distance, and the number of samples in the new cluster, and then constructs a complete binary clustering tree. In this tree structure, the leaf nodes represent the original clients, and the internal nodes record each merging event. Let the connection matrix generated in the clustering process be The i-th row represents the information of the merging, which is in the form of:

[0074] Li=[c1,c2,d,s]

[0075] Among them, c1 and c2 are the indexes of the merged clusters, d is the distance value at the time of merging, and s is the number of clients in the merged cluster.

[0076] After the clustering tree is constructed, the subtree is recursively traversed from the root node to identify and remove abnormal clusters that are too small. For this purpose, a minimum cluster size threshold τ is set. If the number of leaf nodes |C| of a subtree satisfies:

[0077] |C|<τ

[0078] This subcluster is then considered an outlier, and the clients it contains are temporarily removed from the clustering structure to prevent extremely heterogeneous data from disrupting the overall structure. It's important to emphasize that clients in such outlier clusters are not necessarily malicious; they may be mistakenly excluded from the main cluster because their local data distribution deviates from the main modality. Therefore, this strategy employs a conservative removal mechanism, and their behavioral characteristics will be further evaluated in subsequent multidimensional feature analysis.

[0079] Preferably, the step S5, based on the clustering results, performs cluster-level numerical feature analysis on the client model update, and integrates statistical features with frequency domain features to identify potential malicious behaviors, including:

[0080] After completing the hierarchical clustering of the clients, the server updates the features based on the model of the clients in each cluster, extracts multi-dimensional statistical indicators and performs comprehensive discriminant analysis to identify potential malicious client behaviors. The feature analysis module includes the following steps:

[0081] First, the average gradient norm of each cluster is calculated to measure the overall strength of the client model update within the cluster, which is defined as:

[0082]

[0083] Where N represents the number of clients in the cluster, g i The gradient vector uploaded by the i-th client. This metric can reveal the behavior of malicious clients amplifying the influence of the model through scaling operations in local updates.

[0084] Secondly, we extract the gradient distribution anomaly index to measure the consistency of the distribution structure of client updates within the cluster. This index is calculated based on the variance of skewness and kurtosis and is defined as follows:

[0085]

[0086] in and are the variances of the skewness and kurtosis within the cluster, respectively, and γ is a modulating factor. Benign clients exhibit higher variance in their high-order statistical features due to large differences in data distribution. However, malicious clients controlled by the same strategy tend to update more consistently, resulting in a significant decrease in variance and an increase in the gradient distribution anomaly index.

[0087] Next, calculate the update direction entropy and measure the consistency of the gradient update direction of each client in the cluster based on the information entropy principle. Let the probability distribution of the update direction angle be {p k}, then the updated direction entropy is expressed as:

[0088]

[0089] Where K represents the number of discrete bins of direction angles, {p k} is the probability of falling into the kth interval, and ∈ is a stability term. The lower the entropy value, the more concentrated the update direction is, which usually indicates highly coordinated malicious behavior.

[0090] Finally, calculate the spectrum normalization and obtain the gradient vector g for each client i in the cluster. i , and calculate its one-dimensional fast Fourier transform:

[0091]

[0092] Then normalize it to get the power spectral density distribution:

[0093]

[0094] where ε is a numerical stability constant.

[0095] Then calculate the cosine similarity between any two client spectra:

[0096]

[0097] The spectrum similarity matrix is obtained by calculating the mean μ and standard deviation σ of all spectrum similarities and constructing the spectrum consistency index:

[0098] SCI=μ·(1-σ)

[0099] This indicator encourages the coexistence of overall high consistency and structural stability, reflecting the overall synchronization degree of the client cluster in the spectrum structure.

[0100] To improve the accuracy and robustness of malicious client identification after clustering, we designed a cluster-level discrimination mechanism based on multi-feature priority judgment. This mechanism leverages three statistical features—average gradient norm, gradient distribution anomaly index, and calculated update direction entropy—as well as an auxiliary feature spectral normalization index. This mechanism comprehensively models client update behavior at the cluster level and, combined with hierarchical rules, enables accurate identification of malicious clusters. Compared to traditional heuristic methods that rely on cluster size, this mechanism exhibits greater attack adaptability and maintains effective discrimination, especially when the proportion of malicious clients is high.

[0101] Preferably, the step S6, based on the feature result analysis and judgment, classifies the clients into benign clients and malicious clients, retains the model updates of the benign clients, and removes the updates uploaded by the malicious clients, includes:

[0102] First, if the ratio of the average gradient norms of two subclusters is within a set threshold, it is considered that there is significant gradient amplification behavior. In this case, the cluster with the larger norm is directly judged as malicious.

[0103] If the norm ratio is insufficient to provide effective differentiation, the system further checks for extreme values in three characteristic metrics (gradient distribution anomaly index, computational update direction entropy, and spectral normalization metric). The system counts the number of metrics reaching the extreme value threshold within each cluster, and considers the cluster with the most extreme values malicious. If two clusters have the same number of extreme values, the system proceeds to further fine-grained analysis.

[0104] If the number of extreme values is equal or neither cluster has an extreme value, the two clusters are further evaluated for their gradient distribution anomaly index. This metric reflects the consistency of the distribution structure by measuring the skewness and kurtosis variance of the client gradients within the cluster. A higher gradient distribution anomaly index generally indicates highly similar update patterns within the cluster, indicating a potential coordinated attack. When the difference in the gradient distribution anomaly index between two clusters exceeds a set threshold, the cluster with the higher anomaly is considered malicious.

[0105] If none of the above three metrics are significantly discriminative, the computational update direction entropy is introduced as the final basis. This metric measures the uncertainty of the gradient direction distribution within a cluster. Lower computational update direction entropy values indicate more consistent gradient directions, and thus a higher likelihood of organized malicious behavior. In particular, if a cluster's computational update direction entropy is extremely low, it can be directly considered malicious; otherwise, the cluster with the lower computational update direction entropy is selected as the final judgment.

[0106] In summary, the present invention proposes a federated learning anti-poisoning method for a high proportion of malicious clients. Specifically, principal component analysis is first applied to reduce the dimensionality of the model parameters, followed by hierarchical clustering. Finally, an evaluation-based detection module is designed to accurately distinguish between malicious and benign clients. Even in complex environments with a high proportion of malicious clients, this invention can achieve robust federated learning by accurately detecting malicious clients.

[0107] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A federated learning anti-poisoning method for a high proportion of malicious clients, characterized by: The following steps are involved: 1) The server performs principal component analysis and dimensionality reduction on the model update data uploaded by the client to remove redundant information and retain the main features; 2) Based on the client model parameters after dimensionality reduction in step 1), a hierarchical agglomerative clustering algorithm is used to perform cluster analysis on the client model parameters. Divide clients into two clusters to achieve a preliminary distinction between normal and potentially malicious clients, and lay the foundation for subsequent identification and elimination of malicious clusters; 3) Based on the clustering results in step 2), perform cluster-level numerical feature analysis on the client model update, integrating statistical features with frequency domain features to identify potential malicious behaviors; 4) Based on the feature analysis results of step 3), the clients are divided into benign clients and malicious clients. The model updates of benign clients are retained, and the updates detected as uploaded by malicious clients are eliminated.

2. The method for preventing poisoning by federated learning for high-proportion malicious clients according to claim 1, characterized in that: The client model parameters after dimensionality reduction are clustered using a hierarchical agglomerative clustering algorithm to divide the clients into two clusters, thereby achieving a preliminary distinction between normal and potentially malicious clients and laying the foundation for subsequent identification and elimination of malicious clusters, including: Take each client as the initial single cluster and build a cluster label list; A bottom-up strategy is used to gradually merge the most similar clusters until a complete clustering structure is generated and a connection matrix is formed; Extract the distance and cluster size information of each merge in the connection matrix and construct a clustering tree; Starting from the root node of the clustering tree, recursively traverse each subtree and count the number of clients it contains; Determine whether each subtree meets the minimum cluster size requirement. If the number of clients in a subtree is less than the preset threshold, mark the subcluster as abnormal. The clients corresponding to the abnormal clusters are removed from the clustering structure, and the two main clusters are retained as input for subsequent malicious identification.

3. The method for preventing poisoning by federated learning for high-proportion malicious clients according to claim 1, characterized in that: Based on the clustering results, cluster-level numerical feature analysis is performed on the client model update, and statistical features are integrated with frequency domain features to identify potential malicious behaviors, including: Extract the updated gradient features of each cluster from the clustered clients; The average gradient norm, gradient distribution anomaly index, update direction entropy and spectrum consistency index of each cluster are calculated in turn. First, the average gradient norm of each cluster is calculated to measure the overall strength of the client model update in the cluster, which is defined as: Where N represents the number of clients in the cluster, g i is the gradient vector uploaded by the i-th client. This metric can reveal malicious clients' behavior of amplifying the model's influence through scaling operations during local updates. Secondly, the gradient distribution anomaly index is extracted to measure the consistency of the distribution structure of client updates within the cluster. This metric is calculated based on the variance of skewness and kurtosis and is defined as follows: in and are the variances of skewness and kurtosis within the cluster, and γ is the adjustment factor. Benign clients have large differences in data distribution, so their high-order statistical characteristics exhibit a higher variance. However, malicious clients controlled by the same strategy tend to update in a consistent manner, and the variance decreases significantly, leading to an increase in the gradient distribution anomaly index. Next, the update direction entropy is calculated to measure the consistency of the gradient update direction of each client in the cluster based on the principle of information entropy. Let the probability distribution of the update direction angle be {p k }, then the updated direction entropy is expressed as: Where K represents the number of discrete bins of direction angles, {p k } is the probability of falling into the kth interval, and ∈ is a stable term. The lower the entropy value, the more concentrated the update direction is, which usually indicates highly coordinated malicious behavior. Finally, the spectrum normalization is calculated. For each client i in the cluster, its gradient vector g is obtained. i , and calculate its one-dimensional fast Fourier transform: Then normalize it to get the power spectral density distribution: Where ε is a numerical stability constant. Then calculate the cosine similarity between any two client spectra: The spectrum similarity matrix is obtained by calculating the mean μ and standard deviation σ of all spectrum similarities and constructing the spectrum consistency index: SCI = μ·(1-σ).

4. The method for preventing poisoning by federated learning for high-proportion malicious clients according to claim 1, characterized in that: The aforementioned process of distinguishing clients into benign and malicious clients based on the feature analysis results, retaining model updates from benign clients, and removing updates uploaded by malicious clients includes: Determine whether the norm ratio of the two clusters falls within a preset threshold range. If so, the cluster with the larger norm is determined to be malicious. If the norm ratio judgment is invalid, the number of index values in each cluster that reach the extreme value threshold is counted, and the cluster with more extreme value items is judged as a malicious cluster; If the number of extreme values is the same, the gradient distribution anomaly index values are compared. If the difference exceeds the preset threshold, the cluster with the higher gradient distribution anomaly index is considered a malicious cluster; If it is still impossible to distinguish, the update direction entropy values are compared. The cluster with the lower entropy value is judged as a malicious cluster. If the update direction entropy of a cluster is extremely low, it is directly marked as a malicious cluster.

5. The method for preventing poisoning by federated learning for high-proportion malicious clients according to claim 4 is characterized in that: The updates detected as uploaded by malicious clients are eliminated, and the model update results of the benign clients screened out are subjected to model aggregation.

Citation Information

Cited By

  • Federal learning model poisoning attack defense method and system

    CN120880801A

  • Defence method and system for federal learning model poisoning attack

    CN120880801B

  • Defense method for power system poisoning attack

    CN121509115A

  • Multi-instance learning ovarian cancer classification method capable of resisting single-side tag noise

    CN121686446A

  • A multi-instance learning ovarian cancer classification method resistant to one-sided label noise

    CN121686446B