Federal learning aggregation defense method and device based on hierarchical clustering
By adopting a hierarchical clustering method in the federated learning framework, clustering and contribution evaluation of gradients is solved, and the problem of difficulty in eliminating malicious gradients in the existing technology is improved, and the defense ability of poisoning attacks is improved.
Patent Information
- Application Number
- CN202510038130.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-30
AI Technical Summary
When facing poisoning attacks, the existing federated learning framework is difficult to accurately eliminate malicious gradients, resulting in defense failure.
The hierarchical clustering method is used to cluster the gradient from bottom to top to separate the malicious gradients that are much different from the benign gradients, and the gradient is contributed by the auxiliary data set to eliminate malicious gradients similar to the benign gradients.
Improves the robustness of defense and the ability to defend against poisoning attacks, and can eliminate malicious gradients more accurately to ensure the security of the model.
Smart Images

Figure CN120069003A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence security, and particularly relates to a federated learning aggregation defense method, which can be used as a defense means for a federated learning framework to cope with poisoning attacks. Background Art
[0002] Federated learning is a distributed machine learning framework used to solve the data silo problem. Its implementation mainly includes the following steps: initializing the federated learning framework and setting model parameters; the central server sending the global model to each client node; each client node training to obtain local gradients using local data and the global model; the client uploading the trained gradients to the central server; the central server aggregating the gradients of the clients and updating the global model; repeating the above steps until the global model converges.
[0003] Poisoning attack means that the attacker designs malicious data to disrupt the model training process and cause the model to go wrong. It can be subdivided into targeted attacks and untargeted attacks. The former specifies the target class as the attack object, only causing errors in the target class and having no impact on other classes. The latter destroys the overall performance of the model, resulting in poor performance of the model on all classes and being more harmful.
[0004] The security issues of the federated learning framework have gradually attracted people's attention. As a typical attack means, researchers have proposed many defense solutions for poisoning attacks.
[0005] The patent document with the application number CN202410776571.8 discloses a robust aggregation method for federated learning based on backdoor attack defense. By analyzing the similarity of the key parameters of the federated model, it reduces the dimension of the model parameters, performs unsupervised clustering, and calculates to obtain local proxy models. It divides malicious models and benign models through the cosine distance between the model parameters after dimension reduction, ensuring the performance of outlier detection and unsupervised clustering. It trims the local proxy models through the Euclidean distance to effectively resist high-amplitude malicious backdoor attacks. Although this method improves the robustness of aggregation to a certain extent, due to ignoring the attack characteristic that the attacker may upload malicious gradients similar to benign gradients, it cannot accurately eliminate malicious gradients, resulting in the failure of defense. Summary of the Invention
[0006] The purpose of the present invention is to propose a federated learning aggregation defense method and device based on hierarchical clustering in view of the deficiencies of the above technologies, so as to improve the robustness of the defense and the ability to cope with poisoning attacks.
[0007] The technical idea for achieving the object of the present invention includes: hierarchically clustering gradients from bottom to top, separating malicious gradients that are significantly different from benign gradients, and then evaluating the contributions of gradients based on an auxiliary dataset to eliminate malicious gradients similar to benign gradients, thereby improving the defense robustness and the ability to cope with poisoning attacks.
[0008] According to the above idea, the technical solution of the present invention is as follows:
[0009] 1. A federated learning aggregation defense method based on hierarchical clustering, characterized by comprising:
[0010] Initialize the federated learning framework and set the global model parameters;
[0011] The central server sends the global model to each client node;
[0012] Each client node uses local data and the global model to train local gradients and uploads the local gradients to the central server;
[0013] In the aggregation stage of federated learning, hierarchically cluster the gradients from bottom to top to obtain a gradient structure tree;
[0014] Construct an auxiliary dataset according to the task objective of federated learning;
[0015] Use the auxiliary dataset to evaluate the contributions of the gradient structure tree;
[0016] Eliminate malicious gradients according to the gradient contribution scores, aggregate the remaining gradients and update the global model to achieve the defense against poisoning attacks.
[0017] Further, the step of hierarchically clustering the gradients from bottom to top in the aggregation stage of federated learning to obtain a gradient structure tree includes:
[0018] 2a) After the central server receives the gradient sets {g 1 , g 2 , … g i , …, g n} from n clients, for each gradient g i , calculate the similarity d j between it and the remaining gradients g ij ;
[0019] 2b) Establish a similarity set D and a gradient structure tree M, form the client set C with all clients, and store the result of step 2a) in the similarity set D;
[0020] 2c) For each client node C x in the client set C, find the client node Cy , merge C x , C y becomes a cluster node C xy , and add C x , C y as child nodes, C xy as the parent node and store them together in the gradient structure tree M. Delete the client nodes C x and C y from the client set C and insert the cluster node C xy to update the client set C;
[0021] 2d) Take the larger similarity in C x , C y as the similarity of C xy , and calculate the similarity d xy between the cluster C k and the client C xyk ;
[0022] 2e) Delete the similarities d x and C y corresponding to the client nodes C xk , d yk from the similarity set D, and insert the similarity d xy of the cluster node C xyk to update the similarity set D;
[0023] 2f) Repeat steps 2c)-2e) until all sample points can be merged into one cluster to obtain the final gradient structure tree M.
[0024] Furthermore, the contribution evaluation of the gradient structure tree using the auxiliary dataset includes:
[0025] 5a) Set the contribution difference of the root node N of the gradient structure tree M to 0;
[0026] 5b) Divide the root node N into two left and right nodes N L and N R . For all gradients of these two child nodes, perform an aggregation operation to obtain the aggregated gradient g L of the left node and the aggregated gradient g R of the right node;
[0027] 5c) Copy the global model into two copy models and add them to the central server. Update each copy model according to the aggregated gradient of one child node, and calculate the cross-entropy losses loss L and loss L of these two copy models and their difference l;
[0028] 5d) Set the contribution difference threshold δ and compare it with the difference l:
[0029] If l < δ, it is considered that there is no contribution difference between the two child nodes, that is, the contributions of the two child nodes to the global model are similar and no further division is required;
[0030] If l ≥ δ, it is considered that there is a contribution difference between the two child nodes, and they have different training objectives, and the child nodes need to be further divided;
[0031] 5e) Repeat steps 5b)-5d) until there is no contribution difference in the gradients within all nodes, and a gradient structure tree M based on gradient evaluation is obtained. Each leaf node on this tree corresponds to a gradient cluster with similar contributions, and all gradients within each cluster correspond to the same training objective.
[0032] 2. A federated learning aggregation defense device based on hierarchical clustering, comprising:
[0033] A federated initialization module for initializing the federated learning framework, including setting its global model parameters, the number of clients, and selecting a suitable aggregation algorithm;
[0034] A global model distribution module for the central server to send the global model to each client;
[0035] A training data module for each client to train the global model using local data to obtain local gradients;
[0036] An uploaded gradient module for each client to upload the local gradients to the central server;
[0037] An aggregated gradient module for the central server to deploy an aggregation defense scheme and update the global model;
[0038] Furthermore, the aggregated gradient module includes:
[0039] A coarse-grained gradient division sub-module for building a gradient structure tree, clustering gradient nodes into a cluster from bottom to top, and performing hierarchical clustering on similar gradient sample points to improve the tightness between gradient structure trees;
[0040] A fine-grained gradient evaluation sub-module for evaluating the contributions of the gradient structure tree, evaluating the nodes of the gradient structure tree by constructing an auxiliary data set, and obtaining a gradient contribution tree based on gradient contribution evaluation;
[0041] A malicious gradient elimination sub-module for finding and eliminating malicious gradient clusters with low contribution degrees, obtaining the remaining clean gradients, aggregating them, and updating the global model to achieve the purpose of defending against poisoning attacks.
[0042] The present invention has the following advantages compared with the prior art:
[0043] First, when initially dividing the gradient, the present invention uses the maximum distance method between clusters to distinguish gradient clusters, ignoring the redundant features of high-dimensional data, thus reducing the influence of outliers, achieving a tighter clustering effect, and improving the ability of clustering to adapt to heterogeneous federated learning scenarios.
[0044] Second, since the present invention introduces an auxiliary data set and distinguishes malicious gradients similar to benign gradients by evaluating gradient contributions, it can achieve more precise elimination of malicious gradients and improve the defense effect against poisoning attacks. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 is a flowchart of an embodiment of the federated learning aggregation defense method based on hierarchical clustering according to the present invention;
[0046] Figure 2 is Figure 1 a sub-flowchart of coarsely dividing gradients in
[0047] Figure 3 is Figure 1 a sub-flowchart of finely evaluating gradients in
[0048] Figure 4 a block diagram of an embodiment of the federated learning aggregation defense system based on hierarchical clustering according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0050] Embodiment 1, a federated learning aggregation defense method based on hierarchical clustering.
[0051] This example refers to the federated learning process framework and deploys a defense solution in the aggregation stage.
[0052] The federated learning framework consists of two parts: a central server and clients. Among them, the central server holds the global model and undertakes the update operation of the global model; the clients hold local data and are responsible for the training work of the model.
[0053] The local data refers to the private data of the clients, which cannot communicate with other clients, and the private data held by the clients is non-independent and identically distributed.
[0054] Referring to Figure 1 , the implementation steps of this example are as follows:
[0055] Step 1, initialize the federated framework.
[0056] Initialize the federated learning framework, which includes setting the number of clients, selecting a suitable aggregation algorithm, setting the model structure, number of training rounds, learning rate, and other parameters necessary for machine learning model training. In this example, the settings are as follows but are not limited to:
[0057] Set the total number of clients to 50, select the aggregation algorithm and global model, and set the learning rate to 0.01.
[0058] Existing aggregation algorithms include secure aggregation algorithms such as FedAvg, FedProx, SCAFFOLD, etc. In this example, the FedAvg algorithm is selected but is not limited to it;
[0059] Existing global model structures include neural network model structures such as ResNet, LeNet, VGG, etc. In the example, the ResNet network structure is selected but is not limited to it, and it is represented as the global model G.
[0060] Step 2, the central server distributes the global model.
[0061] The data of the clients cannot leave the local area and can only interact with the central server through model gradients. Considering the balance between communication and client computing power, federated learning does not communicate with all clients in each round of training, but selects a part of the clients as active clients.
[0062] As an example, the central server randomly selects 10 clients from 50 clients and sends the global model G to these 10 clients.
[0063] Step 3, the clients train the local models.
[0064] When each client receives the global model G sent by the server, it needs to train the local model. Its implementation includes the following:
[0065] 3.1) Copy the global model G as a copy model L and keep it locally;
[0066] 3.2) Input the local data into the copy model L to obtain the cross-entropy loss value, and the model will automatically adjust the values of the neurons according to the cross-entropy loss value, that is, update the neuron parameters of the copy model L;
[0067] 3.3) Repeat step 3.2). After reaching the number of local model training rounds set by the central server, obtain the trained local model L';
[0068] 3.4) After the training is completed, calculate the change between the copy model L' and the global model G, that is, the difference between the two, to obtain the local gradient g:
[0069] g = L' - G.
[0070] Step 4, the client uploads the local gradient.
[0071] Since gradients are vulnerable to being stolen or damaged by attackers during transmission, after the training is completed, the client first needs to encrypt the local gradient g. Existing gradient encryption methods include homomorphic encryption, differential privacy, and secure multi-party computation. In this example, homomorphic encryption is selected but not limited to for encrypting the gradient. Among them, homomorphic encryption algorithms include PySyft, Helib, and TenSEAL. In this example, the PySyft algorithm is used but not limited to for encrypting the gradient.
[0072] The implementation of this step includes the following:
[0073] 4.1) Each client uses the public key provided by the server to encrypt the gradient to generate the encrypted gradient, ensuring that the private key is only managed on the server or in a secure encryption area and cannot be accessed by the client, so that the encrypted gradient supports direct mathematical operations and the server can perform aggregation operations without decrypting.
[0074] 4.2) The client uploads the encrypted gradient to the central server.
[0075] Step 5, the central server makes a coarse-grained division of the local gradients to obtain the final gradient structure tree M.
[0076] Refer to Figure 2 , the implementation of this step includes the following:
[0077] 5.1) Calculate the similarity between gradients:
[0078] After the central server receives the gradient sets {g 1 , g 2 , … g i , …, g n} from n clients, it calculates the gradient similarity.
[0079] Existing similarity calculation methods include cosine distance, Euclidean distance, and Manhattan distance. In this example, the cosine distance is used to measure the similarity between gradients. The cosine distance is a direction-based distance and can ignore the influence of high-dimensional redundant data. For each gradient g i , its similarity d j with the remaining gradients g ij is calculated as follows:
[0080]
[0081] Among them, g i represents the i-th client among n clients, i ∈ (1, n); g j represents the gradient of the j-th client among the remaining n - 1 clients, j ∈ (1, n)\i; gik , g jk respectively represent the neuron components of gradient g i , g j ;
[0082] 5.2) Establish a similarity set D, form a client set C with all clients, and store the result of step 5.1) in the similarity set D. The similarity set D corresponds to the client set C, that is, each client in the client set C can find a similarity in the similarity set D;
[0083] 5.3) Establish a bottom-up generated binary tree, use the gradient of a single client as the leaf node of the binary tree, use the gradient cluster composed of several clients as its non-leaf node, and use the gradient cluster of all clients as its root node to form a gradient structure tree M;
[0084] 5.4) Merge client nodes:
[0085] For each client node C in the client set C x , find the client node C y in the similarity set D that is closest to it, merge C x , C y into a cluster node C xy , and use C x , C y as child nodes and C xy as the parent node and store them together in the gradient structure tree M. Delete the client nodes C x and C y from the client set C and insert the cluster node C xy to update the client set C;
[0086] 5.5) Calculate the similarity between clusters:
[0087] Existing methods for calculating the similarity between clusters include: complete-link method, single-link method, average-link method. In this example, the complete-link method is selected. Since this method regards the maximum similarity of the sample points in the cluster as the similarity between clusters, compared with other methods, it is less sensitive to outliers and will group closer gradients together to produce a tighter clustering, which can ensure that this example can handle heterogeneous data distributions. The specific calculation is as follows:
[0088] Take the larger similarity in the said C x and C y as the similarity d xy of C xyk :
[0089] d xyk = max(d xk , dyk )
[0090] Among them, d xk is the similarity between client C x and client C k , and d yk is the similarity between client C y and client C k ;
[0091] 5.6) Delete the client node C x and C y from the similarity set D, and the corresponding similarities d xk , d yk , and insert the similarity d xy of the cluster node C xyk to update the similarity set D;
[0092] 5.7) Repeat steps 5.2)-5.6) until all sample points can be merged into one cluster to obtain the final gradient structure tree M.
[0093] Step 6, perform a fine-grained evaluation on the central server gradient structure tree M to obtain a gradient structure tree M' after gradient evaluation.
[0094] Refer to Figure 3 , and the implementation of this step includes the following:
[0095] 6.1) Initialize the contribution difference:
[0096] Since the gradient structure tree includes m nodes including the root node, it is necessary to evaluate the contribution difference between two child nodes of the same parent node at each layer. At the same time, since the root node N is the largest parent node in the gradient structure tree M, it is necessary to process the root node N first, that is, set the contribution difference of the root node N of the gradient structure tree to 0;
[0097] 6.2) Divide the root node N into two left and right nodes N L and N R , and use the FedAvg aggregation algorithm to perform aggregation operations on all gradients of these two child nodes respectively to obtain the aggregated gradient g L of the left node and the aggregated gradient g R of the right node;
[0098] 6.3) Update the replica model:
[0099] Copy the global model into two replica models and add them to the central server, and update each replica model according to the aggregated gradient of a child node;
[0100] Construct a small auxiliary dataset that is similar in average size to a single client and similar to the overall client data distribution, and input this auxiliary dataset into the two replica models to obtain the cross-entropy loss loss L and loss R , and calculate the difference l:
[0101] l = |loss L - loss R |;
[0102] 6.4) Evaluate the contribution difference:
[0103] Set the contribution difference threshold δ, and compare it with the difference l:
[0104] If l < δ, it is considered that there is no contribution difference between the two child nodes, that is, the contributions of the two child nodes to the global model are similar and no further division is required;
[0105] If l ≥ δ, it is considered that there is a contribution difference between the two child nodes, and they have different training objectives, and the child nodes need to be further divided.
[0106] In the example, the threshold is set to 0.01. When the threshold is less than the change in the loss value at model convergence, the model contribution difference can be ignored.
[0107] 6.5) Repeat steps 6.2)-6.4) until there is no contribution difference in the gradients within all nodes, and obtain a gradient structure tree M' based on gradient evaluation. Each leaf node on this tree corresponds to a gradient cluster with similar contributions, and all gradients within each cluster correspond to the same training objective.
[0108] Step 7, aggregate the clean gradient clusters of the gradient structure tree M' and update the global model.
[0109] 7.1) Calculate the contribution score s i of each cluster g i in the gradient structure tree M':
[0110]
[0111] where Test() represents the function for testing the cross-entropy loss of the model, and g k is the k-th gradient cluster among the m gradient clusters in the gradient structure tree M';
[0112] 7.2) Calculate the contribution score s mal of the malicious gradient cluster g mal :
[0113]
[0114] 7.3) Remove the gradient cluster g mal with a contribution score of s mal , obtaining the remaining m - 1 clean gradients;
[0115] 7.4) Update the global model:
[0116] Use the Federated Averaging aggregation algorithm FedAvg to aggregate the remaining m - 1 clean gradients and update the model, obtaining the updated global model G':
[0117]
[0118] Step 8, determine whether the global model G' converges:
[0119] If the model loss value is still changing significantly, in this example, the change in the loss value is greater than 0.01, then it is considered that the model has not converged, and return to step 2 to continue execution;
[0120] Otherwise, it is considered that the model has converged, the federated learning training ends, obtaining a clean global model, achieving the purpose of successfully defending against poisoning attacks.
[0121] Example 2, Federated Learning Aggregation Defense Device Based on Hierarchical Clustering
[0122] Refer to Figure 4 , this example includes: Federated Initialization Module 1, Global Model Distribution Module 2, Training Data Module 3, Gradient Upload Module 4, Aggregate Gradient Module 5. Its working principle is as follows:
[0123] The Federated Initialization Module 1 is used to initialize the federated learning framework, including setting its global model parameters, the number of clients, and selecting a suitable aggregation algorithm;
[0124] The central server uses the Global Model Distribution Module 2 to randomly select several clients and send the global model of the central server to these clients;
[0125] After receiving the global model, the client uses the Training Data Module 3 to train a copy of the global model using local data to obtain local gradients; and input the local gradients into the Gradient Upload Module 4, and upload the local gradients to the central server through encryption means;
[0126] After receiving the output of the Gradient Upload Module 4, the central server inputs it into the Aggregate Gradient Module 5, deploys the aggregation defense scheme and updates the global model.
[0127] The aggregation gradient module 5 includes a coarse-grained division gradient sub-module 51, a fine-grained evaluation gradient sub-module 52, and a malicious gradient elimination sub-module 53. The upload gradient module 4 inputs the client gradient into the coarse-grained division gradient sub-module 51 for establishing a gradient structure tree. By clustering gradient nodes into a cluster from bottom to top and hierarchically clustering similar gradient sample points, a gradient structure tree with a tight internal structure is obtained. The coarse-grained division sub-module 51 inputs the gradient structure tree into the fine-grained evaluation gradient sub-module 52, and evaluates the nodes of the gradient structure tree by constructing an auxiliary data set to obtain a gradient structure tree based on contribution evaluation. The malicious gradient elimination sub-module 53 searches for and eliminates malicious gradient clusters with low contribution degrees in the gradient structure tree based on contribution evaluation, obtains the remaining clean gradients, aggregates them, and updates the global model to achieve the purpose of defending against poisoning attacks.
[0128] The above description is only several specific examples of the present invention and does not constitute any limitation to the present invention. Obviously, for professionals in the field, after understanding the content and principle of the present invention, various modifications and changes in form and details may be made without departing from the principle and structure of the present invention. However, these corrections and changes based on the idea of the present invention are still within the protection scope of the claims of the present invention.
[0129] It should be noted that the step numbers in the specification and claims of the present invention are only for a clear description of the implementation embodiments of the present invention for easy understanding, and their sequence numbers are not limited.
Claims
1. A federated learning aggregation defense method based on hierarchical clustering, characterized in that: include: Initialize the federated learning framework and set global model parameters; The central server sends the global model to each client node; Each client node uses local data and global model training to obtain local gradients, and uploads the local gradients to the central server; In the aggregation phase of federated learning, the gradient structure tree is obtained by performing hierarchical clustering of gradients from bottom to top; Construct auxiliary datasets based on the task objectives of federated learning; Use auxiliary datasets to evaluate the contribution of the gradient structure tree; Malicious gradients are removed according to the gradient contribution score, the remaining gradients are aggregated and the global model is updated to defend against poisoning attacks.
2. The method according to claim 1, characterized in that: In the aggregation stage of federated learning, the gradients are hierarchically clustered from bottom to top to obtain a gradient structure tree, including: 2a) The central server receives the gradient set {g1,g2,…g i ,…,g n }, for each gradient g i , calculate its difference with the residual gradient g j The similarity between ij ; 2b) Establish a similarity set D and a gradient structure tree M, form all clients into a client set C, and store the result of step 2a) in the similarity set D; 2c) For each client node C in the client set C x , find the client node C that is closest to it from the similarity set D y , merge C x , C y Become a cluster node C xy , and C x , C y As a child node, C xy Store them together as parent nodes in the gradient structure tree M, and delete the client node C from the client set C x and C y And insert cluster node C xy To update the client set C; 2d) Take C x , C y The larger similarity is C xy , calculate the similarity of cluster C xy With client C k The similarity d xyk ; 2e) Delete client node C from similarity set D x and C y The corresponding similarity d xk d yk , and insert cluster node C xy The similarity d xyk To update the similarity set D; 2f) Repeat steps 2c)-2e) until all sample points can be merged into a cluster to obtain the final gradient structure tree M.
3. The method according to claim 2, characterized in that: The calculation is the same as the residual gradient g j The similarity between ij , and the calculation cluster C xy With client C k The similarity d xyk , the formulas are as follows: d xyk =max(d xk ,d yk ) Among them, g i represents the i-th client among n clients, i∈(1,n); g j represents the gradient of the jth client among the remaining n-1 clients, j∈(1,n)\i; g ik , g jk Represent the gradient g i , g j The neuronal component of xk It is client C x With client C k The similarity, d yk It is client C y With client C k The similarity.
4. The method according to claim 1, characterized in that: The construction of the auxiliary data set according to the task objective of federated learning is to establish an auxiliary data set with a similar distribution to the entire sample set through task analysis of federated learning, and the size of the auxiliary data set is consistent with the average size of the data set held by the client.
5. The method according to claim 1, characterized in that: The use of the auxiliary data set to evaluate the contribution of the gradient structure tree includes: 5a) Setting the contribution difference of the root node N of the gradient structure tree M to 0; 5b) Split the root node N into two nodes N on the left and right L and N R , perform aggregation operations on all the gradients of these two child nodes respectively, and obtain the aggregated gradient g of the left node L and the aggregate gradient g of the right node R ; 5c) Copy the global model into two replica models and add them to the central server. Update each replica model according to the aggregated gradient of a child node, and use the constructed auxiliary dataset to calculate the cross entropy loss of the two replica models respectively. L and loss R and its difference l; 5d) Set the contribution difference threshold δ and compare it with the difference l: If l<δ, it is considered that there is no difference in contribution between the two child nodes, that is, the two child nodes contribute similarly to the global model and no further division is required; If l ≥ δ, it is considered that there is a contribution difference between the two sub-nodes, and the two have different training objectives, and the sub-nodes need to be further divided; 5e) Repeat steps 5b)-5d) until there is no difference in the contribution of the gradients in all nodes, and obtain a gradient structure tree M based on gradient evaluation. Each leaf node on the tree corresponds to a gradient cluster with similar contribution, and all gradients in each cluster correspond to the same training target.
6. The method according to claim 1, characterized in that: The method removes malicious gradients according to the gradient contribution score, aggregates the remaining gradients and updates the global model, including: 6a) Calculate the gradient structure M for each cluster g i Contribution score i : Among them, Test() represents the function of testing the cross entropy loss of the model, g k is the kth gradient in the gradient structure tree M; 6b) Eliminate the cluster with the smallest contribution score: 6c) Aggregate the gradients using the FedAvg algorithm and update the model.
7. The method according to claim 2, characterized in that: Step 2b) The gradient structure tree M is a bottom-up generated binary tree, in which a leaf node is a gradient of a single client and a non-leaf node is a gradient cluster composed of several clients.
8. The method according to claim 1, characterized in that: Initializing the federated learning framework and setting model parameters, including setting the number of clients, selecting an appropriate aggregation algorithm, setting the model structure, training rounds, learning rate, and other parameters necessary for machine learning model training; The local data refers to the private data of the client, which cannot be communicated with other clients, and the private data held by the client is not independent and identically distributed.
9. A federated learning aggregation defense device based on hierarchical clustering, comprising: The federated initialization module is used to initialize the federated learning framework, including setting its global model parameters, the number of clients, and selecting a suitable aggregation algorithm; The global model distribution module is used by the central server to send the global model to each client; The training data module is used by each client to train the global model using local data to obtain local gradients; Upload gradient module, used by each client to upload local gradient to the central server; Aggregate gradient module, used by the central server to deploy aggregate defense solutions and update the global model.
10. The device according to claim 9, characterized in that The aggregation gradient module comprises: The coarse-grained gradient submodule is used to establish a gradient structure tree. It clusters gradient nodes into a cluster from bottom to top and hierarchically clusters similar gradient sample points to improve the compactness between gradient structure trees. The fine-grained evaluation gradient submodule is used for the contribution evaluation of the gradient structure tree. It evaluates the nodes of the gradient structure tree by constructing an auxiliary data set to obtain a gradient contribution tree based on the gradient contribution evaluation. The malicious gradient elimination submodule is used to find and eliminate malicious gradient clusters with low contribution, obtain the remaining clean gradients, aggregate them and update the global model to achieve the purpose of defending against poisoning attacks.
Citation Information
Patent Citations
Federal learning robust aggregation method based on backdoor attack defense
CN118965415A
Cited By
Federal learning-oriented label perception distribution dynamic weighted backdoor defense method
CN122475940A