Federal learning robustness aggregation method based on lens detection spectral clustering
By using the lens detection spectral clustering method to identify and eliminate malicious clients in federated learning, the problem that traditional robustness algorithms are unable to cope with Byzantine attacks is solved, ensuring the normal convergence of the model and improving the robustness of model training.
Patent Information
- Application Number
- CN202510755035.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-10-03
AI Technical Summary
Traditional robustness algorithms cannot cope with the problem that the model cannot converge when the Euclidean distance updated by the Byzantine attacker is hidden in the benign client.
Through the lens detection spectral clustering method, the Gaussian kernel function is used to map the client parameters to the infinite-dimensional Hilbert space, construct a similarity matrix, eliminate malicious parameters, form a new similarity matrix, and calculate the Laplace matrix through the adjacency matrix and degree matrix. The normalized cut method is used to cluster and identify benign update clusters, and clusters with a larger number are selected as benign update clusters for model parameter update.
Effectively identify and eliminate malicious clients in federated learning, ensure normal model convergence, improve the accuracy of malicious client identification, and enhance the robustness of model training.
Smart Images

Figure CN120745872A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a robust aggregation method for federated learning based on shot detection spectral clustering, belonging to the technical field of federated learning attack defense. Background Art
[0002] In the current era of rapid development of general-purpose large-scale AI models, coupled with the advancement of hardware and iteration of algorithms, large models are consuming data at an ever-increasing rate. Statistics show that the global stock and consumption curves of publicly available data (including open source datasets, web pages, and books) available for large-scale model pre-training will intersect in 2028. This means that public data available for large-scale model training will be depleted by 2028. High-quality private data is crucial for large-scale model training. To ensure the security of private data while ensuring high-quality data for large-scale model training, federated learning, with its unique characteristics of "data remains static, model remains dynamic, and data is available but invisible," can effectively address the data silo problem.
[0003] During federated learning, each client simultaneously trains its local dataset using a global model broadcast by the server. After several rounds of stochastic gradient descent, the updated parameters are transmitted back to the server. Upon receiving the updated parameters, the server executes a parameter aggregation algorithm to update the global model. These steps are then repeated until the model converges. However, due to its distributed nature, malicious attackers may manipulate clients to upload carefully crafted malicious updates to disrupt or prevent model convergence. This tactic is known as a Byzantine attack. This attack has been shown to have a significant impact on federated learning, as even a single Byzantine device can prevent the model from convergence.
[0004] To combat Byzantine attacks, extensive research has focused on robust aggregation schemes for federated learning. Defense methods such as Krum, Median, and Centered Clipping, which use Euclidean distance and median to detect outliers, can effectively detect Byzantine attacks and mitigate the impact of malicious attacks on the model. However, when the Euclidean distance updated by a Byzantine attacker is hidden in a benign client, the defense algorithm becomes ineffective, preventing the model from converging. Summary of the Invention
[0005] The purpose of the present invention is to provide a robust aggregation method for federated learning based on shot detection spectral clustering, aiming to solve the technical problem that traditional robustness algorithms cannot cope with the same-value attack, that is, the Euclidean distance updated by the Byzantine attacker loses its defensive effect when hidden in the benign client, resulting in the failure of the model to converge.
[0006] To achieve the above objectives, the present invention provides a robust aggregation method for federated learning based on shot detection spectral clustering. When faced with identical attacks that traditional robustness algorithms cannot cope with, this method traverses the similarity matrix to find unevenly distributed points in the infinite-dimensional Hilbert space mapped by the Gaussian kernel, thereby identifying identical attack clients and ensuring normal model convergence. The method includes the following steps:
[0007] Step 1: In the aggregated parameters uploaded by the federated learning client, the uploaded model parameters are treated as a point in the data space. A similarity matrix is constructed using the Gaussian kernel function. The similarity matrix is then traversed to remove malicious parameters and form a new similarity matrix.
[0008] Step 2: Construct the adjacency matrix and degree matrix based on the new similarity matrix to obtain the Laplace matrix, which is then normalized and cut using the normalized cutting method to obtain two clusters.
[0009] Step 3: Calculate the number of clients in the two clusters, select the cluster with more clients as the benign update cluster, and use the mean of all client updates in the benign update cluster as the final update parameter to upload to the server for model parameter update.
[0010] The step 1 specifically includes the following steps:
[0011] Step 1.1: Treat the updated parameters uploaded by multiple clients of federated learning as a point in the data space, and then construct an undirected weighted graph for multiple clients ,in is the parameter uploaded by the i-th client, is the weight between the i-th and j-th clients;
[0012] Step 1.2: Construct similarity matrix by Gaussian kernel function ,in is the similarity between clients i and j, is the bandwidth parameter, which controls the radial range of the Gaussian kernel function;
[0013] Step 1.3: Traverse the similarity matrix S and set the threshold ,when Value greater than When , it means that the current two clients are closely distributed in the space. The i, j clients are regarded as malicious parameters and removed from S. Finally, a new similarity matrix is formed. .
[0014] The step 2 specifically includes the following steps:
[0015] Step 2.1: Based on the new similarity matrix Construct the adjacency matrix W and transform the similarity matrix The similarity between each element in As the weight between each point ,get ,in ;
[0016] Step 2.2: Assign connectivity from each point to other points based on the weights between the points and calculate the degree of each point , find the degree matrix ;
[0017] Step 2.3: Calculate the Laplacian matrix based on the degree matrix and adjacency matrix ,in is the Laplace matrix element with row i and column j;
[0018] Step 2.4: Normalize the Laplacian matrix , and obtain the normalized Laplace matrix ;
[0019] Step 2.5: Using the normalized cut method, find the eigenvectors corresponding to the smallest k eigenvalues of the normalized Laplacian matrix;
[0020] Step 2.6: After solving the eigenvectors corresponding to the two smallest eigenvalues of the normalized Laplace matrix, the corresponding eigenvectors are The matrix is normalized by row, and finally formed dimensional feature matrix F;
[0021] Step 2.7: Treat each row in the feature matrix F as a 2D sample, with a total of n samples, and cluster them to obtain two clusters , .
[0022] The step 3 specifically includes the following steps:
[0023] Step 3.1: Calculation , The number of clients in each of the two clusters, select the cluster with more clients as the benign update cluster, and obtain the update weights corresponding to all clients in the benign update cluster ;
[0024] Step 3.2: Calculate the mean of the update weights in the benign update cluster And upload to the server;
[0025] Step 3.3: The server receives the mean Then update the model parameters , and then As the model parameters of a new round of iterations, it is distributed to all clients, where is the next round of global model parameters, is the global model parameter of the current round, is the learning rate.
[0026] The beneficial effects of the present invention are:
[0027] (1) The robust aggregation method for federated learning based on shot detection spectral clustering proposed in this paper can ensure the normal convergence of the model during the entire federated learning training process by identifying malicious clients in federated learning;
[0028] (2) The present invention uses a lens detection method to capture the characteristics of the distribution of update parameters in the data space from local to global, which can well identify the same-value attacks that are hidden in benign clients and cause distribution anomalies;
[0029] (3) The present invention uses the Gaussian kernel function mapping method to map the parameters uploaded by the client into the infinite-dimensional Hilbert space, thereby enhancing the distinguishability of the high-dimensional features of the data and improving the accuracy of malicious client identification and elimination. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a schematic flow chart of the present invention;
[0031] Figure 2 This is a schematic diagram of the present invention eliminating malicious parameters during the federated learning process. DETAILED DESCRIPTION
[0032] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0033] The diagrams provided in the following examples and the settings of specific parameter values in the model are mainly for illustrating the basic concept of the present invention and for simulation verification of the present invention. In specific application environments, appropriate adjustments can be made according to actual scenarios and needs.
[0034] Example 1: A robust aggregation method for federated learning based on shot detection spectral clustering, such as Figure 1 Shown, including:
[0035] Step 1: In the aggregated parameters uploaded by the federated learning client, the uploaded model parameters are regarded as a point in the data space. A similarity matrix is constructed using the Gaussian kernel function. The similarity matrix is traversed to remove malicious parameters and form a new similarity matrix.
[0036] Step 1.1: Treat the updated parameters uploaded by multiple clients of federated learning as a point in the data space, and then construct an undirected weighted graph for multiple clients ,in is the parameter uploaded by the i-th client, is the weight between the i-th and j-th clients;
[0037] Step 1.2: Construct similarity matrix by Gaussian kernel function ,in is the similarity between clients i and j, is the bandwidth parameter, which controls the radial range of the Gaussian kernel function;
[0038] Step 1.3: Traverse the similarity matrix S and set the threshold ,when Value greater than When , it means that the current two clients are closely distributed in the space. The i, j clients are regarded as malicious parameters and removed from S. Finally, a new similarity matrix is formed. .
[0039] Specifically, in this embodiment, Figure 2 As shown, there are 5 clients participating in the training in federated learning, and the global model parameters at this time are for , learning rate is 0.1:
[0040] The uploaded model parameters are
[0041] The uploaded model parameters are
[0042] The uploaded model parameters are
[0043] The uploaded model parameters are
[0044] The uploaded model parameters are
[0045] Furthermore, the Gaussian kernel function is used to construct the similarity matrix, specifically:
[0046]
[0047] Furthermore, in this embodiment =0.5, =0.9995.
[0048] Step 2: Construct the adjacency matrix and degree matrix based on the new similarity matrix to obtain the Laplace matrix, which is then normalized and cut using the normalized cutting method to obtain two clusters.
[0049] Step 2.1: Based on the new similarity matrix Construct the adjacency matrix W and transform the similarity matrix The similarity between each element in As the weight between each point ,get ,in ;
[0050] Step 2.2: Assign connectivity from each point to other points based on the weights between the points and calculate the degree of each point , find the degree matrix ;
[0051] Step 2.3: Calculate the Laplacian matrix based on the degree matrix and adjacency matrix ,in is the Laplace matrix element with row i and column j;
[0052] Step 2.4: Normalize the Laplacian matrix , and obtain the normalized Laplace matrix ;
[0053] Step 2.5: Using the normalized cut method, find the eigenvectors corresponding to the smallest k eigenvalues of the normalized Laplacian matrix;
[0054] Step 2.6: After solving the eigenvectors corresponding to the two smallest eigenvalues of the normalized Laplace matrix, the corresponding eigenvectors are The matrix is normalized by row, and finally formed dimensional feature matrix F;
[0055] Step 2.7: Treat each row in the feature matrix F as a 2D sample, with a total of n samples, and cluster them to obtain two clusters , .
[0056] Specifically, in this embodiment, all elements except the main diagonal do not exceed the threshold value. , so the similarity matrix S does not change, and we get W=S. Then calculate the degree of each point , find the degree matrix D:
[0057]
[0058] Furthermore, the Laplace matrix L is obtained based on W and D:
[0059]
[0060] Furthermore, L is normalized:
[0061]
[0062] Furthermore, the eigenvectors corresponding to the two smallest eigenvalues of the normalized Laplace matrix are obtained:
[0063]
[0064] Furthermore, the corresponding eigenvectors The matrix is normalized by row, and finally formed The characteristic matrix F of dimension:
[0065]
[0066] Furthermore, the feature matrix F is clustered to obtain two clusters , In this embodiment, the clustering process uses the k-means algorithm.
[0067] Step 3: Calculate the number of clients in the two clusters, select the cluster with more clients as the benign update cluster, and use the mean of all client updates in the benign update cluster as the final update parameter to upload to the server for model parameter update.
[0068] Step 3.1: Calculation , The number of clients in each of the two clusters, select the cluster with more clients as the benign update cluster, and obtain the update weights corresponding to all clients in the benign update cluster ;
[0069] Step 3.2: Calculate the mean of the update weights in the benign update cluster And upload to the server;
[0070] Step 3.3: The server receives the mean Then update the model parameters = , and then As the model parameters of a new round of iterations, it is distributed to all clients, where is the next round of global model parameters, is the global model parameter of the current round, is the learning rate.
[0071] Specifically, in this embodiment, There are two clients in There are three clients in The number of clients is greater than ,choose As a benign update cluster, according to In the cluster Client side, calculates the mean as the update parameter ;
[0072] Furthermore, the server receives the updated parameters And update the global model parameters:
[0073]
[0074] .
[0075] The above describes the specific embodiments of the present invention in detail with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Various changes can be made within the knowledge of ordinary technicians in this field without departing from the scope of the present invention.
Claims
1. A robust aggregation method for federated learning based on shot detection spectral clustering, characterized by: The following steps are involved: Step 1: In the aggregated parameters uploaded by the federated learning client, the uploaded model parameters are treated as a point in the data space. A similarity matrix is constructed using the Gaussian kernel function. The similarity matrix is then traversed to remove malicious parameters and form a new similarity matrix. Step 2: Construct the adjacency matrix and degree matrix based on the new similarity matrix to obtain the Laplace matrix, which is then normalized and cut using the normalized cutting method to obtain two clusters. Step 3: Calculate the number of clients in the two clusters, select the cluster with more clients as the benign update cluster, and use the mean of all client updates in the benign update cluster as the final update parameter to upload to the server for model parameter update.
2. The robust aggregation method based on shot detection spectral clustering federated learning according to claim 1 is characterized in that: The step 1 specifically includes the following steps: Step 1.1: Treat the updated parameters uploaded by multiple clients of federated learning as a point in the data space, and then construct an undirected weighted graph for multiple clients ,in is the parameter uploaded by the i-th client, is the weight between the i-th and j-th clients; Step 1.2: Construct similarity matrix by Gaussian kernel function ,in is the similarity between clients i and j, is the bandwidth parameter, which controls the radial range of the Gaussian kernel function; Step 1.3: Traverse the similarity matrix S and set the threshold ,when Value greater than When , it means that the current two clients are closely distributed in the space. The i, j clients are regarded as malicious parameters and removed from S. Finally, a new similarity matrix is formed. .
3. The robust aggregation method based on shot detection spectral clustering federated learning according to claim 1 is characterized in that: The step 2 specifically includes the following steps: Step 2.1: Based on the new similarity matrix Construct the adjacency matrix W and transform the similarity matrix The similarity between each element in As the weight between each point ,get ,in ; Step 2.2: Assign connectivity from each point to other points based on the weights between the points and calculate the degree of each point , find the degree matrix ; Step 2.3: Calculate the Laplacian matrix based on the degree matrix and adjacency matrix ,in is the Laplace matrix element with row i and column j; Step 2.4: Normalize the Laplacian matrix , and obtain the normalized Laplace matrix ; Step 2.5: Using the normalized cut method, find the eigenvectors corresponding to the smallest k eigenvalues of the normalized Laplacian matrix; Step 2.6: After solving the eigenvectors corresponding to the two smallest eigenvalues of the normalized Laplace matrix, the corresponding eigenvectors are The matrix is normalized by row, and finally formed dimensional feature matrix F; Step 2.7: Treat each row in the feature matrix F as a 2D sample, with a total of n samples, and cluster them to obtain two clusters , .
4. The robust aggregation method based on shot detection spectral clustering federated learning according to claim 1 is characterized in that: The step 3 specifically includes the following steps: Step 3.1: Calculation , The number of clients in each of the two clusters, select the cluster with more clients as the benign update cluster, and obtain the update weights corresponding to all clients in the benign update cluster ; Step 3.2: Calculate the mean of the update weights in the benign update cluster And upload to the server; Step 3.3: The server receives the mean Then update the model parameters , and then As the model parameters of a new round of iterations, it is distributed to all clients, where is the next round of global model parameters, is the global model parameter of the current round, is the learning rate.