Heterogeneous distributed robust learning gradient aggregation method combining SVD and K-means
By combining the comprehensive scoring mechanism of SVD and K-means, identifying and eliminating Byzantine nodes in heterogeneous distributed machine learning environments, the problem of difficulty in effectively identifying and eliminating Byzantine nodes in the existing technology is solved, and the robustness and performance of the system are improved.
Patent Information
- Application Number
- CN202510266896.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-07
AI Technical Summary
In a heterogeneous distributed machine learning environment, it is difficult for the existing technology to effectively identify and eliminate Byzantine nodes, resulting in information pollution during model training and reducing the overall performance of the system.
Combining singular value decomposition (SVD) and K-means clustering algorithms, a comprehensive scoring mechanism is designed to evaluate the contribution of each node through the combination of SVD scores and K-means scores, and identify and eliminate Byzantine nodes.
It significantly improves the Byzantine robustness of distributed learning systems under heterogeneous data distribution, ensures the accuracy and stability of the model, and can effectively resist Byzantine attacks in complex data environments.
Smart Images

Figure CN120106247A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of distributed machine learning, and more specifically, relates to a heterogeneous distributed robust learning gradient aggregation method combining SVD and K-means. Background Art
[0002] As data scale and complexity continue to increase, traditional stand-alone learning methods can no longer meet increasingly complex computing needs. Distributed machine learning distributes data and computing tasks to multiple nodes, which not only improves computing efficiency, but also speeds up the model training process, and is highly scalable and fault-tolerant. This has led to the widespread application of distributed machine learning in Internet services, financial analysis, scientific research and other fields.
[0003] However, data in practical applications usually presents heterogeneous distribution, that is, there are significant differences in the data distribution owned by each node. This heterogeneity exacerbates the difficulty of identifying Byzantine nodes, which affect the accuracy of the global model by uploading malicious or erroneous information. Due to the existence of heterogeneous data distribution, traditional aggregation methods often have difficulty effectively distinguishing normal nodes from malicious nodes when facing Byzantine attacks, resulting in information pollution during model training, thereby reducing the overall performance of the system. Therefore, designing effective methods to resist Byzantine attacks and achieve Byzantine robustness is an important research direction in heterogeneous distributed machine learning.
[0004] In order to solve the problem of Byzantine robustness under heterogeneous data distribution, existing studies have proposed a variety of methods, but they still have certain limitations. The NBS method only considers the gradient norm information and ignores other important features, which may not effectively identify malicious nodes in complex data environments. The Krum method aggregates by selecting updates similar to other gradients, but it is difficult to effectively distinguish normal nodes when the proportion of malicious nodes is high, and its performance will be significantly reduced when the data is highly heterogeneous. The Bucketing method divides the nodes into multiple buckets for independent aggregation. Although it can reduce the impact of malicious nodes in a single bucket, if the data distribution in the bucket is too special, the model may not be well generalized to the entire data set. Although the objective function combined with the regularization term can enhance the model's resistance to outliers, the selection and adjustment of the regularization term increases the complexity of model training and may cause the model's performance to be unstable under different data distributions. Existing scoring mechanism-based methods such as ByGARS and GAS each have some shortcomings in improving Byzantine robustness under heterogeneous data distribution. The ByGARS method uses reputation scores to aggregate gradients. Although it can effectively deal with any number of Byzantine nodes, it relies on additional auxiliary data sets to calculate reputation scores, which may lead to performance degradation when data is scarce or unbalanced. In addition, the GAS method aggregates the gradient by dividing it into multiple sub-vectors and calculates the sum of the recognition scores of each sub-vector to eliminate malicious updates, which may cause the updates of normal nodes to be misjudged, especially when the data distribution is extremely uneven. Overall, the above methods have difficulty in effectively identifying and eliminating Byzantine nodes when processing complex heterogeneous data, resulting in reduced robustness and performance of the model in heterogeneous distributed learning environments, and difficulty in dealing with Byzantine attacks in heterogeneous environments. Therefore, in heterogeneous distributed machine learning, designing a method that can effectively identify and eliminate Byzantine nodes is still a key issue that needs to be solved.
[0005] Chinese patent document CN111967015A discloses a defense proxy method for improving the Byzantine robustness of a distributed learning system. The invention uses an adaptive credibility evaluation module based on a neural network structure to dynamically evaluate the credibility of each submitted gradient, update the global classifier parameters maintained on the current master node, generate a reward signal, and adjust the parameters of the adaptive credibility evaluation module under the framework of reinforcement learning according to the reward signal; dynamically adjust the feasibility evaluation value of each working node during the training process to alleviate the impact of the tampered gradient submitted by the malicious working node on the system training process, so as to improve the Byzantine robustness of the distributed learning system. However, the patent relies on additional auxiliary data sets to calculate the reputation score, which may lead to performance degradation in the case of scarce or unbalanced data, especially in the environment of heterogeneous data distribution, and it is difficult to ensure the accuracy of the evaluation. Chinese patent document CN111898763A discloses a robust Byzantine fault-tolerant distributed gradient descent algorithm, which includes the following steps: Step 1: Initialize the structure and hyperparameters of the model to be trained; Step 2: Each working node in the parameter server framework calculates the local gradient according to the gradient descent method and sends it to the parameter server; Step 3: The parameter server first uses the multi-Krum aggregation algorithm to aggregate the local gradient at the beginning of the training, and after the global model is trained to a certain effect, the parameter server uses the Acc-based aggregation algorithm to aggregate the local gradient; Step 4: The parameter server updates the global model according to the aggregated gradient, and sends the global model to the working node for the next iteration. When the number of training iterations reaches the set threshold T, the training is stopped. However, the patent mainly relies on the multi-Krum and Acc-based algorithms for gradient aggregation. These two methods are mainly based on the distance or norm between gradients for screening, lack a comprehensive evaluation of gradient features (such as gradient direction, singular values, etc.), and it is difficult to effectively identify and eliminate carefully designed Byzantine attacks.
[0006] In view of this, the present invention designs a heterogeneous distributed robust learning gradient aggregation method combining SVD and K-means. Summary of the invention
[0007] The present invention aims to overcome at least one defect of the above-mentioned prior art and provide a heterogeneous distributed robust learning gradient aggregation method combining SVD and K-means, which can effectively identify and eliminate Byzantine nodes in a heterogeneous distributed machine learning system, thereby improving the performance of the heterogeneous distributed machine learning system.
[0008] The detailed technical scheme of the present invention is as follows:
[0009] A heterogeneous distributed robust learning gradient aggregation method combining SVD and K-means, the method comprising:
[0010] S1. Build a distributed learning system with a parameter server and n nodes, where the nodes are divided into honest nodes and Byzantine nodes;
[0011] S2. The honest node extracts some data samples from the local classification data set and calculates the gradient using the stochastic gradient descent algorithm, and uploads the calculated gradient information to the parameter server;
[0012] S3. The parameter server performs singular value decomposition on the gradient matrices submitted by all nodes, extracts singular values and corresponding eigenvectors, and calculates the SVD score of the gradient to evaluate the effectiveness and importance of each gradient, providing a basis for subsequent aggregation steps.
[0013] S4. The parameter server uses the K-means clustering algorithm to cluster the gradients of all nodes, obtain k cluster centers, and record the distance between each node and its cluster center. Based on the distance, the K-means score of the gradient is calculated for subsequent comprehensive scoring and gradient aggregation.
[0014] S5. Combine the SVD score and K-means score to calculate the comprehensive score of each node, and select nf gradients with the smallest comprehensive score for average aggregation to obtain the final global gradient. The aggregated global gradient is used to update the global model parameters.
[0015] Preferably, according to the present invention, the specific steps of step S1 are as follows:
[0016] Construct a distributed learning system with a parameter server and n nodes; the n nodes include G honest nodes and nG Byzantine nodes, and the Byzantine nodes send false information to the parameter server to disrupt the normal operation of the system;
[0017] Next, we will optimize the system. The optimization goal is to find the parameter x that minimizes the expected value of the average loss function of all honest nodes. * :
[0018]
[0019] In formula (1), x * is the final goal of system optimization and the final parameters of the model after training. It represents the model parameters that perform best on the data of honest nodes; i represents the local data set of node i, and the data distribution among nodes is heterogeneous; F i Represents node i about its local data set ξ i The loss function of i Represents local dataset ξ i Calculate the loss function F iThe expected value of , x represents the model parameters, which are used to calculate the predicted value of the model and are continuously updated through the optimization algorithm to minimize the loss function.
[0020] Preferably, according to the present invention, the specific steps of step S2 are as follows:
[0021] S21, honest node i first receives the current global model parameter x from the parameter server t ;
[0022] S22, honest node i from its local classification data set ξ i Extract some data samples And calculate the global model parameter x t Under local stochastic gradient
[0023]
[0024] In formula (2), t represents the tth iteration, x t Represents the current global model parameters;
[0025] S23, honest node i calculates the local stochastic gradient Upload to the parameter server, and the Byzantine node will send malicious gradients to the parameter server.
[0026] Preferably, according to the present invention, the specific steps of step S3 are as follows:
[0027] S31. After receiving the gradients uploaded by all nodes, the parameter server constructs all the gradients into a matrix M:
[0028]
[0029] In formula (3), represents the random gradient sent by node n at the tth iteration;
[0030] S32, perform singular value decomposition on the gradient matrix M to obtain:
[0031] M=UΣV T (4)
[0033] In formula (4), U is the left singular vector matrix, Σ is the diagonal matrix containing singular values, V T is the right singular vector matrix;
[0034] S33. Extract the singular value σ corresponding to node i from the diagonal matrix Σ i , the singular value reflects the importance of node i. Subsequently, the SVD score of each node is calculated based on the extracted singular value:
[0035]
[0036] In formula (5), represents the SVD score of node i, σ i represents the singular value corresponding to node i, σ j represents the singular value corresponding to node j, n represents node n, Represents the sum of the singular values corresponding to all nodes. Through the above steps, the contribution of each node can be effectively evaluated, providing a basis for the subsequent calculation of the comprehensive score and aggregation process.
[0037] Preferably, according to the present invention, the specific steps of step S4 are as follows:
[0038] S41, the parameter server uses the K-means clustering algorithm to cluster the gradients of all nodes and randomly selects k cluster centers c q , where q∈{1,2,…,k};
[0039] S42, assign the gradient of each node to the cluster with the nearest cluster center, and record the distance between each node and its corresponding cluster center:
[0040]
[0041] In formula (6), d i Represents the cluster center c of node i and its cluster q The distance between represents the gradient sent by node i at the tth iteration, c q represents the cluster center of the cluster to which node i belongs;
[0042] S43, based on distance d i , calculate the K-means score for each node:
[0043]
[0044] In formula (7), Represents the K-means score of node i. The closer the node is to the cluster center, The higher the score, the more effective the K-means score of each node is calculated through the above steps, providing a basis for subsequent comprehensive scoring and aggregation.
[0045] Preferably, according to the present invention, the specific steps of step S5 are as follows:
[0046] S51. Combine the SVD score and the K-means score to get the comprehensive score of each node:
[0047]
[0048] In formula (8), S i represents the comprehensive score of node i, represents the SVD score of node i, represents the K-means score of node i, represents the SVD score of node j, represents the K-means score of node j;
[0049] S52: Select the gradients sent by nf nodes with the smallest comprehensive scores for average aggregation and calculate the global gradient g t :
[0050]
[0051] S53, use global gradient g t Update global model parameters:
[0052] x t+1 =x t -γ·g t (10)
[0054] In formula (10), x t+1 represents the global model parameters at the t+1th iteration, x t is the global model parameter at the tth iteration, and γ represents the learning rate;
[0055] S54. Use this method to continuously iterate the global model parameters until the objective function is satisfied or the set maximum number of iterations is reached.
[0056] Compared with the prior art, the present invention has the following beneficial effects:
[0057] The present invention combines the comprehensive scoring mechanism of singular value decomposition SVD and K-means clustering to effectively evaluate the contribution of each node, giving priority to nodes with higher comprehensive scores. Compared with the prior art, the present invention comprehensively considers the effectiveness of nodes, accurately identifies and screens Byzantine nodes, and significantly improves the robustness of distributed learning systems against Byzantine attacks under heterogeneous data distribution. Ultimately, accurate aggregation of information from honest nodes in complex data environments is achieved, thereby improving the performance and stability of the global model. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 It is a flow chart of a heterogeneous distributed robust learning gradient aggregation method combining SVD and K-means described in the present invention.
[0059] Figure 2This is a comparison of the test accuracy of the SK Score method of the present invention and other clustering methods in the non-attack scenario in Example 1 of the present invention.
[0060] Figure 3 This is a comparison of the test accuracy of the SK Score method of the present invention and other clustering methods under the Bit Flipping attack scenario in Example 1 of the present invention.
[0061] Figure 4 This is a comparison of the test accuracy of the SK Score method of the present invention and other clustering methods under the IPM attack scenario in Example 1 of the present invention. DETAILED DESCRIPTION
[0062] The present disclosure is further described below in conjunction with the accompanying drawings and embodiments.
[0063] Embodiment 1,
[0064] Ginseng Figure 1 This embodiment provides a heterogeneous distributed robust learning gradient aggregation method combining SVD and K-means, the method comprising the following steps:
[0065] S1. Build a heterogeneous distributed learning system. The specific steps are as follows:
[0066] A distributed learning system with a parameter server and n nodes is constructed, wherein the n nodes include G honest nodes and nG Byzantine nodes. The Byzantine nodes send false information to the parameter server in order to disrupt the normal operation of the system.
[0067] Next, we will optimize the system. The optimization goal is to find the parameter x that minimizes the expected value of the average loss function of all honest nodes. * :
[0068]
[0069] In formula (1), x * is the final goal of system optimization and the final parameters of the model after training. It represents the model parameters that perform best on the data of honest nodes; i represents the local data set of node i, and the data distribution among nodes is heterogeneous; F i Represents node i about its local classification dataset ξ i The loss function of i Represents a local classification dataset ξ i Calculate the loss function F i , x represents the model parameter, which is used to calculate the predicted value of the model and is continuously updated through the optimization algorithm to minimize the loss function. Preferably, the local classification data set is an image data set.
[0070] S2. The honest node extracts some data samples from its local classification data set and calculates the gradient using the stochastic gradient descent algorithm, and uploads the calculated gradient information to the parameter server;
[0071] S21. At each iteration, honest node i first receives the current global model parameters x from the parameter server. t , preferably, the global model is an MLP model;
[0072] S22, honest node i from its local classification data set ξ i Extract some data samples And calculate the global model parameter x t Under local stochastic gradient
[0073]
[0074] In formula (2), t represents the tth iteration, x t Represents the current global model parameters;
[0075] S23, honest node i calculates the local stochastic gradient Upload to the parameter server, and the Byzantine node will send malicious gradients to the parameter server.
[0076] S3. The parameter server performs singular value decomposition on the gradient matrices submitted by all nodes, extracts singular values and corresponding eigenvectors, and calculates the SVD score of the gradient to evaluate the effectiveness and importance of each gradient, providing a basis for subsequent aggregation steps. The specific steps are as follows:
[0077] S31. After receiving the gradients uploaded by all nodes, the parameter server constructs all the gradients into a matrix M:
[0078]
[0079] In formula (3), represents the random gradient sent by node n at the tth iteration;
[0080] S32, perform singular value decomposition on the gradient matrix M to obtain:
[0081] M=UΣV T (4)
[0083] In formula (4), U is the left singular vector matrix, Σ is the diagonal matrix containing singular values, V T is the right singular vector matrix;
[0084] S33. Extract the singular value σ corresponding to node i from the diagonal matrix Σ i , the singular value reflects the importance of node i. Subsequently, the SVD score of each node is calculated based on the extracted singular value:
[0085]
[0086] In formula (5), represents the SVD score of node i, σ i represents the singular value corresponding to node i, σ j represents the singular value corresponding to node j, n represents node n, Represents the sum of the singular values corresponding to all nodes. Through the above steps, the contribution of each node can be effectively evaluated, providing a basis for the subsequent calculation of the comprehensive score and aggregation process.
[0087] S4. The parameter server uses the K-means clustering algorithm to cluster the gradients of all nodes, obtain k cluster centers, and record the distance between each node and its cluster center. Based on the distance, the K-means score of the gradient is calculated for subsequent comprehensive scoring and gradient aggregation. The specific steps are as follows:
[0088] S41, the parameter server uses the K-means clustering algorithm to cluster the gradients of all nodes and randomly selects k cluster centers c q , where q∈{1,2,…,k};
[0089] S42, assign the gradient of each node to the cluster with the nearest cluster center, and record the distance between each node and its corresponding cluster center:
[0090]
[0091] In formula (6), d i Represents the cluster center c of node i and its cluster q The distance between represents the gradient sent by node i at the tth iteration, c q represents the cluster center of the cluster to which node i belongs;
[0092] S43, based on distance d i , calculate the K-means score for each node:
[0093]
[0094] In formula (7), Represents the K-means score of node i. The closer the node is to the cluster center, The higher the score, the more effective the K-means score of each node is calculated through the above steps, providing a basis for subsequent comprehensive scoring and aggregation.
[0095] S5. Combine the SVD score and K-means score to calculate the comprehensive score of each node, and select nf gradients with the smallest comprehensive score for average aggregation to obtain the final global gradient. The aggregated global gradient is used for parameter update. The specific steps are as follows:
[0096] S51. Combine the SVD score and the K-means score to get the comprehensive score of each node:
[0097]
[0098] In formula (8), S i represents the comprehensive score of node i, represents the SVD score of node i, represents the K-means score of node i, represents the SVD score of node j, represents the K-means score of node j;
[0099] S52: Select the gradients sent by nf nodes with the smallest comprehensive scores for average aggregation and calculate the global gradient g t :
[0100]
[0101] S53, use global gradient g t Update global model parameters:
[0102] x t+1 =x t -γ·g t (10)
[0104] Among them, x t+1 represents the global model parameters at the t+1th iteration, x t is the global model parameter at the tth iteration, and γ represents the learning rate;
[0105] S54. Use this method to continuously iterate the global model parameters until the objective function is satisfied or the set maximum number of iterations is reached.
[0106] Experimental example
[0107] The experiment uses the FEMNIST image dataset, and the loss function considers the cross entropy loss to evaluate the training effect of the model. The cross entropy loss function is defined as follows:
[0108]
[0109] Where N represents the total number of data samples, y i represents the true label, Indicates that the model predicts the input data sample as y i probability.
[0110] This method (SK Score) is compared with Median, RFA, Bulyan, and Multi-Krum aggregation methods, and three experimental scenarios are considered: no attack, Bit Flipping attack, and IPM attack. In the Bit Flipping attack scenario, after the Byzantine node calculates the random gradient, it takes a negative value for the gradient and then In the IPM attack scenario, the Byzantine node first calculates the average value of the gradients of all honest nodes, then multiplies this average gradient by a negative coefficient -∈, and finally sends this gradient to the parameter server. Figure 2 It can be seen that in the non-attack scenario, the method described in the present invention achieves the highest test accuracy, which is better than other methods. Figure 3 It can be seen that in the Bit Flipping attack scenario, the method described in the present invention achieves the highest test accuracy, which is better than other methods. Figure 4 It can be seen that in the IPM attack scenario, the method described in the present invention achieves the highest test accuracy, which is better than other methods.
[0111] Embodiment 2,
[0112] This embodiment applies the method described in the present invention to a personalized recommendation system in Internet services and a fraud detection system in financial analysis, as follows:
[0113] In Internet services, personalized recommendation systems, such as e-commerce platforms, video sites, and social media, usually rely on distributed machine learning to train recommendation models. These systems need to process massive amounts of data from users around the world, and user behavior data is usually heterogeneously distributed. For example, user preferences in different regions vary greatly, which will lead to uneven distribution of gradient matrices. In addition, there may be malicious nodes in the system, such as attacked servers or malicious users, which may upload false data or malicious gradients to interfere with the training of recommendation models.
[0114] The method of the present invention first extracts the global features of user behavior data through singular value decomposition and evaluates the contribution of each node. Then, through cluster analysis, normal nodes and abnormal nodes are distinguished and malicious nodes are identified. Finally, the SVD score and K-means score are combined to accurately identify and eliminate malicious nodes to ensure the robustness of the recommendation model. By effectively aggregating the gradients from normal nodes, the accuracy of the recommendation model is significantly improved. In the presence of malicious nodes, the system can still run stably to avoid interference with the recommendation results.
[0115] In the field of financial analysis, fraud detection systems often rely on distributed machine learning to train fraud detection models. These systems need to process heterogeneous data from multiple financial institutions (such as transaction records, user behavior, etc.), and the data distribution may vary greatly, resulting in uneven distribution of gradient matrices. In addition, there may be malicious nodes in the system (such as attacked financial institutions or malicious users), which may upload false data or malicious gradients, interfere with the training of fraud detection models, and lead to inaccurate detection results.
[0116] The method of the present invention first extracts the global features of transaction data through singular value decomposition and evaluates the contribution of each node. Then, through cluster analysis, normal nodes and abnormal nodes are distinguished and malicious nodes are identified. Finally, the SVD score and K-means score are combined to accurately identify and eliminate malicious nodes to ensure the robustness of the fraud detection model. By effectively aggregating the gradients from normal nodes, the accuracy of the fraud detection model is significantly improved. In the presence of malicious nodes, the system can still run stably to avoid interference with the detection results.
[0117] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solution of the present invention, and are not intended to limit the specific implementation methods of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the claims of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A heterogeneous distributed robust learning gradient aggregation method combining SVD and K-means, characterized in that: The method comprises: S1. Build a distributed learning system with a parameter server and n nodes, where the nodes are divided into honest nodes and Byzantine nodes; S2. The honest node extracts some data samples from the local classification data set and calculates the gradient using the stochastic gradient descent algorithm, and uploads the calculated gradient information to the parameter server; S3. The parameter server performs singular value decomposition on the gradient matrices submitted by all nodes, and calculates the SVD score of the gradient by extracting the singular values and the corresponding eigenvectors. S4. The parameter server uses the K-means clustering algorithm to cluster the gradients of all nodes, obtain k cluster centers, and record the distance between each node and its cluster center. Based on the distance, the K-means score of the gradient is calculated; S5. Combine the SVD score and K-means score to calculate the comprehensive score of each node, and select nf gradients with the smallest comprehensive score for average aggregation to obtain the final global gradient. The aggregated global gradient is used to update the global model parameters.
2. According to claim 1, a heterogeneous distributed robust learning gradient aggregation method combining SVD and K-means is characterized in that: The specific steps of step S1 are as follows: Construct a distributed learning system with a parameter server and n nodes; the n nodes include G honest nodes and nG Byzantine nodes; Next, we will optimize the system. The optimization goal is to find the parameter x that minimizes the expected value of the average loss function of all honest nodes. * : In formula (1), x * is the final goal of system optimization and the final parameters of the model after training. It represents the model parameters that perform best on the data of honest nodes; i represents the local data set of node i, and the data distribution among nodes is heterogeneous; F i Represents node i about its local data set ξ i The loss function of i Represents local dataset ξ i Calculate the loss function F i The expected value of , x represents the model parameters, which are used to calculate the predicted value of the model and are continuously updated through the optimization algorithm to minimize the loss function.
3. The heterogeneous distributed robust learning gradient aggregation method combining SVD and K-means according to claim 1, characterized in that: The specific steps of step S2 are as follows: S21, honest node i first receives the current global model parameter x from the parameter server t ; S22, honest node i from its local classification data set ξ i Extract some data samples And calculate the global model parameter x t Under local stochastic gradient In formula (2), t represents the tth iteration, x t Represents the current global model parameters; S23, honest node i calculates the local stochastic gradient Upload to the parameter server, and the Byzantine node will send malicious gradients to the parameter server.
4. The heterogeneous distributed robust learning gradient aggregation method combining SVD and K-means according to claim 1, characterized in that: The specific steps of step S3 are as follows: S31. After receiving the gradients uploaded by all nodes, the parameter server constructs all the gradients into a matrix M: In formula (3), represents the random gradient sent by node n at the tth iteration; S32, perform singular value decomposition on the gradient matrix M to obtain: M=UΣV T (4) In formula (4), U is the left singular vector matrix, Σ is the diagonal matrix containing singular values, V T is the right singular vector matrix; S33. Extract the singular value σ corresponding to node i from the diagonal matrix Σ i , the singular value reflects the importance of node i. Subsequently, the SVD score of each node is calculated based on the extracted singular value: In formula (5), represents the SVD score of node i, σ i represents the singular value corresponding to node i, σ j represents the singular value corresponding to node j, n represents node n, Represents the sum of the singular values corresponding to all nodes.
5. The heterogeneous distributed robust learning gradient aggregation method combining SVD and K-means according to claim 1, characterized in that: The specific steps of step S4 are as follows: S41, the parameter server uses the K-means clustering algorithm to cluster the gradients of all nodes and randomly selects k cluster centers c q , where q∈{1,2,…,k}; S42, assign the gradient of each node to the cluster with the nearest cluster center, and record the distance between each node and its corresponding cluster center: In formula (6), d i Represents the cluster center c of node i and its cluster q The distance between represents the gradient sent by node i at the tth iteration, c q represents the cluster center of the cluster to which node i belongs; S43, based on distance d i , calculate the K-means score for each node: In formula (7), Represents the K-means score of node i. The closer the node is to the cluster center, The higher the score.
6. The heterogeneous distributed robust learning gradient aggregation method combining SVD and K-means according to claim 1, characterized in that: The specific steps of step S5 are as follows: S51. Combine the SVD score and the K-means score to get the comprehensive score of each node: In formula (8), S i represents the comprehensive score of node i, represents the SVD score of node i, represents the K-means score of node i, represents the SVD score of node j, represents the K-means score of node j; S52: Select the gradients sent by nf nodes with the smallest comprehensive scores for average aggregation and calculate the global gradient g t : S53, use global gradient g t Update global model parameters: x t+1 =x t -γ·g t (10) In formula (10), x t+1 represents the global model parameters at the t+1th iteration, x t is the global model parameter at the tth iteration, and γ represents the learning rate; S54. Use this method to continuously iterate the global model parameters until the objective function is satisfied or the set maximum number of iterations is reached.
Citation Information
Patent Citations
Defense agent method for improving Byzantine robustness of distributed learning system
CN111967015A
Robust Byzantine fault-tolerant distributed gradient descent algorithm
CN111898763A
Gradient heterogeneous dual optimization method and device in distributed machine learning system, electronic equipment and storage medium
CN118070929A
Distributed learning aggregation method applied to attack scene, storage medium and program product
CN119089982A
System for automated malicious software detection
US11436330B1