Federal learning defense method for high-density label flipping attack
By using CNN models and unsupervised learning cluster analysis in federated learning to identify malicious nodes, the federated learning accuracy and robustness problems under high-density tag flip attacks are solved, and higher attack resistance and stability are achieved.
Patent Information
- Application Number
- CN202510041529.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art is difficult to effectively detect and defend against high-density tag flip attacks, resulting in a decrease in accuracy and robustness of federated learning models.
By introducing a convolutional neural network CNN model in federated learning, and using unsupervised learning to cluster the outlier gradient of user nodes on the central server side, the cluster density size is compared using cosine similarity to identify suspected malicious nodes, and then weight empowerment and model aggregation are performed.
It significantly improves the accuracy and robustness of federated learning under high-density tag flip attacks, can effectively resist more than 30% of high-density attacks, and maintains good stability.
Smart Images

Figure CN119962618A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of federated learning, and in particular relates to a federated learning defense method against high-density label flipping attacks. Background Art
[0002] In the era of artificial intelligence, data has become a key factor in promoting scientific and technological progress and social development. Traditional centralized data processing methods face a series of challenges such as serious privacy leakage and low communication efficiency. Distrust between devices has made data island conflicts increasingly prominent. Federated learning can ensure that all users can jointly train effective machine learning models without leaving the local data, realize knowledge sharing without data sharing, and better solve the problems of data islands and privacy protection.
[0003] However, due to the differences in security defense capabilities among user nodes in federated learning, federated learning faces many security issues and attack threats in practical applications, among which the most typical one is the label flipping attack initiated by malicious user nodes.
[0004] Label flipping attacks are attacks that poison the data source by tampering with the labels of local data while ensuring that the data's characteristic attributes remain unchanged. For example, an image originally labeled "dog" is mistakenly labeled as "cat", thereby misleading model training. The dirty data created by label flipping attacks generates incorrect model parameters by participating in local model training. During the federated aggregation process, the incorrect model parameters are uploaded to the central server, reducing the accuracy and robustness of the global model.
[0005] In order to defend against the harm of label flipping attacks to federated learning, there are currently a variety of defense mechanisms. However, existing algorithms often only target low-density attacks, and lack research on higher-density label flipping attacks. For example, the Byzantine fault-tolerant machine learning algorithm (Krum algorithm) can only defend against 40% of label flipping attacks at most, and when the attack ratio reaches 40%, the model stability is poor and the robustness decreases significantly. Therefore, how to improve the accuracy of federated learning under high-density label flipping attacks is an important problem that has yet to be solved. Summary of the invention
[0006] In order to solve the problem that the prior art is difficult to detect and defend against high-density label flipping attacks, the present invention provides a federated learning defense method for high-density label flipping attacks, which specifically includes the following steps:
[0007] S1. Establish a centralized federated architecture, which includes a central server and M user nodes. The central server initializes the global model parameters to W0;
[0008] The convolutional neural network CNN model is selected for implementation;
[0009] S2. The central server selects local users participating in the training and broadcasts the global model to the local users participating in the training;
[0010] The details are as follows:
[0011] S21. Assume that federated learning performs T rounds of model updates. In each round of training t, the central server randomly selects m = max(α·M,1) users to participate in the training, where M represents the total number of user nodes, α∈[0,1] represents the selection ratio each time, and “·” represents multiplication.
[0012] S22. The central server distributes the global model to the selected m users;
[0013] S3. User local model training, calculation and extraction of the last layer of neuron outlier gradients, and upload to the central server;
[0014] The details are as follows:
[0015] S31. Each user node w(w∈m) receives the global model W sent by the central server t , using the stochastic gradient descent SGD algorithm on the local dataset D w Perform E rounds of model training on the local model to obtain the local model parameters W t w ;
[0016] S32. After each user completes training, calculate the gradient size of the last layer of neurons in, represents the gradient of the i-th neuron in the output layer of the w-th user node during the t-th round of training, and ||·||2 is the L2 norm;
[0017] S33. Each user sorts the gradient size of the last layer of neurons;
[0018] S34. Each user has a gradient size of neurons after sorting Extract large outlier gradients, denoted as And upload to the central server;
[0019] S4. The central server uses unsupervised learning to learn outlier gradients of all users Perform cluster analysis and use cosine similarity to compare cluster density to find suspected malicious nodes;
[0020] The details are as follows:
[0021] S41. The central server receives outlier gradients from each user and forms an outlier gradient set l represents the total number of neurons in the last layer;
[0022] S42. The central server uses the K-means algorithm to perform unsupervised learning cluster analysis on the outlier gradient set outgradients, sets the cluster center K=2, and obtains two clusters, cluster1 and cluster2;
[0023] S43. The central server uses cosine similarity to calculate the cluster density sizes CluD1 and CluD2 of cluster 1 and cluster 2 respectively. The calculation formula is: CluD[i] = mean(CosSim[i]), where mean(.) is the mean function, and CosSim[i] represents the cosine similarity between the gradients in each cluster. The calculation formula is:
[0024]
[0025] CluD[i] represents the cluster density of different clusters, which is the result of taking the average value of cosine similarity CosSim[i];
[0026] S44. The central server compares the sizes of cluster densities CluD1 and CluD2 to distinguish the cluster cluster_malicious where the malicious gradient is located and the cluster cluster_honest where the honest gradient is located. The larger the cluster density, the more likely it is that the malicious user node gradient is updated in the cluster. The smaller the cluster density, the more likely it is that the honest user node gradient is updated in the cluster. The central server finds the cluster cluster_malicious where the malicious gradient is located by comparing the cluster densities, and then finds the poisoned user suspected of being attacked by label flipping based on the cluster index where the malicious gradient is located.
[0027] S5. The central server assigns weights to the user nodes participating in the training according to the rules of steps S51 and S52, completing a federation aggregation;
[0028] S51. The users who are attacked by label flipping are called poisoned users, and the users who are not attacked by label flipping are called honest users. The total number of honest users and poisoned users is m. The central server assigns an aggregate weight of δ to the honest user nodes. honest =1, giving a small weight to the suspected poisoned users detected in S4;
[0029] S52. The central server performs weighted averaging according to the weights assigned to each node, completes federation aggregation using the following formula, and updates the global model parameters. The calculation formula is:
[0030]
[0031] Among them, δ w ∈{δhonest ,δ malicious}, represents the weight value of the w-th user node, W t w represents the model parameters of the w-th user node in the t-th round of training;
[0032] S6. The central server sends the updated global model parameters W t , the system repeats steps S2 to S5 and performs iterative training until T rounds of training are completed.
[0033] In step S1 of a specific embodiment of the present invention, the convolutional neural network CNN model has a total of M=50 user nodes.
[0034] In step S22 of one embodiment of the present invention, when the federated learning process starts training, the central server generates the initial global model parameter W0 and sends it to the users participating in the training. After the users complete the local training, they upload the local model parameter W0. t w , the central server aggregates all local model parameters and generates new global model parameters W after updating according to the rules defined in step S5 t .
[0035] In step S34 of another specific embodiment of the present invention, each user is sorted according to the size of the neuron gradient. extract Outlier gradients in the middle first third.
[0036] In step S44 of another embodiment of the present invention, the specific indexing method used is as follows:
[0037] Assume that the gradient element in the cluster cluster_malicious where the malicious gradient is located is Right now: Where p∈m means that among all the m users participating in federated learning, the pth user is a poisoned user, thereby completing the index and finding the poisoned user suspected of being attacked by label flipping.
[0038] In step S51 of another specific embodiment of the present invention, δ malicious =0.5.
[0039] This paper proposes a federated learning defense method for high-density label flipping attacks, which has the following advantages over existing methods:
[0040] (1) Effectiveness: Under low-density label flipping attacks, the accuracy of the proposed method is basically consistent with that of the Federated Averaging algorithm (fedavg) and the median algorithm (median), indicating that the proposed algorithm can effectively complete federated learning and prevent the model from being attacked by label flipping.
[0041] (2) Accuracy: Under low-density label flipping attacks, the accuracy of the method of the present invention is better than that of the median algorithm. Under high-density attacks above 30%, the present invention is significantly better than the fedavg and median algorithms.
[0042] (3) Robustness: Under high-density label flipping attacks, the training volatility of the fedavg and median algorithms is large and the training is very unstable, but the method of the present invention can still maintain good stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 This is a diagram of the federated learning architecture used in the present invention;
[0044] Figure 2 This is a comparative analysis of the accuracy of each algorithm of the present invention under a 10% attack ratio;
[0045] Figure 3 This is a comparative analysis of the accuracy of each algorithm of the present invention under a 20% attack ratio;
[0046] Figure 4 This is a comparative analysis of the accuracy of each algorithm of the present invention under a 30% attack ratio;
[0047] Figure 5 This is a comparative analysis of the accuracy of each algorithm of the present invention under a 40% attack ratio;
[0048] Figure 6 This is a comparative analysis of the accuracy of each algorithm of the present invention under a 50% attack ratio. DETAILED DESCRIPTION
[0049] The technical solution of the present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0050] The present invention proposes a federated learning defense method against high-density label flipping attacks, and the specific implementation steps include:
[0051] S1. Establish Figure 1 The centralized federated architecture shown includes a central server and M user nodes. The central server initializes the global model parameters to W0 (including the model's weights and biases).
[0052] This embodiment uses a convolutional neural network (CNN) model for implementation, and a total of M=50 user nodes are set.
[0053] S2. The central server selects local users to participate in the training and broadcasts the global model to the local users participating in the training.
[0054] S21. Assume that federated learning performs T rounds of model updates. In each round of training t, the central server randomly selects m = max(α·M,1) users to participate in the training, where M represents the total number of user nodes, α∈[0,1] represents the selection ratio each time, and “·” represents multiplication.
[0055] S22. The central server distributes the global model to the selected m users. (During the first distribution, the initialized global model W0 is distributed to the selected m users)
[0056] The federated learning process is an iterative training process. When training starts, the central server generates the initial global model parameters W0 and sends them to the users participating in the training. After the users complete the local training, they upload the local model parameters W0. t w , the central server aggregates all local model parameters and generates new global model parameters W after updating according to the rules defined later (rules defined in step S5) t This technology is well known to those skilled in the art and will not be described again.
[0057] S3. User local model training, calculation and extraction of the last layer of neuron outlier gradients, and upload to the central server.
[0058] S31. Each user node w(w∈m) receives the global model W sent by the central server t , using the Stochastic Gradient Descent (SGD) algorithm on the local dataset D w Perform E rounds of model training on the local model to obtain the local model parameters W t w ;
[0059] The SGD algorithm is a classic optimization algorithm used in machine learning to find the minimum value of a function. It calculates the gradient of the loss function by randomly selecting one (or a batch of samples) and updates the model parameters in the opposite direction of the gradient to gradually approach the global minimum or local minimum. This algorithm is well known to those skilled in the art.
[0060] S32. After each user completes training, calculate the gradient size of the last layer of neurons The formula is: in, represents the gradient of the i-th neuron in the output layer of the w-th user node during the t-th round of training, and ||·||2 is the L2 norm.
[0061] S33. Each user ranks the gradient size of the last layer of neurons.
[0062] S34. Each user has a gradient size of neurons after sorting Extract the larger outlier gradient, denoted as And upload to the central server.
[0063] In the specific implementation of the present invention, "larger" is usually taken as The top third of the value.
[0064] S4. The central server uses unsupervised learning to learn outlier gradients of all users Perform cluster analysis and use cosine similarity to compare cluster density to find suspected malicious nodes.
[0065] Unsupervised learning is a machine learning paradigm that allows algorithms to explore data sets without pre-labeled training data to discover hidden patterns, structures or relationships in the data, and is often used for tasks such as clustering, dimensionality reduction, and anomaly detection. This algorithm is well known to those skilled in the art.
[0066] S41. The central server receives outlier gradients from each user and forms an outlier gradient set l represents the total number of neurons in the last layer.
[0067] S42. The central server uses the K-means algorithm to perform unsupervised learning cluster analysis on the outlier gradient set outgradients, sets the cluster center K=2, and obtains two clusters, cluster1 and cluster2.
[0068] K-means algorithm is a popular unsupervised learning clustering algorithm, whose goal is to divide data points into K clusters, so that the points within the cluster are as similar as possible, and the points between clusters are as different as possible, usually by iteratively selecting cluster centers, assigning data points to the nearest cluster center, and updating the position of the cluster center to achieve this goal. This algorithm is well known to those skilled in the art.
[0069] S43. The central server uses cosine similarity to calculate the cluster density CluD1 and CluD2 of cluster1 and cluster2 respectively, and the calculation formula is: CluD[i] = mean(CosSim[i]), where mean(.) is a mean function well known to those skilled in the art (see mean(.) provided by Pyphon). Among them, CosSim[i] represents the cosine similarity between the gradients in each cluster, and the calculation formula is:
[0070]
[0071] Among them, CluD[i] represents the cluster density of different clusters, which is the result of taking the average value of cosine similarity CosSim[i].
[0072] S44. The central server compares the size of cluster densities CluD1 and CluD2 to distinguish the cluster where the malicious gradient is located, cluster_malicious, and the cluster where the honest gradient is located, cluster_honest. Since label flipping attacks reduce the similarity of gradients, the larger the cluster density, the more likely it is that the malicious user node gradient is updated in the cluster; the smaller the cluster density, the more likely it is that the honest user node gradient is updated in the cluster. Therefore, the central server can find the cluster where the malicious gradient is located, cluster_malicious, by comparing the cluster density, and then find the poisoned user suspected of being attacked by label flipping based on the cluster index where the malicious gradient is located. The specific indexing method is as follows:
[0073] Assume that the gradient element in the cluster cluster_malicious where the malicious gradient is located is Right now: Where p∈m means that among all the m users participating in federated learning, the pth user is a poisoned user, thereby completing the index and finding the poisoned user suspected of being attacked by label flipping.
[0074] S5. The central server assigns weights to the user nodes participating in the training according to the rules of steps S51 and S52 to complete a federation aggregation.
[0075] S51. The users who are attacked by label flipping are called poisoned users, and the users who are not attacked by label flipping are called honest users. The total number of honest users and poisoned users is m. The central server assigns an aggregate weight of δ to the honest user nodes. honest =1, a smaller weight is assigned to the suspected poisoned user detected in S4. In a specific embodiment of the present invention, δ malicious =0.5.
[0076] S52. The central server performs weighted averaging according to the weights assigned to each node, completes federation aggregation using the following formula, and updates the global model parameters. The calculation formula is:
[0077]
[0078] Among them, δ w ∈{δ honest ,δ malicious}, represents the weight value of the w-th user node, W t w Represents the model parameters of the w-th user node in the t-th round of training.
[0079] S6. The central server sends the updated global model parameters W t , the system repeats steps S2 to S5 and performs iterative training until T rounds of training are completed.
[0080] The following are specific experiments and descriptions of the embodiments:
[0081] This embodiment uses the classic MNIST handwritten digital image dataset, which contains 60,000 training samples and 10,000 test samples. Each sample is a grayscale image of 28x28 pixels, representing ten numbers from 0 to 9, which is commonly used in image classification tasks in machine learning. In the embodiment, the MNIST dataset samples are randomly distributed to each user node participating in federated learning to simulate actual federated learning.
[0082] The embodiment uses a CNN model for implementation, with a total of M=50 user nodes, and global iteration updates for 100 times. The experiment changes the data label "5" of the poisoned node (the user node attacked by the label flipping attack) to "3" to create poisonous data; the label flipping attack ratios are set to 10%, 20%, 30%, 40% and 50% in sequence.
[0083] This example defines global accuracy (ACC) as the evaluation index of the algorithm. The higher the global accuracy, the better the algorithm performance. This example compares with fedavg, median and other algorithms to verify the superiority of the algorithm proposed in the present invention from three dimensions: effectiveness, accuracy and robustness. The results are as follows: Figures 2 to 6 shown.
[0084]
Claims
1. A federated learning defense method against high-density label flipping attacks, characterized in that: The specific steps include: S1. Establish a centralized federated architecture, which includes a central server and M user nodes. The central server initializes the global model parameters to W0; The convolutional neural network CNN model is selected for implementation; S2. The central server selects local users participating in the training and broadcasts the global model to the local users participating in the training; The details are as follows: S21. Assume that federated learning performs T rounds of model updates. In each round of training t, the central server randomly selects m = max(α·M,1) users to participate in the training, where M represents the total number of user nodes, α∈[0,1] represents the selection ratio each time, and "·" represents multiplication; S22. The central server distributes the global model to the selected m users; S3. User local model training, calculation and extraction of the last layer of neuron outlier gradients, and upload to the central server; The details are as follows: S31. Each user node w(w∈m) receives the global model W sent by the central server t , using the stochastic gradient descent SGD algorithm on the local dataset D w Perform E rounds of model training on the local model to obtain the local model parameters W t w ; S32. After each user completes training, calculate the gradient size of the last layer of neurons in, represents the gradient of the i-th neuron in the output layer of the w-th user node during the t-th round of training, and ||·||2 is the L2 norm; S33. Each user sorts the gradient size of the last layer of neurons; S34. Each user has a gradient size of neurons after sorting Extract large outlier gradients, denoted as And upload to the central server; S4. The central server uses unsupervised learning to learn outlier gradients of all users Perform cluster analysis and use cosine similarity to compare cluster density to find suspected malicious nodes; The details are as follows: S41. The central server receives outlier gradients from each user and forms an outlier gradient set l represents the total number of neurons in the last layer; S42. The central server uses the K-means algorithm to perform unsupervised learning cluster analysis on the outlier gradient set outgradients, sets the cluster center K=2, and obtains two clusters, cluster1 and cluster2; S43. The central server uses cosine similarity to calculate the cluster density sizes CluD1 and CluD2 of cluster 1 and cluster 2 respectively. The calculation formula is: CluD[i] = mean(CosSim[i]), where mean(.) is the mean function, and CosSim[i] represents the cosine similarity between the gradients in each cluster. The calculation formula is: CluD[i] represents the cluster density of different clusters, which is the result of taking the average value of cosine similarity CosSim[i]; S44. The central server compares the sizes of cluster densities CluD1 and CluD2 to distinguish the cluster cluster_malicious where the malicious gradient is located and the cluster cluster_honest where the honest gradient is located. The larger the cluster density, the more likely it is that the malicious user node gradient is updated in the cluster. The smaller the cluster density, the more likely it is that the honest user node gradient is updated in the cluster. The central server finds the cluster cluster_malicious where the malicious gradient is located by comparing the cluster densities, and then finds the poisoned user suspected of being attacked by label flipping based on the cluster index where the malicious gradient is located. S5. The central server assigns weights to the user nodes participating in the training according to the rules of steps S51 and S52, completing a federation aggregation; S51. The users who are attacked by label flipping are called poisoned users, and the users who are not attacked by label flipping are called honest users. The total number of honest users and poisoned users is m. The central server assigns an aggregate weight of δ to the honest user nodes. honest =1, giving a small weight to the suspected poisoned users detected in S4; S52. The central server performs weighted averaging according to the weights assigned to each node, completes federation aggregation using the following formula, and updates the global model parameters. The calculation formula is: Among them, δ w ∈{δ honest ,δ malicious }, represents the weight value of the w-th user node, W t w represents the model parameters of the w-th user node in the t-th round of training; S6. The central server sends the updated global model parameters W t , the system repeats steps S2 to S5 and performs iterative training until T rounds of training are completed.
2. The federated learning defense method against high-density label flipping attacks according to claim 1, characterized in that: In step S1, the convolutional neural network CNN model has a total of M=50 user nodes.
3. The federated learning defense method against high-density label flipping attacks according to claim 1, characterized in that: In step S22, when the federated learning process starts training, the central server generates the initial global model parameter W0 and sends it to the users participating in the training. After the users complete the local training, they upload the local model parameter W0. t w , the central server aggregates all local model parameters and generates new global model parameters W after updating according to the rules defined in step S5 t .
4. The federated learning defense method for high-density label flipping attacks according to claim 1, characterized in that: In step S34, each user is assigned a neuron gradient based on the sorted neuron gradient. extract Outlier gradients in the middle first third.
5. The federated learning defense method against high-density label flipping attacks according to claim 1, characterized in that: In step S44, the specific indexing method used is as follows: Assume that the gradient element in the cluster cluster_malicious where the malicious gradient is located is Right now: Where p∈m means that among all the m users participating in federated learning, the pth user is a poisoned user, thereby completing the index and finding the poisoned user suspected of being attacked by label flipping.
6. The federated learning defense method against high-density label flipping attacks according to claim 1, characterized in that: In step S51, δ is assigned malicious =0.5.
Citation Information
Cited By
Intelligent system evaluation method for laboratory data identification data credibility
CN121637369A
An intelligent system evaluation method for data credibility of laboratory data recognition data
CN121637369B