A Method and System for Defending Against Byzantine Attacks in the Context of Federated Learning

By using methods such as preprocessing and expansion, gradient anomaly checking and cluster detection in federated learning, the defense performance problem of federated learning under heterogeneous data distribution is solved, and effective defense against Byzantine attacks is achieved.

CN116389093BActive Publication Date: 2025-07-25HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310308504.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-28
Publication Date
2025-07-25
Estimated Expiration
2043-03-28

AI Technical Summary

Technical Problem

The existing federated learning defense method is not effective in heterogeneous data distribution scenarios and relies on the same verification data set as the client data distribution, resulting in a degradation of defense performance.

Method used

The server obtains data sets different from the client for preprocessing and expansion, and the client divides and mixes the data sets, combines four gradient anomaly checks, uses cosine similarity and clustering algorithm to detect and eliminate malicious gradients, and designs other modal aggregation strategies to correct the global model.

Benefits of technology

It effectively reduces the impact of heterogeneous data distribution on defense effects, improves the detection accuracy of malicious updates, and enhances the defense capabilities in actual scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116389093B_ABST
    Figure CN116389093B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for defending against Byzantine attacks in the context of federated learning, including: the server obtains a second dataset D different from the first dataset obtained by the client aux and the number of classifications of the server model, preprocesses the second dataset to obtain a preprocessed second dataset, expands the obtained number of classifications of the server model to obtain an expanded server model, and sends the preprocessed second dataset and the expanded server model to the client. The client obtains the first dataset and divides the dataset into a training set and a test set according to a ratio of 8:2. The client obtains a second dataset that is completely different from the preprocessed first dataset from the server and modifies the size of the second dataset to the input size of the client model. The present invention can solve the technical problem that the existing FL defense method is only effective when the datasets of the clients are independent and identically distributed, thus greatly affecting the application of the method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of distributed machine learning security, and more specifically, relates to a method and system for defending against Byzantine attacks in a federated learning scenario. Background Art

[0002] Federated Learning (FL), which has received much attention in recent years, is a promising distributed learning paradigm that enables numerous resource-constrained clients to jointly learn a global model without sharing privacy-sensitive data. Federated learning has been widely applied to various real-world applications such as finance, healthcare, and smart cities. Specifically, in federated learning, thousands of clients (such as Internet of Things devices, edge devices) perform local model training on their private datasets and upload updates (or local models), while a central server (e.g., Google, Apple) provides an aggregation strategy to publish the aggregated global model. However, federated learning is vulnerable to poisoning attacks, where an attacker (e.g., an affected client) may tamper with the local training data (referred to as data poisoning attack) or the local model (referred to as model poisoning attack), thus undermining the global model. For example, an attacker can modify the labels of the data to make the global model make incorrect predictions, or randomly generate a local model to save computing resources while enjoying the benefits of using the global model.

[0003] To prevent poisoning attacks on FL, a great deal of research work has been done in enhancing the Byzantine robustness of FL. The key idea of these works is to design a robust aggregation strategy on the server side so that the impact of poisoned updates on the global model can be negligible.

[0004] However, there are some non-negligible technical problems in existing FL defense methods: First, existing work has not fully considered heterogeneous data distribution (non-IID), which is one of the most important and fundamental characteristics of FL. Most existing defenses are only effective when the datasets of clients are independent and identically distributed, which greatly affects the application of this method. Second, in order to more effectively detect malicious updates, some defense methods rely on a clean validation dataset that is close to or even the same as the client data distribution. However, for the FL scenario originally designed to protect the privacy of clients, such an assumption is very unrealistic and will lead to a reduction in defense performance. Summary of the Invention

[0005] In view of the above deficiencies or improvement requirements of the prior art, the present invention provides a method and system for defending against Byzantine attacks in the context of federated learning, aiming to solve the technical problems that existing FL defense methods are only effective when the datasets of clients are independent and identically distributed, thus greatly affecting the application of this method, and the technical problem that the defense performance is reduced due to over-reliance on a clean validation dataset that is close to or even the same as the client data distribution.

[0006] To achieve the above object, according to one aspect of the present invention, a method for defending against Byzantine attacks in the context of federated learning is provided, including the following steps:

[0007] (1) The server obtains a second dataset D that is different from the first dataset obtained by the client, preprocesses the second dataset to obtain a preprocessed second dataset, expands the number of classifications of the obtained server model to obtain an expanded server model, and sends the preprocessed second dataset and the expanded server model to the client. aux And the number of classifications of the server model, preprocess the second dataset to obtain a preprocessed second dataset, expand the number of classifications of the obtained server model to obtain an expanded server model, and send the preprocessed second dataset and the expanded server model to the client.

[0008] (2) The client obtains the first dataset, divides the dataset into a training set and a test set according to a ratio of 8:2. Obtain a second dataset that is completely different from the first dataset and preprocessed from the server, modify the size of the second dataset to the input size of the client model, and also label the labels of the second dataset as L, L + 1, L + 2,..., L + K; where L represents the number of data categories of the client, and K is the expanded number of classifications. Then uniformly mix the training set of the first dataset and the second dataset to obtain a mixed training set, and initialize the weights of the client model according to the weights of the expanded server model to obtain an initialized client model.

[0009] (3) The client inputs the mixed training set into the client model initialized in step (2) to obtain a client gradient set C = {g1, g2,..., g m} for updating the client model, and sends it to the server, where the range of m is 1 ≤ m ≤ n, and n is the total number of clients participating in federated learning, represents the gradient of the i1-th client among all clients participating in federated learning, and i1 ∈ [1, m].

[0010] (4) The server performs four client gradient anomaly checks and malicious client gradient removal processes on the client gradients obtained in step (3) in sequence, and aggregates the remaining client gradients to obtain updated client gradients. Then, the server uses the updated client gradients to perform a second-round update on the server model to obtain the server model of the second round, …, repeating this process for T rounds to obtain the final server model, where T ≥ 1000.

[0011] Preferably, in step (1), first, the size of the second dataset is modified to the input size of the server model. Then, the labels of the second dataset are marked as L, L + 1, L + 2, ..., L + K; where L represents the number of data categories of the client, and K is the number of extended classifications. Finally, the number of classifications of the server model is set to L + K;

[0012] The server model is the same deep learning model as the client model.

[0013] Preferably, step (4) includes the following sub-steps:

[0014] (4-1) The server randomly selects m clients and receives the set of client gradients C = {g1, g2, …, g m}, which represents the gradient of the i1-th client among all clients participating in federated learning, and i1 ∈ [1, m];

[0015] (4-2) The server performs a filtering process on the set of client gradients C = {g1, g2, …, g m} obtained in step (4-1) to remove malicious client gradients with similar or identical directions, and obtains the set of client gradients C1 after the first detection;

[0016] (4-3) The server performs a second detection process on the set of client gradients C1 obtained in step (4-2) to obtain the set of client gradients C2 after removing malicious client gradients whose directions are not similar to those of normal client gradients;

[0017] (4-4) The server performs a third detection process on the set of client gradients C2 obtained in step (4-3) to obtain the set of client gradients C3 after further removing malicious client gradients whose directions are closer to those of normal client gradients;

[0018] (4-5) The server performs a fourth detection on the set of normal client gradients C3 obtained in step (4-4) to correct the norm of each client gradient therein, thereby obtaining the corrected client gradient where i8 ∈ [1, |C3|];

[0019] (4-6) The server performs an aggregation process on all the corrected client gradients obtained in step (4-5) to obtain the latest server gradient g, and updates the latest server model accordingly;

[0020] (4-7) Repeat steps (4-1) to (4-6) for T rounds to obtain the final server model.

[0021] Preferably, step (4) includes the following sub-steps:

[0022] (4-2-1) The server calculates the cosine similarity between all client gradients in the client gradient set {g1, g2,..., g m} to obtain a similarity score set {s1, s2,..., s m};

[0023] Specifically, the calculation formula for the similarity score in this step is:

[0024]

[0025] where represents the set of the top client gradients in the client gradient set C that have the highest similarity to . The similarity between and is measured by cosine similarity; <,> represents the inner product, and j ∈ [1, m].

[0026] (4-2-2) The server calculates the median m of the similarity score set {s1, s2,..., s

[0027] (4-2-3) The server uses the median obtained in step (4-2-2) to obtain the distance set between the similarity score and the median and obtains the median

[0028] (4-2-4) Set the counter i1 = 1 and initialize an empty malicious client gradient set C mal ;

[0029] (4-2-5) Determine whether i1 is less than or equal to m. If so, go to step (4-2-6); otherwise, go to step (4-2-9);

[0030] (4-2-6) Determine whether the similarity score is greater than If yes, go to step (4-2-7), otherwise go to step (4-2-8);

[0031] (4-2-7) Score the similarity Corresponding client gradient Put in the malicious client gradient set C mal ;

[0032] (4-2-8) Set i1=i1+1 and return to step (4-2-5);

[0033] (4-2-9) The server uses the malicious client gradient set C obtained in step (4-2-7) mal To remove the malicious client gradients in the client gradient set C obtained in step (4-1), so as to obtain the client gradient set C1=CC after the first detection mal .

[0034] Preferably, step (4-3) comprises the following sub-steps:

[0035] (4-3-1) The server checks all client gradients in the client gradient set C1 after the first detection. Perform cosine similarity calculations between pairs to obtain a similarity score set in represents the i2th client gradient in the client gradient set C1 after the first detection, Represents a similarity score set The i2th similarity score in , and i2∈[1,|C1|];

[0036] Specifically, the similarity score in this step The calculation formula is:

[0037]

[0038] in Indicates that the client gradient set C1 is neutralized after the first detection Most similar to A collection of client-side gradients.

[0039] (4-3-2) The server obtains the similarity score set according to step (4-3-1) Mean shift clustering is performed on the client gradients in the client gradient set C1 to obtain h clusters {Ω1,Ω2,…,Ω h}; where h can be determined by the mean shift clustering algorithm. represents the i3th cluster set, the elements in this cluster set are client gradients, and i3∈[1,h];

[0040] (4-3-3) The server obtains the number of elements a of the cluster with the most elements among the h clusters obtained in step (4-3-2).

[0041] (4-3-4) Set the counter i4 = 1 and initialize an empty set.

[0042] (4-3-5) Determine whether i4 is greater than or equal to a. If so, go to step (4-3-11); otherwise, go to step (4-3-6).

[0043] (4-3-6) Set the counter i3 = 1.

[0044] (4-3-7) Determine whether i3 is greater than or equal to h. If so, go to step (4-3-10); otherwise, go to step (4-3-8).

[0045] (4-3-8) Determine whether is empty. If so, go to step (4-3-9); otherwise, add the client gradient to the empty set and then go to step (4-3-9).

[0046] (4-3-9) Set i3 = i3 + 1 and return to step (4-3-7).

[0047] (4-3-10) Set i4 = i4 + 1 and return to step (4-3-5).

[0048] (4-3-11) Initialize three empty sets D1, D2, D mal ;

[0049] (4-3-12) Set the counter i4 = 1.

[0050] (4-3-13) Determine whether i4 is greater than or equal to a. If so, go to step (4-3-16); otherwise, go to step (4-3-14).

[0051] (4-3-14) Set D1[i4] as the average value of all client gradients in the group ;

[0052] (4-3-15) Set i4 = i4 + 1 and return to step (4-3-13).

[0053] (4-3-16) The server calculates the cosine similarity between all client gradients in the client gradient set D1 pairwise to obtain a similarity score set where ​​Denote the \(i_5\)-th client gradient in the client gradient set \(D_1\). Denote the set of relative cosine similarity scores and the \(i_5\)-th similarity score in it, where \(i_5\in[1,|D_1|]\);

[0054] Specifically, the calculation formula for this step is:

[0055]

[0056] where denotes the set of the first client gradients in the client gradient set \(D_1\) that are most similar to ;

[0057] (4 - 3 - 17) The server performs mean shift clustering on the obtained in step (4 - 3 - 16) to obtain the index corresponding to the largest cluster, and marks the client gradient set corresponding to this index as \(\varPsi\). Note that at this time, the client gradients in \(\varPsi\) are the most similar part in \(D_1\);

[0058] (4 - 3 - 18) Treat all client gradients in \(D_1\) but not in \(\varPsi\) as malicious client gradients and put them into the malicious client gradient set \(D\) mal ;

[0059] (4 - 3 - 19) Set the counter \(i_4 = 1\);

[0060] (4 - 3 - 20) Determine whether \(i_4\) is greater than or equal to \(a\). If so, go to step (4 - 3 - 24); otherwise, go to step (4 - 3 - 21);

[0061] (4 - 3 - 21) Determine whether \(D_1[i_4]\) is a client gradient in \(D\) mal . If so, go to step (4 - 3 - 22); otherwise, go to step (4 - 3 - 23);

[0062] (4 - 3 - 22) Put all client gradients in the group into the set \(D_2\);

[0063] (4 - 3 - 23) Set \(i_4 = i_4 + 1\) and return to step (4 - 3 - 20);

[0064] (4 - 3 - 24) The server calculates the cosine similarity between all client gradients in the client gradient set \(C_1\) obtained in step (4 - 3 - 7) and the client gradient set \(D_2\) obtained in step (4 - 3 - 22) pairwise to obtain the set of similarity scores where denotes the set of similarity scores the i6-th similarity score in it, and i6 ∈ [1, |D2|];

[0065] Specifically, the calculation formula for this step is:

[0066]

[0067] where represents the set of the top client gradients in the client gradient set C1 that are most similar to D2[i6].

[0068] (4 - 3 - 25) Re-initialize an empty malicious client gradient set C mal , and set a hyperparameter β, where 0 ≤ β ≤ 1. In this solution, β = 0.2 is preferably selected;

[0069] (4 - 3 - 26) Put the first β·|D2| client gradients with the lowest similarity scores from the client gradient set D2 into the malicious client gradient set C ; mal in;

[0070] (4 - 3 - 27) The server uses the malicious client gradient set C mal obtained in step (4 - 3 - 26) to remove the malicious client gradients in the client gradient set C1, so as to obtain the client gradient set C2 = C1 - C mal ;

[0071] Preferably, step (4 - 4) includes the following sub-steps:

[0072] (4 - 4 - 1) The server performs gradient descent processing on each client gradient in the client gradient set C2 obtained in step (4 - 3 - 27) to obtain the client model corresponding to each client gradient where represents the i7-th client gradient in the client gradient set C2, represents the corresponding client model, and i7 ∈ [1, |C2|];

[0073] Specifically, the calculation formula for this step is as follows:

[0074]

[0075] where w is the server model and η is the learning rate.

[0076] (4 - 4 - 2) The server uses the second dataset D aux to process the client model w iPerform loss calculation to obtain the evaluated loss change value ρ i ;

[0077] Specifically, the calculation formula for this step is as follows:

[0078]

[0079] where is the loss function.

[0080] (4-4-3) The server performs anomaly selection on the loss change value obtained in step (4-4-2) to obtain the malicious client gradient set C with a change less than the threshold λ mal ;

[0081] (4-4-4) The server uses the malicious client gradient set C obtained in step (4-4-3) mal to screen the client gradient set C2 obtained in (4-3-26) to obtain the normal client gradient set C3. The calculation formula for this step is C3 = C2 - C mal .

[0082] Preferably, step (4-5) includes the following sub-steps

[0083] (4-5-1) The server calculates the norm of the client gradient set C3 obtained in step (4-4-4) to obtain the median f of the norms of all client gradients;

[0084] (4-5-2) The server divides the norms greater than the median f in step (4-5-1) by f and puts them into the empty set D. For the norms less than the median f, it divides f by these norms and puts them into the empty set D to obtain the set D for judging norm anomalies;

[0085] (4-5-3) The server statistically processes the norms in D to obtain the standard deviation s and the mean 1 of all norms;

[0086] (4-5-4) The server removes all client gradients with norms greater than 8(s + l) after the processing in step (4-5-2) from C3 to obtain the client gradient set C3 excluding malicious client gradients with very large and very small norms;

[0087] (4-5-5) The server calculates the norm of the client gradient set C3 obtained in step 4-5-4 to obtain the average norm e of all client gradients in C3;

[0088] (4-5-6) The server performs norm correction on all client gradients in C3 using e obtained in step (4-5-5);

[0089] Specifically, the norm correction formula for this step is as follows, for all client gradients in C3 All have:

[0090]

[0091] According to another aspect of the present invention, there is provided a system for defending against Byzantine attacks in a federated learning scenario, including:

[0092] A first module, which is set on the server side and is used to obtain a second data set D that is different from the first data set obtained by the client aux and the number of classifications of the server model, preprocess the second data set to obtain a preprocessed second data set, expand the number of classifications of the obtained server model to obtain an expanded server model, and send the preprocessed second data set and the expanded server model to the client.

[0093] A second module, which is set on the client side and is used to obtain the first data set, divide the data set into a training set and a test set according to a ratio of 8:2. Obtain a second data set that is completely different from the first data set and preprocessed from the server, modify the size of the second data set to the input size of the client model, and also label the second data set with labels L, L + 1, L + 2,..., L + K; where L represents the number of data categories of the client, and K is the expanded number of classifications. Then evenly mix the training set of the first data set and the second data set to obtain a mixed training set, and initialize the weights of the client model according to the weights of the expanded server model to obtain an initialized client model.

[0094] A third module, which is set on the client side and is used to input the mixed training set into the client model initialized by the second module to obtain a client gradient set C = {g1, g2,..., g m}, and send it to the server, where the range of m is 1 ≤ m ≤ n, and n is the total number of clients participating in federated learning, represents the gradient of the i1-th client among all clients participating in federated learning, and i1 ∈ [1, m].

[0095] A fourth module, which is set on the server side and is used to perform four client gradient anomaly checks and malicious client gradient elimination processing on the client gradients obtained by the third module in sequence, and aggregate the remaining client gradients to obtain updated client gradients, and use the updated client gradients to perform a second-round update on the server model to obtain a second-round server model,..., repeat this process for T rounds to obtain a final server model, where T ≥ 1000.

[0096] Generally speaking, compared with the prior art through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:

[0097] (1) Since the present invention adopts step (4) and designs a strategy of clustering first and then grouping, the degree of data heterogeneity (non-IID) in FL is greatly reduced, the similarity between normal updates is significantly improved, so that the present invention can use the similarity detection method to filter out the updates inconsistent with the majority. Therefore, the technical problem that the existing FL defense methods are only effective when the datasets of clients are independent and identically distributed, resulting in poor applicability of these defense methods can be solved;

[0098] (2) Since the present invention adopts step (1) and step (2), the server first obtains a second dataset different from the first dataset obtained by the client, preprocesses it and then sends it to the client. The second dataset is used as an auxiliary dataset, which is quite small and does not need to be related to the local dataset. The present invention uses the second dataset to evaluate the performance of each update to limit the normal direction space, and any update with poor performance on the auxiliary dataset will be discarded. Therefore, the technical problem that the existing FL defense methods rely on a clean validation dataset that is close to or even the same as the client data distribution, resulting in low defense efficiency can be solved;

[0099] (3) The present invention uses four independent modules to define the spatial characteristics of maliciousness. Any update belonging to these direction spaces will be rejected. The fourth module proposes an equal modulus aggregation strategy, which can further correct malicious updates that can escape the first three modules, and can greatly reduce the influence of malicious directions with a small deviation from the normal update direction on the global update direction. BRIEF DESCRIPTION OF THE DRAWINGS

[0100] Figure 1 is the system architecture diagram of the method for defending against Byzantine attacks in the federated learning scenario of the present invention;

[0101] Figure 2 is the flowchart of the method for defending against Byzantine attacks in the federated learning scenario of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0102] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0103] The present invention proposes a new FL Byzantine defense scheme, which can defend against the most advanced poisoning attacks against FL in a data heterogeneous scenario that is more in line with the real world. The core idea of the present invention is to compress the normal direction space of local updates as much as possible and correct the aggregated global model update direction towards the optimal global model update, thereby greatly reducing the deflection degree of malicious updates on the global update direction. Specifically, the present invention proposes four independent modules to define the spatial characteristics of maliciousness, and any update belonging to these direction spaces will be rejected. First, the present invention proposes a Sybil direction elimination module to prevent malicious attackers from colluding; secondly, in order to reduce the non-iid degree of FL, the present invention designs a strategy of clustering first and then grouping, so that the similarity between normal updates is greatly improved, thereby allowing the present invention to use a similarity detection method to filter out updates that are inconsistent with the majority. In addition, the present invention uses a relatively small auxiliary dataset that does not need to be related to the local dataset to evaluate the performance of each update to limit the normal direction space. Any update with poor performance on the auxiliary dataset will be discarded. Finally, the present invention proposes an equal modulus aggregation module to further correct malicious updates that can escape the first three modules, thereby greatly reducing the impact of malicious directions with a small deviation from the normal update direction on the global update direction.

[0104] The present invention is an aggregation scheme that can resist Byzantine attacks in the federated learning data heterogeneous scenario. The functions of this scheme are realized by a cloud server and a client, a total of 2 participants and 4 steps. The system architecture diagram of the scheme applicable to resisting Byzantine attacks in the federated learning data heterogeneous scenario involved in the present invention is as Figure 1 shown, and the specific implementation steps of the scheme described in the present invention will be introduced below in combination with Figure 2 the following.

[0105] As Figure 2 shown, the present invention provides a method for defending against Byzantine attacks in the federated learning scenario, including the following steps:

[0106] (1) The server obtains a second dataset D different from the first dataset obtained by the client aux and the number of classifications of the server model, preprocesses the second dataset to obtain a preprocessed second dataset, expands the obtained number of classifications of the server model to obtain an expanded server model, and sends the preprocessed second dataset and the expanded server model to the client.

[0107] Specifically, in this step, first, the size of the second dataset is modified to the input size of the server model, and then the labels of the second dataset are marked as L, L+1, L+2,..., L+K; where L represents the number of data categories of the client, and K is the number of extended classifications. Finally, the number of classifications of the server model is set to L+K. In the solution of the present invention, generally setting K to 1 can effectively detect malicious client models.

[0108] The server model in this step is a deep learning model, such as CNN, ResNet, etc.

[0109] (2) The client obtains the first dataset and divides this dataset into a training set and a test set according to a ratio of 8:2. The client obtains a second dataset that is completely different from the preprocessed first dataset from the server, modifies the size of the second dataset to the input size of the client model, and similarly marks the labels of the second dataset as L, L+1, L+2,..., L+K; where L represents the number of data categories of the client, and K is the number of extended classifications. Then, the training set of the first dataset and the second dataset are evenly mixed to obtain a mixed training set, and the weights of the client model are initialized according to the weights of the extended server model to obtain an initialized client model.

[0110] The client model in this step is the same deep learning model as the server model, such as CNN, ResNet, etc.

[0111] (3) The client inputs the mixed training set into the client model initialized in step (2) to obtain a client gradient set C={g1, g2,…, g m}, and sends it to the server, where the range of m is 1≤m≤n (n is the total number of clients participating in federated learning), preferably n (corresponding to the case where all clients participate), represents the gradient of the i1-th client among all clients participating in federated learning, and i1∈[1, m].

[0112] (4) The server performs four client gradient anomaly checks and malicious client gradient elimination processing on the client gradients obtained in step (3) in sequence, and aggregates the remaining client gradients to obtain updated client gradients. The server model is updated for the second time using the updated client gradients to obtain the server model of the second round,... This process is repeated for T rounds (where T≥1000, the larger its value, the better the convergence of the server model, but the corresponding computational amount will also increase) to obtain the final server model.

[0113] Specifically, step (4) includes the following sub-steps:

[0114] (4-1) The server randomly selects m clients and receives the client gradient set C = {g1, g2, …, g m}, where represents the gradient of the i1-th client among all the clients participating in federated learning, and i1 ∈ [1, m];

[0115] (4-2) The server filters the client gradient set C = {g1, g2, …, g m} obtained in step (4-1) to remove malicious client gradients with similar or identical directions, and obtains the client gradient set C1 after the first detection;

[0116] Specifically, this step includes the following sub-steps:

[0117] (4-2-1) The server calculates the cosine similarity between all pairs of client gradients in the client gradient set {g1, g2, …, g m} to obtain the similarity score set {s1, s2, …, s m};

[0118] Specifically, the similarity score in this step is calculated by the formula:

[0119]

[0120] where represents the set of the top client gradients in the client gradient set C that have the highest similarity with . The similarity between and is measured by cosine similarity; <,> represents the inner product, and j ∈ [1, m].

[0121] (4-2-2) The server calculates the median m of the similarity score set {s1, s2, …, s

[0122] (4-2-3) The server uses the median obtained in step (4-2-2) to obtain the distance set of the similarity scores from the median and obtains the median

[0123] (4-2-4) Set the counter i1 = 1 and initialize an empty malicious client gradient set C mal ;

[0124] (4-2-5) Determine whether i1 is less than or equal to m. If so, go to step (4-2-6); otherwise, go to step (4-2-9);

[0125] (4-2-6) Determine the similarity score whether it is greater than If so, go to step (4-2-7); otherwise, go to step (4-2-8);

[0126] The purpose of this step is to obtain a suitable threshold based on the median of the similarity scores and the median of the distance between the similarity scores and the median to indicate the range of normal similarity scores. Specifically, in the present invention, the similarity scores less than or equal to the threshold are considered normal similarity scores, and the corresponding client gradients are normal, while the similarity scores greater than the threshold are considered abnormal similarity scores, and the corresponding client gradients are determined as malicious client gradients.

[0127] (4-2-7) Put the client gradient corresponding to the similarity score into the malicious client gradient set C mal ;

[0128] (4-2-8) Set i1 = i1 + 1, and return to step (4-2-5);

[0129] (4-2-9) The server uses the malicious client gradient set C mal obtained in step (4-2-7) to remove the malicious client gradients in the client gradient set C obtained in step (4-1), so as to obtain the client gradient set C1 = C - C mal ;

[0130] The advantages of the above sub-steps (4-2-1) to sub-step (4-2-6) are that by using the relative cosine similarity and their real-time median and distance, it is possible to completely remove the exactly the same or relatively similar malicious client gradients without wrongly removing the normal client gradients;

[0131] (4-3) The server performs a second detection process on the client gradient set C1 obtained in step (4-2) to obtain a client gradient set C2 after removing the malicious client gradients that are not similar to the normal client gradient direction;

[0132] Specifically, this step includes the following sub-steps:

[0133] (4-3-1) The server calculates the cosine similarity between all client gradients in the client gradient set C1 after the first detection to obtain a set of similarity scores where represents the i2-th client gradient in the client gradient set C1 after the first detection, represents the i2-th similarity score in the set of similarity scores and i2 ∈ [1, |C1|];

[0134] Specifically, in this step, the similarity score is calculated as follows:

[0135]

[0136] where represents the set of the top client gradients in the client gradient set C1 after the first detection that are most similar to ;

[0137] (4-3-2) The server performs Mean shift clustering on the client gradients in the client gradient set C1 according to the set of similarity scores obtained in step (4-3-1) to obtain h clusters {Ω1, Ω2, …, Ω h}; where h can be determined by the mean shift clustering algorithm itself, represents the i3-th cluster set, and the elements in this cluster set are client gradients, and i3 ∈ [1, h];

[0138] (4-3-3) The server obtains the number of elements a of the cluster with the most elements among the h clusters obtained in step (4-3-2);

[0139] (4-3-4) Set the counter i4 = 1 and initialize an empty set

[0140] (4-3-5) Determine whether i4 is greater than or equal to a. If so, go to step (4-3-11); otherwise, go to step (4-3-6);

[0141] (4-3-6) Set the counter i3 = 1;

[0142] (4-3-7) Determine whether i3 is greater than or equal to h. If so, go to step (4-3-10); otherwise, go to step (4-3-8);

[0143] (4-3-8) Determine whether is empty. If so, go to step (4-3-9); otherwise, the client gradient Add the empty set Then go to step (4-3-9);

[0144] (4-3-9) Set i3 = i3 + 1, and return to step (4-3-7);

[0145] (4-3-10) Set i4 = i4 + 1, and return to step (4-3-5);

[0146] (4-3-11) Initialize three empty sets D1, D2, D mal ;

[0147] (4-3-12) Set the counter i4 = 1;

[0148] (4-3-13) Determine whether i4 is greater than or equal to a. If so, go to step (4-3-16); otherwise, go to step (4-3-14);

[0149] (4-3-14) Set D1[i4] to the average value of all client gradients in the group ;

[0150] (4-3-15) Set i4 = i4 + 1, and return to step (4-3-13);

[0151] (4-3-16) The server calculates the cosine similarity between all client gradients in the client gradient set D1 pairwise to obtain a similarity score set where represents the i5th client gradient in the client gradient set D1, represents the i5th similarity score in the relative cosine similarity score set and i5 ∈ [1, |D1|]; Specifically, the calculation formula for this step is:

[0152]

[0153]

[0154] where represents the set of the first client gradients in the client gradient set D1 that are most similar to ;

[0155] (4-3-17) The server performs mean shift clustering on the result obtained in step (4-3-16) to obtain the index corresponding to the largest cluster, and marks the client gradient set corresponding to this index as Ψ. Note that at this time, the client gradients in Ψ are the most similar part in D1; ​​

[0156] (4-3-18) All client gradients in D1 but not in Ψ are regarded as malicious client gradients and put into the malicious client gradient set D mal middle;

[0157] The advantage of the above sub-steps (4-3-14) to (4-3-18) is that the cosine similarity between the centroids (i.e., average values) of each group of client gradients can improve the success rate of detecting malicious client gradients. This is because the centroids of normal client gradients usually have more consistent update targets, that is, they all have good test accuracy on the first data set of all clients, and thus have similar update directions. Therefore, as long as there is a group of client gradients that want to destroy the test accuracy on the first data set, it means that its centroid must be updated in a malicious direction, so it is easy to be dissimilar in direction to the centroid of the normal group.

[0158] (4-3-19) Set counter i4=1;

[0159] (4-3-20) Determine whether i4 is greater than or equal to a. If so, proceed to step (4-3-24); otherwise, proceed to step (4-3-21);

[0160] (4-3-21) Determine whether D1[i4] is D mal A client gradient in , if yes, go to step (4-3-22), otherwise go to step (4-3-23);

[0161] (4-3-22) Group All client gradients are put into set D2;

[0162] (4-3-23) Set i4=i4+1 and return to step (4-3-20);

[0163] (4-3-24) The server performs a cosine similarity calculation between each of the client gradient set C1 obtained in step (4-3-7) and all the client gradients in the client gradient set D2 obtained in step (4-3-22) to obtain a similarity score set. in Represents a similarity score set The i6th similarity score in , and i6∈[1,|D2|];

[0164] Specifically, the calculation formula for this step is:

[0165]

[0166] in represents the most similar predecessor in the client gradient set C1 to D2[i6] A set of client gradients.

[0167] (4-3-25) Re-initialize an empty set C of malicious client gradients mal , and set a hyperparameter β, where 0 ≤ β ≤ 1, and β = 0.2 is preferably used in this solution;

[0168] (4-3-26) Put the first β·|D2| client gradients with the lowest similarity scores from the client gradient set D2 into the set C of malicious client gradients mal ;

[0169] (4-3-27) The server uses the set C of malicious client gradients obtained in step (4-3-26) mal to remove the malicious client gradients from the client gradient set C1 to obtain the client gradient set C2 = C1 - C after the second detection mal ;

[0170] (4-4) The server performs a third detection process on the client gradient set C2 obtained in step (4-3) to obtain the client gradient set C3 after further removing the malicious client gradients whose directions are closer to the normal client gradients;

[0171] Specifically, the steps of the third detection are as follows:

[0172] (4-4-1) The server performs gradient descent processing on each client gradient in the client gradient set C2 obtained in step (4-3-27) to obtain the client model corresponding to each client gradient where represents the i7-th client gradient in the client gradient set C2, represents the corresponding client model, and i7 ∈ [1, |C2|];

[0173] Specifically, the calculation formula for this step is as follows:

[0174]

[0175] where w is the server model and η is the learning rate.

[0176] (4-4-2) The server uses the second dataset D aux to calculate the loss of the client model w obtained in step (4-4-1) i to obtain the evaluated loss change value ρ i ;

[0177] Specifically, the calculation formula for this step is as follows:

[0178]

[0179] wherein is the loss function.

[0180] The advantage of this step is that using the change value of the loss instead of the loss value can be more sensitive and robust to model poisoning manipulation.

[0181] (4-4-3) The server performs an outlier selection on the loss change value obtained in step (4-4-2) to obtain a set C of malicious client gradients with a change less than the threshold λ mal ;

[0182] (4-4-4) The server uses the set C of malicious client gradients obtained in step (4-4-3) mal to screen the set C2 of client gradients obtained in (4-3-26) to obtain a set C3 of normal client gradients. The calculation formula for this step is C3 = C2 - C mal ;

[0183] (4-5) The server performs a fourth detection on the set C3 of normal client gradients obtained in step (4-4) to correct the norm of each client gradient therein, thereby obtaining the corrected client gradient where i8 ∈ [1, |C3|];

[0184] Specifically, the execution steps of the fourth detection are as follows:

[0185] (4-5-1) The server calculates the norm of the set C3 of client gradients obtained in step (4-4-4) to obtain the median f of the norms of all client gradients;

[0186] (4-5-2) The server divides the norm greater than the median f in step (4-5-1) by f and puts it into the empty set D, and for the norm less than the median f, divides f by these norms and puts it into the empty set D to obtain a set D for judging norm outliers;

[0187] The advantage of this step is that using the median as a benchmark to judge the outliers of the norms of other client gradients is more accurate, and at the same time, converting all norms less than the median into multiples of the median norm can simplify the number of hyperparameters used in the removal process.

[0188] (4-5-3) The server statistically processes the norms in D to obtain the standard deviation s and the mean 1 of all norms;

[0189] (4-5-4) The server removes all client gradients with norms greater than 8(s + l) after being processed in step (4-5-2) from C3 to obtain a set C3 of client gradients excluding malicious client gradients with very large and very small norms;

[0190] The advantage of this step is that it can adaptively eliminate abnormal modulus lengths according to the mean and standard deviation.

[0191] (4-5-5) The server calculates the modulus lengths of the client gradient set C3 obtained in step 4-5-4 to obtain the average modulus length e of all client gradients in C3;

[0192] (4-5-6) The server corrects the modulus lengths of all client gradients in C3 using e obtained in step (4-5-5);

[0193] Specifically, the modulus length correction formula for this step is as follows, for all client gradients in C3 There is:

[0194]

[0195] The advantage of this step is that by correcting the modulus lengths to the same length, the ability of normal client gradients to deflect the direction of malicious client gradients can be greatly improved, thus realizing corrective aggregation.

[0196] (4-6) The server performs aggregation processing on all the corrected client gradients obtained in step (4-5) to obtain the latest server gradient g, and thereby update to obtain the latest server model;

[0197] Specifically, the aggregation formula adopted in this step is as follows:

[0198]

[0199] (4-7) Repeat steps (4-1) to (4-6) for T rounds to obtain the final server model.

[0200] Those skilled in the art can easily understand that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for defending against Byzantine attacks in a federated learning scenario, characterized in that, It includes the following steps: (1) The server obtains a second dataset D that is different from the first dataset obtained by the client aux and the number of classifications of the server model, preprocesses the second dataset to obtain a preprocessed second dataset, expands the obtained number of classifications of the server model to obtain an expanded server model, and sends the preprocessed second dataset and the expanded server model to the client; (2) The client obtains the first data set and divides the data set into a training set and a test set according to a ratio of 8:

2. The client obtains a second data set that is completely different from the first data set after preprocessing from the server, modifies the size of the second data set to the input size of the client model, and also marks the labels of the second data set as L, L+1, L+2,..., L+K; where L represents the number of data categories of the client, and K is the number of extended classifications; then evenly mixes the training set of the first data set and the second data set to obtain a mixed training set, and initializes the weights of the client model according to the weights of the extended server model to obtain an initialized client model. (3) The client inputs the mixed training set into the client model initialized in step (2) to obtain a client gradient set C = {g1, g2, …, g m} for updating the client model, and sends it to the server, where the range of m is 1 ≤ m ≤ n, and n is the total number of clients participating in federated learning. represents the gradient of the i1-th client among all clients participating in federated learning, and i1 ∈ [1, m]; (4) The server performs four client gradient anomaly checks and malicious client gradient elimination processing on the client gradients obtained in step (3) in sequence, and aggregates the remaining client gradients to obtain updated client gradients, and uses the updated client gradients to perform a second-round update on the server model to obtain the second-round server model,..., repeat this process for T rounds to obtain the final server model, where T≥1000; step (4) includes the following sub-steps: (4-1) The server randomly selects m clients and receives the client gradient set C = {g1, g2, …, g m} represents the gradient of the i1-th client among all the clients participating in federated learning, and i1 ∈ [1, m]; (4-2) The server filters the client gradient set C = {g1, g2, …, g m} to remove malicious client gradients with similar or identical directions, and obtains the client gradient set C1 after the first detection; (4-3) The server performs a second detection process on the set C1 of client gradients after the first detection obtained in step (4-2) to obtain a set C2 of client gradients after eliminating malicious client gradients that are not similar to the normal client gradient direction. (4-4) The server performs a third detection process on the set C2 of client gradients obtained in step (4-3) to obtain a set C3 of client gradients after further eliminating malicious client gradients whose directions are closer to the normal client gradient. (4-5) The server performs a fourth detection on the normal client gradient set C3 obtained in step (4-4) to correct the norm of each client gradient therein, thereby obtaining the corrected client gradients where i8 ∈ [1, |C3|]; (4-6) The server performs an aggregation process on all the corrected client gradients obtained in step (4-5) to obtain the latest server gradient g, and thereby updates to obtain the latest server model; (4-7) Repeat steps (4-1) to (4-6) for T rounds to obtain the final server model.

2. The method for defending against Byzantine attacks in the federated learning scenario according to claim 1, wherein In step (1), first, the size of the second data set is modified to the input size of the server model, and then the labels of the second data set are marked as L, L+1, L+2,..., L+K; where L represents the number of data categories of the client, and K is the number of extended classifications; finally, the number of classifications of the server model is set to L+K. The server model is the same deep learning model as the client model.

3. The method for defending against Byzantine attacks in the federated learning scenario according to claim 1 or 2, characterized in that, Step (4) includes the following sub-steps: (4-2-1) The server calculates the cosine similarity between every pair of client gradients in the client gradient set {g1, g2, …, g m} to obtain a similarity score set {s1, s2, …, s m}; Specifically, the similarity score in this step is calculated as follows: Among them represents the set of the top client gradients in the client gradient set C that have the highest similarity to ; where the similarity to is measured by cosine similarity; <,> represents the inner product, and j ∈ [1, m]; (4-2-2) The median of the set of similarity scores {s1, s2, …, s m} obtained in the server calculation step (4-2-1) (4-2-3) Steps for the server to use the median obtained in (4-2-2) Obtain the similarity score and the median of the distance set and obtain the median of this distance set (4-2-4) Set the counter i1 = 1 and initialize an empty malicious client gradient set C mal ; (4-2-5) Determine whether i1 is less than or equal to m. If so, go to step (4-2-6); otherwise, go to step (4-2-9). (4-2-6) Determine the similarity score whether it is greater than If yes, go to step (4-2-7); otherwise, go to step (4-2-8). (4-2-7) Put the similarity score The corresponding client gradient into the malicious client gradient set C mal ; (4-2-8) Set i1 = i1 + 1 and return to step (4-2-5). (4-2-9) The server uses the malicious client gradient set C obtained in step (4-2-7) mal to remove the malicious client gradients in the client gradient set C obtained in step (4-1), so as to obtain the client gradient set C1 = C - C after the first detection mal .

4. The method for defending against Byzantine attacks in the federated learning scenario according to claim 3, wherein, Step (4-3) includes the following sub-steps: (4-3-1) The server calculates the cosine similarity between all client gradients in the client gradient set C1 after the first detection to obtain a similarity score set where represents the i2-th client gradient in the client gradient set C1 after the first detection, represents the i2-th similarity score in the similarity score set and i2 ∈ [1, |C1|]; Specifically, in this step, the similarity score is calculated as follows: Among them represents the set of the top most similar client gradients in the client gradient set C1 after the first detection; client gradients (4-3-2) The server obtains the similarity score set according to step (4-3-1). Perform Mean shift clustering on the client gradients in the client gradient set C1 to obtain h clusters {Ω1, Ω2, …, Ω h}; where h can be determined by the mean shift clustering algorithm itself. represents the i3-th cluster set, the elements in this cluster set are client gradients, and i3 ∈ [1, h]; (4-3-3) The server obtains the number of elements a in the cluster with the most elements among the h clusters obtained in step (4-3-2). (4-3-4) Set the counter i4 = 1 and initialize an empty set (4-3-5) Determine whether i4 is greater than or equal to a. If so, go to step (4-3-11); otherwise, go to step (4-3-6). (4-3-6) Set the counter i3 = 1. (4-3-7) Determine whether i3 is greater than or equal to h. If so, go to step (4-3-10); otherwise, go to step (4-3-8); (4-3-8) Judgment Check if it is empty. If so, go to step (4-3-9); otherwise, add the client gradient to the empty set and then go to step (4-3-9). (4-3-9) Set i3 = i3 + 1 and return to step (4-3-7); (4-3-10) Set i4 = i4 + 1 and return to step (4-3-5); (4-3-11) Initialize three empty sets D1, D2, D mal ; (4-3-12) Set the counter i4 = 1; (4-3-13) Determine whether i4 is greater than or equal to a. If so, go to step (4-3-16); otherwise, go to step (4-3-14); (4-3-14) Set D1[i4] to the average of all client gradients in the group ; (4-3-15) Set i4 = i4 + 1 and return to step (4-3-13); (4-3-16) The server calculates the cosine similarity between all client gradients in the client gradient set D1 to obtain a similarity score set where represents the i5-th client gradient in the client gradient set D1, represents the i5-th similarity score in the relative cosine similarity score set and i5 ∈ [1, |D1|]; ​ Specifically, the calculation formula for this step is: Among them represents the set of the top client gradients in the client gradient set D1 that are most similar; (4-3-17) The server performs mean shift clustering on the obtained in step (4-3-16) to obtain the index corresponding to the largest cluster, and marks the client gradient set corresponding to this index as Ψ; Note that at this time, the client gradient in Ψ is the most similar part in D1; (4-3-18) Treat all client gradients that are in D1 but not in Ψ as malicious client gradients and put them into the malicious client gradient set D mal ; (4-3-19) Set the counter i4 = 1; (4-3-20) Determine whether i4 is greater than or equal to a. If so, go to step (4-3-24); otherwise, go to step (4-3-21); (4-3-21) Determine whether D1[i4] is a client gradient in D mal If it is, go to step (4-3-22); otherwise, go to step (4-3-23). (4-3-22) Put all the client gradients in the group into the set D2; (4-3-23) Set i4 = i4 + 1 and return to step (4-3-20); (4-3-24) The server calculates the cosine similarity between all pairs of client gradients in the client gradient set C1 obtained in step (4-3-7) and the client gradient set D2 obtained in step (4-3-22) to obtain a similarity score set where represents the i6-th similarity score in the similarity score set and i6 ∈ [1, |D2|]; Specifically, the calculation formula for this step is: Among them represents the set of the top client gradients in the client gradient set C1 that are most similar to D2[i6]; (4-3-25) Re-initialize an empty malicious client gradient set C mal , and set a hyperparameter β, where 0 ≤ β ≤ 1; (4-3-26) Put the first β·|D2| client gradients with the lowest similarity scores from the client gradient set D2 into the malicious client gradient set C mal ; (4-3-27) The server uses the malicious client gradient set C obtained in step (4-3-26) mal to remove the malicious client gradients in the client gradient set C1, so as to obtain the client gradient set C2 = C1 - C after the second detection mal .

5. The method for defending against Byzantine attacks in the federated learning scenario according to claim 4, wherein Step (4-4) includes the following sub-steps: (4-4-1) The server performs gradient descent processing on each client gradient in the client gradient set C2 obtained in step (4-3-27) to obtain the client model corresponding to each client gradient where represents the i7-th client gradient in the client gradient set C2, represents the corresponding client model, and i7 ∈ [1, |C2|]; Specifically, the calculation formula for this step is as follows: Where w is the server model and η is the learning rate; (4-4-2) The server uses the second dataset D aux to calculate the loss of the client model w obtained in step (4-4-1) i so as to obtain the evaluated loss change value ρ i ; Specifically, the calculation formula for this step is as follows: wherein is the loss function; (4-4-3) The server performs abnormal selection on the loss change value obtained in step (4-4-2) to obtain a set C of malicious client gradients with a change less than the threshold λ mal ; (4-4-4) The server uses the malicious client gradient set C obtained in step (4-4-3). mal Perform screening processing on the client gradient set C2 obtained in (4-3-26) to obtain a normal client gradient set C3; the calculation formula for this step is C3 = C2 - C. mal .

6. The method for defending against Byzantine attacks in the federated learning scenario according to claim 5, wherein Step (4-5) includes the following sub-steps (4-5-1) The server calculates the norm of the client gradient set C3 obtained in step (4-4-4) to obtain the median f of the norms of all client gradients; (4-5-2) The server divides the norms greater than the median f in step (4-5-1) by f and puts them into the empty set D. For the norms less than the median f, it divides f by these norms and puts them into the empty set D to obtain the set D for judging norm anomalies; (4-5-3) The server statistically processes the norms in D to obtain the standard deviation s and the mean l of all norms; (4-5-4) The server removes all client gradients with norms greater than 8(s + l) after the processing in step (4-5-2) from C3 to obtain the client gradient set C3 after removing malicious client gradients with very large and very small norms; (4-5-5) The server calculates the norm of the client gradient set C3 obtained in step 4-5-4 to obtain the average norm e of all client gradients in C3; (4-5-6) The server uses e obtained in step (4-5-5) to correct the norms of all client gradients in C3; Specifically, the formula for norm correction in this step is as follows. For all client gradients in C3 we have:

7. A system for defending against Byzantine attacks in a federated learning scenario, characterized in that, Includes: The first module, which is set on the server side, is used to obtain a second data set D different from the first data set obtained by the client aux and the number of classifications of the server model, preprocess the second data set to obtain a preprocessed second data set, expand the obtained number of classifications of the server model to obtain an expanded server model, and send the preprocessed second data set and the expanded server model to the client; The second module, which is set on the client side and is used to obtain the first data set and divide this data set into a training set and a test set according to a ratio of 8:2; Obtain a second data set that is completely different from the first data set after preprocessing from the server side, modify the size of the second data set to the input size of the client model, and also label the labels of the second data set as L, L + 1, L + 2,..., L + K; where L represents the number of data categories of the client, and K is the number of extended classifications; then evenly mix the training set of the first data set and the second data set to obtain a mixed training set, and initialize the weights of the client model according to the weights of the extended server model to obtain an initialized client model; The third module, which is set on the client side, is used to input the mixed training set into the client model initialized by the second module to obtain a client gradient set C = {g1, g2, …, g m}, and send it to the server, where the range of m is 1 ≤ m ≤ n, and n is the total number of clients participating in federated learning. represents the gradient of the i1-th client among all clients participating in federated learning, and i1 ∈ [1, m]; The fourth module, which is set on the server side, is used to perform four client gradient anomaly checks and malicious client gradient elimination processing on the client gradients obtained by the third module in sequence, and aggregate the remaining client gradients to obtain updated client gradients, and use the updated client gradients to perform a second-round update on the server model to obtain the server model of the second round,..., repeat this process for T rounds to obtain the final server model, where T ≥ 1000; the fourth module includes the following sub-modules: The first sub-module is set on the server side and is used to randomly select m clients and receive the client gradient set C = {g1, g2, …, g m}, represents the gradient of the i1-th client among all the clients participating in the federated learning, and i1 ∈ [1, m]; The second sub-module, which is set on the server side, is used to filter the client gradient set C = {g1, g2, …, g m} to remove malicious client gradients with similar or identical directions, and obtain the client gradient set C1 after the first detection; The third sub-module, which is set on the server side, is used to perform a second detection process on the client gradient set C1 after the first detection obtained by the second sub-module to obtain a client gradient set C2 after eliminating malicious client gradients that are not similar to the normal client gradient direction; The fourth sub-module, which is set on the server side, is used to perform a third detection process on the client gradient set C2 obtained by the third sub-module to obtain a client gradient set C3 after further eliminating malicious client gradients that are closer to the normal client gradient direction; The fifth sub-module, which is arranged on the server side, is used to perform a fourth detection on the normal client gradient set C3 obtained by the fourth sub-module, so as to perform norm correction on each client gradient therein, thereby obtaining the corrected client gradient where i8 ∈ [1, |C3|]; The sixth sub-module, which is set on the server side and is used to aggregate all the corrected client gradients obtained by the fifth sub-module to obtain the latest server-side gradient g, and thereby update to obtain the latest server-side model; (4-7) Repeat steps (4-1) to (4-6) for T rounds to obtain the final server model.