Security federal aggregation algorithm for solving data heterogeneity

Through dynamic local aggregation and adaptive hierarchical gradient cutting methods, poor generalization and privacy leakage problems under Non-IID data heterogeneity are solved, and efficient and accurate model training and privacy protection are achieved.

CN120373420APending Publication Date: 2025-07-25LIAONING UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510438934.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the case of Non-IID data heterogeneity, existing federated learning algorithms have problems such as poor generalization of model, low accuracy, privacy leakage and high communication cost, and existing methods cannot effectively capture the unique data characteristics of the client or ignore privacy protection when processing Non-IID data.

Method used

Dynamic local aggregation and adaptive hierarchical gradient cropping methods are adopted to dynamically adjust local aggregation weights and hierarchical gradient cropping, and combined with noise processing, optimize model training efficiency and accuracy to protect user privacy.

Benefits of technology

It improves the efficiency and accuracy of model training, solves the problem of data heterogeneity, and improves the generalization ability and convergence speed of the model while protecting user privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373420A_ABST
    Figure CN120373420A_ABST
Patent Text Reader

Abstract

The invention discloses a security federal aggregation algorithm for solving data heterogeneity, which comprises the following steps that: step 1, a server initializes a global model and distributes the global model to clients, and the clients respectively carry out a round of iteration on the global model to obtain a model gradient after the first round of updating and upload the model gradient to the server; step 2, the server distributes the model gradient after the first round of updating to a client subset, the client subset independently uses own local data to carry out iterative training on the model, and an updated local model is obtained after the iterative training is finished; and step 3, after iteration of the clients is completed, the updated local gradient is sent to the server, and the server aggregates the local model and sends the obtained global model gradient to all the clients for a new round of client local training until the global model meeting the convergence condition is obtained. The method has the advantages that key features can be captured, privacy protection is improved, and operation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data privacy protection. More specifically, the present invention relates to a secure federated aggregation algorithm for solving data heterogeneity. Background Art

[0002] In the Non-IID data heterogeneity scenario, if all the information of the global model is used to update the local model parameters, this may cause problems such as poor generalization of the model and low accuracy of the model. In addition, in CNN, higher layers know more client-specific data than lower layers, and attackers can easily attack user data. Therefore, there are problems of local model aggregation and privacy leakage in the Non-IID data heterogeneity scenario. Moreover, when processing Non-IID data, if the gradient parameters of the previous round are used to overwrite the local model parameters of the current round, this will not only lead to a decline in model performance, but also bring a large communication cost.

[0003] In the face of the challenge of non-independent and identically distributed (Non-IID) data distribution, some research uses the basic federated learning algorithm (FedAvg). Due to its inherent limitations, it may not be able to fully capture the unique data characteristics of each client, resulting in poor performance of the global model on some clients, and thus increasing the loss value of the model. Some research uses the adaptive personalized federated learning APFL algorithm. By introducing a personalized model learning mechanism, the model update of each client is not only affected by the global model, but also incorporates the adjustment of the local model, so as to more accurately adapt to their respective data distributions. However, it ignores the protection of user privacy data. Some research uses the adaptive differential privacy algorithm ADP-FL by implementing adaptive gradient clipping for each client in different training rounds and dynamically adjusting the noise level to reduce the negative impact of differential privacy on model performance, and successfully improves the accuracy of the model. However, in the face of extremely non-independent and identically distributed (Non-IID) data scenarios, the performance in training is not satisfactory. Some research uses the differential privacy algorithm DP-FL with a fixed level of noise to add a fixed level of noise to the client model parameters to achieve differential privacy. However, this method may lead to the accumulation of noise during the training process, thereby affecting the convergence speed and generalization ability of the model. Summary of the Invention

[0004] The object of the present invention is to design and develop a secure federated aggregation algorithm for solving data heterogeneity. By combining dynamic local aggregation and adaptive hierarchical gradient clipping methods, the training efficiency and accuracy are improved, and thus the privacy of the algorithm is enhanced.

[0005] The technical solution provided by the present invention is as follows:

[0006] A secure federated aggregation algorithm for solving data heterogeneity, comprising the following steps:

[0007] Step 1: The server initializes the global model and distributes it to N clients. The N clients independently use their own local data to perform one round of iteration on the global model, obtain the model gradients after the first round of update, and upload them to the server;

[0008] Step 2: The server distributes the model gradients after the first round of update to a subset of clients. The subset of clients independently use their own local data to perform iterative training on the model. After the iterative training is completed, an updated local model is obtained;

[0009] Step 3: After the client iteration is completed, the updated local gradients are sent to the server. The server aggregates the local models and sends the obtained global model gradients to all clients for a new round of client local training until a global model that meets the convergence condition is obtained.

[0010] Preferably, the specific steps of Step 2 include:

[0011] Step 1: Randomly select γ·N clients to generate a subset of clients Γ;

[0012] where γ is the proportion of local training;

[0013] Step 2: The server distributes the model gradients θ1 after the first round of update to each member in the subset of clients Γ;

[0014] Step 3: Each member in the subset of clients Γ performs iterative training. When t≥2 for the client, in the local aggregation process of the client, aggregation is achieved by dynamically adjusting the weights of data elements;

[0015] Step 4: Obtain the local initialization model of the i-th client according to the local aggregation weights;

[0016] Step 5: The client selects δ of the data from the local dataset for training to obtain local model gradients;

[0017] Step 6: Perform a layering operation on the local training model of the client. Divide the first h layers of the local training model into personalized layers, and the remaining layers are frozen layers;

[0018] Step 7: If the layer of the local model of the client is a personalized layer, update the local gradient of the i-th client;

[0019] If the layer of the local model of the client is a frozen layer, keep the local gradient of the client unchanged.

[0020] Preferably, the local aggregation weights in Step 3 satisfy:

[0021]

[0022] Wherein, is the aggregated weight after the t-th training update of the i-th client, ω i,t is the aggregated weight before the t-th training update of the i-th client, ρ is the local learning rate, is the gradient of the loss function L with respect to ω i,t , L(*) is the loss function, is the locally updated model gradient after adding noise in the (t-1)-th training of the i-th client, D i,t,Γ is the dataset of the i-th client in the t-th training, θ t-1 is the global model gradient at the (t-1)-th time.

[0023] Preferably, step 3 further includes clipping the local aggregated weight;

[0024]

[0025] Wherein, σ(ω) is the clipped aggregated weight, and ω is the aggregated weight before clipping of client i in the t-th training.

[0026] Preferably, the local initialized model is:

[0027]

[0028] Wherein, is the locally initialized updated model gradient of the i-th client in the t-th training, and ⊙ is the Hadamard product (element-wise product).

[0029] Preferably, the local model gradient is:

[0030]

[0031] Wherein, θ i,t is the local model update gradient of the i-th client before adding noise in the t-th training, μ is the global learning rate, is the gradient of the loss function L with respect to the parameter .

[0032] Preferably, the gradient clipping threshold satisfies:

[0033]

[0034] Wherein, C i,t is the gradient clipping threshold of the i-th client in the t-th training, C i,t-1is the gradient clipping threshold for the i-th client in the (t-1)-th training, τ is the gradient clipping factor, θ i,t (x i ) is the local training model of the i-th client in the t-th training, θ i,t-1 (x i ) is the local training model of the i-th client in the t-th training, ||·||2 is the gradient l2 norm, β i,t-1 is the amount of noise for the i-th client in the (t-1)-th training, |Γ| is the number of clients participating in the training.

[0035] Preferably, the amount of noise satisfies:

[0036] β i,t = η·β i,t-1 ;

[0037] In the formula, β i,t is the amount of noise for the i-th client in the t-th training, η is the attenuation rate, η ∈ (0, 1).

[0038] Preferably, the local gradient of the i-th client is updated as:

[0039]

[0040] In the formula, is the local model gradient of the i-th client in the t-th training after adding noise, β t is the noise scale in the t-th training, C i,t is the gradient clipping threshold.

[0041] Preferably, the server aggregates the local models as:

[0042]

[0043] In the formula, θ t represents the global model gradient at the t-th training.

[0044] The beneficial effects of the present invention are:

[0045] (1), A secure federated aggregation algorithm for solving data heterogeneity designed and developed by the present invention not only trains personalized local models according to the unique data characteristics of clients, but also adds noise to the sensitive data of users to prevent it from being attacked, solving the most intractable data heterogeneity problem in federated learning.

[0046] (2) The secure federated aggregation algorithm designed and developed in the present invention to address data heterogeneity adopts a novel hierarchical gradient clipping strategy for the personalized layer of the model, aiming to mitigate the potential impact of the noise introduced by differential privacy on the model performance. While protecting privacy, this strategy effectively improves the training efficiency and accuracy of the model through refined gradient clipping optimization. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 Schematic diagram comparing the model accuracy of the secure federated aggregation algorithm for addressing data heterogeneity of the present invention with baseline algorithm 1 and baseline algorithm 2 in Cifar10.

[0048] Figure 2 Schematic diagram comparing the model loss values of the secure federated aggregation algorithm for addressing data heterogeneity of the present invention with baseline algorithm 1 and baseline algorithm 2 in Cifar10.

[0049] Figure 3 Schematic diagram comparing the model accuracy of the secure federated aggregation algorithm for addressing data heterogeneity of the present invention with baseline algorithm 3 and baseline algorithm 4 in Cifar10. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] The following further elaborates on the present invention in detail to enable those skilled in the art to implement it with reference to the text of the specification.

[0051] A secure federated aggregation algorithm for addressing data heterogeneity (SFA-DH) provided by the present invention includes the following steps:

[0052] Step 1: The server initializes the global model and distributes it to N clients. The N clients independently use their own local data to perform one round of iteration on the global model, obtain the model gradients after the first round of update, and upload them to the server;

[0053] Among them, the global model is:

[0054]

[0055] In the formula, θ0 is the global model, θ i,t is the model gradient of the i-th client in the t-th training, i is the client number, i = 1, 2, 3…, N, t is the iteration number, t = 1, 2, 3,…, T;

[0056] The client performs iterative training on the global model through the CNN network;

[0057] Step 2: The server distributes the model gradients after the first round of update to some selected clients. Each of these clients independently uses its own local data to iteratively train the model. After the iterative training is completed, an updated local model is obtained. The specific training process includes the following steps:

[0058] Step 1: Randomly select γ·N clients to generate a client subset Γ;

[0059] where γ is the proportion of local training;

[0060] Step 2: The server distributes the model gradients θ1 after the first round of update to each member in the client subset Γ;

[0061] Step 3: Each member in the client subset Γ performs iterative training. When t≥2 for a client, during the local aggregation process of the client, aggregation is achieved by dynamically adjusting the weights of data elements. The weight values are adaptively adjusted according to the objective function on the local dataset, that is, when the client iterates once, the local aggregation weight is updated once, so as to more accurately capture information useful for the local model from the global model. The local aggregation weight is updated as:

[0062]

[0063] In the formula, is the aggregated weight after the t-th training update of the i-th client, ω i,t is the aggregated weight before the t-th training update of the i-th client, ρ is the local learning rate, is the gradient of the loss function L with respect to ω i,t L(*) is the loss function, is the locally updated model gradient after adding noise in the (t - 1)-th training of the i-th client, D i,t,Γ is the dataset of the i-th client in the t-th training, θ t-1 is the global model gradient at the (t - 1)-th time;

[0064] That is, using the Hadamard product, the local model is refined and aggregated element by element. During the aggregation process, each client not only downloads the global model but also adaptively adjusts the local aggregation weight. This means that when the client updates its local model, it will retain those global model parameters that are beneficial to improving the performance of the local model, and at the same time adjust or ignore those less relevant parameters. This method is more refined than traditional model-level or hierarchical aggregation and can more accurately capture the desired information in the global model, solving the problems of model performance degradation and high communication cost.

[0065] In this embodiment, to prevent overfitting of the model, a clipping threshold is set for the aggregated weights. The clipping threshold is achieved through a pruning method for element-wise weights, controlling the weight values within the range [0, 1], thereby avoiding excessive influence on the local models during aggregation, maintaining the stability of the model, and ensuring that the model training objectives and weight learning directions are consistent at different stages. Specifically, the clipping function can be expressed as the formula:

[0066]

[0067] In the formula, σ(ω) is the aggregated weight after clipping, and ω is the aggregated weight before clipping of client i in the t-th training;

[0068] Step 4: Obtain the local initialization model of the i-th client according to the aggregated weight after clipping:

[0069]

[0070] In the formula, is the local initialization update model gradient of the i-th client in the t-th training, and ⊙ is the Hadamard product (element-wise product);

[0071] Step 5: The client selects δ of the data from the local dataset for training to obtain the local model gradient:

[0072]

[0073] In the formula, θ i,t is the local model update gradient of the i-th client before adding noise in the t-th training, μ is the global learning rate, is the gradient of the loss function L with respect to the parameter ;

[0074] Step 6: Perform a layering operation on the local training model of the client, dividing the first h layers of the local training model into personalized layers, and the remaining layers into frozen layers;

[0075] Step 7: If the layer of the local model of the client is a personalized layer, calculate the gradient clipping threshold and update the local gradient of the i-th client;

[0076] Among them, the gradient clipping threshold satisfies:

[0077]

[0078] In the formula, C i,t is the gradient clipping threshold of the i-th client in the t-th training, C i,t-1 is the gradient clipping threshold of the i-th client in the (t - 1)-th training, τ is the gradient clipping factor, θ i,t(x i ) is the local training model of the i-th client in the t-th training, and θ i,t-1 (x i ) is the local training model of the i-th client in the t-th training, ||·||2 is the gradient l2 norm, and β i,t-1 is the noise amount of the i-th client in the (t - 1)-th training, and |Γ| is the number of clients participating in the training;

[0079] The noise amount is adding Gaussian noise to the aggregated gradient, and as the number of layers gradually decreases, the scale of adding Gaussian noise also changes, so that the training can achieve better accuracy. For the t-th training process, as the number of layers of the CNN gets lower and lower, the noise scale attenuation of the client is shown as follows:

[0080] β i,t = η·β i,t-1 ;

[0081] In the formula, β i,t is the noise amount of the i-th client in the t-th training, η is the attenuation rate, and η ∈ (0, 1).

[0082] In this embodiment, the initial noise amount size is β0 = 3.

[0083] The local gradient of the i-th client is updated as follows:

[0084]

[0085] In the formula, is the local model gradient of the i-th client in the t-th training after adding noise, β t is the noise scale in the t-th training, and C i,t is the gradient clipping threshold;

[0086] If the layer of the local model of the client is a frozen layer, the local gradient of the client remains unchanged.

[0087] Step 3: After the client iteration is completed, the updated local gradient is sent to the server. The server aggregates the local models and sends the obtained global model gradient to all clients for a new round of client local training until a global model that meets the convergence condition;

[0088] Among them, the server aggregating the local models satisfies:

[0089]

[0090] In the formula, θ t represents the global model gradient at the t-th training.

[0091] The convergence condition is to force termination after reaching a predetermined global iteration number or the global model parameter sequence receiving a unique optimal solution, and finally the obtained global model meets the convergence condition.

[0092] The specific process of the secure federated aggregation algorithm designed and developed by the present invention to solve data heterogeneity is expressed as:

[0093] Input: the number of clients N, the proportion γ of local training, the data sampling rate δ, the local learning rate ρ, the global learning rate μ, the global iteration number T, the gradient clipping factor τ, the noise attenuation rate η;

[0094] Output: model parameter θ t ;

[0095] 1: Initialize the global model θ0, where the global model refers to the aggregation result of the local model parameters of all participating clients in federated learning;

[0096] 2: The server distributes the global model θ0 to N clients for iteration;

[0097] 3: When the model is trained with t = 2, 3,... T;

[0098] 4: Randomly select N×γ clients to generate a client subset Γ;

[0099] 5: The server distributes θ1 to each member in Γ;

[0100] 6: For the i-th client belonging to the client subset Γ, the following operations are executed in parallel:

[0101] 7: When t≥2:

[0102] 8: The i-th client updates the local aggregation weight

[0103] 9: Clip the updated local aggregation weight;

[0104] 10: The i-th client obtains the local initialization model;

[0105] 11: The i-th client selects δ of the data from the local dataset for training;

[0106] 12: The i-th client calculates the local model gradient θ i,t ;

[0107] 13: When the local model of the i-th client is trained in the j-th layer of the CNN:

[0108] 14: If the j-th layer is a personalized layer:

[0109] 15: Calculate the gradient clipping threshold C i,t ;

[0110] 16: Update the local gradient parameters;

[0111] 17: Perform noise clipping on the amount of noise;

[0112] 18: If the j-th layer is a frozen layer:

[0113] 19: The gradient parameters remain unchanged;

[0114] 20: When the i-th client completes the t-th round of training;

[0115] 21: The i-th client sends to the server;

[0116] 22: When federated learning completes one iteration:

[0117] 23: The server aggregates the local model θ t .

[0118] In the above algorithm process, first, the server initializes the global model and distributes the global model to N clients (lines 1 - 2); in lines 6 - 12 of the algorithm, it is the dynamic local adaptive aggregation stage. In the first run, the local aggregation weight value is initialized to 1. Therefore, starting from the second iteration, the client trains the local aggregation weight according to the characteristics of the local data. To prevent the aggregation weight of the model from having too much impact on the local model during aggregation, element-wise weight clipping is performed (line 9); then the local initialized model is trained using the local aggregation weight (line 10); the computational overhead is reduced by randomly selecting δ data from the client's dataset to train the local model and calculating the gradient value (lines 11 - 12); to enhance the security of the client data, a hierarchical gradient clipping method is adopted. Set the first h layers in the CNN as the personalized layers. Then, first calculate the clipping threshold of the gradient. After updating the local model parameters using the clipping threshold, adaptive noise scale reduction is performed to obtain the noise scale of the clients participating in the next round of training. Otherwise, the gradient parameters remain unchanged (lines 14 - 19); finally, the client uploads the gradient value with added noise in the personalized layer to the server (line 21), and the server aggregates the uploaded gradient values to update the global model gradient parameters and distributes them to the clients participating in the training at the beginning of the next iteration (line 23).

[0119] To verify the performance of the SFA - DH described in the present invention, the following experiments are carried out:

[0120] All experiments were conducted using the Python programming language and the model was trained based on the PyTorch 2.3.0 framework. The experiments were run on a local server equipped with a 12th Gen Intel(R) Core(TM) i7-12650H 2.30GHz processor and an Intel(R) UHD Graphics graphics card, with the operating system being Windows 11. To ensure the reliability of the results, all experiments were run independently five times and the average value was taken as the final result.

[0121] To comprehensively evaluate the performance of the SFA-DH algorithm on different types of data and tasks and verify its effectiveness and generalization ability, this experiment focused on the image classification task, selected the Cifar10 dataset as the experimental dataset, set the batch size to 20, and the total number of iterations to 150 to ensure that all methods could converge. In this experiment, the Dirichlet distribution was used to set the Non-IID scenario, and the heterogeneity was set to 0.1. In this experiment, the CNN was set to 4 layers, and the first h = 2 layers were set as the personalized layers.

[0122] In this experiment, the number of clients was 100, the proportion of local training was 20%, the data sampling rate was 80%, the local learning rate was 0.005, the global learning rate was 1, the global number of iterations was 150, the value range of the gradient clipping factor τ was 1 < τ < 10, and the value range of the noise decay rate η was 0 < η < 1.

[0123] Four baseline algorithms were selected for comparison in this experiment, including: Baseline Algorithm 1 - the basic federated learning algorithm FedAvg, Baseline Algorithm 2 - the adaptive personalized federated learning APFL algorithm, Baseline Algorithm 3 - the differential privacy algorithm DP-FL using a fixed level of noise, and Baseline Algorithm 4 - the adaptive gradient clipping algorithm ADP-FL.

[0124] 1. Analyze the effectiveness of the algorithm:

[0125] To verify the effectiveness of the algorithm of the present invention, it was compared with Baseline Algorithm 1 and Baseline Algorithm 2, and the experimental results are as Figure 1As shown, when facing the challenge of non-independent and identically distributed (Non-IID) data distribution, due to its inherent limitations, Baseline Algorithm 1 may not be able to fully capture the unique data characteristics of each client, resulting in poor performance of the global model on some clients, thereby increasing the loss value of the model. In contrast, Baseline Algorithm 2, by introducing a personalized model learning mechanism, enables the model update of each client to be not only affected by the global model but also incorporated with the adjustment of the local model, thus more precisely adapting to their respective data distributions, but it ignores the protection of user privacy data. The algorithm described in the present invention not only trains a personalized local model according to the unique data characteristics of the client but also adds noise to the sensitive data of the user to prevent it from being attacked. As can be clearly seen from the figure, the algorithm described in the present invention has achieved an accuracy rate of 72.08% on the Cifar10 dataset, significantly superior to the FedAvg and APFL algorithms. The SFA-DH algorithm can effectively enable the local model to capture important information element by element from the global model and update the local model, thereby well solving the most intractable data heterogeneity problem in federated learning.

[0126] Meanwhile, in terms of the loss value, the average loss value of the algorithm described in the present invention is lower than that of Baseline Algorithm 1 and Baseline Algorithm 2. As Figure 2 shown, it is respectively 0.327795 and 0.097253 lower in the Cifar10 dataset. In addition, the algorithm described in the present invention already showed signs of convergence when the experiment reached 62 iterations, while Baseline Algorithm 1 and Baseline Algorithm 2 required more rounds to converge, indicating that the SFA-DH algorithm also outperforms these baseline algorithms in terms of convergence speed.

[0127] 2. Analysis of algorithm privacy:

[0128] To verify the effectiveness of the hierarchical gradient adaptive pruning of the training model in the SFA-DH algorithm, SFA-DH was compared with Baseline Algorithm 3 and Baseline Algorithm 4. In the SFA-DH algorithm, SFA-DH adopts a novel hierarchical gradient pruning strategy for the personalized layer of the model, aiming to mitigate the potential impact of the noise introduced by differential privacy on the model performance. While protecting privacy, through refined gradient pruning optimization, it effectively improves the training efficiency and accuracy of the model. Baseline Algorithm 4 dynamically adjusts the noise level by implementing adaptive gradient pruning for each client in different training rounds to reduce the negative impact of differential privacy on the model performance and successfully improves the accuracy of the model. However, in the face of an extremely non-independent and identically distributed (Non-IID) data scenario, the performance of Baseline Algorithm 4 in training is not satisfactory. In contrast, Baseline Algorithm 3 achieves differential privacy by adding a fixed level of noise to the client model parameters, but this method may lead to the accumulation of noise during the training process, thereby affecting the convergence speed and generalization ability of the model.

[0129] From Figure 3 It can be clearly seen that the accuracy of the SFA-DH algorithm is always better than that of baseline algorithm 3 and baseline algorithm 4 on the Cifar10 dataset. In addition, the SFA-DH algorithm shows signs of convergence after 38 iterations on the Cifar10 dataset. It can be concluded that the convergence speed of the SFA-DH algorithm is better than that of baseline algorithm 3 and baseline algorithm 4, and it also demonstrates strong privacy protection capabilities.

[0130] A secure federated aggregation algorithm for solving data heterogeneity designed and developed by the present invention. First, in the dynamic local aggregation stage, users train aggregation weights in each iteration. These weights are responsible for helping the client learn useful information from the global model and initializing the local model. Second, considering that in a convolutional neural network, the layers closer to the output layer contain more user privacy data, differential privacy technology is used to strengthen the protection of these sensitive data. To solve the problem of leakage of user privacy data, an adaptive hierarchical gradient clipping method is adopted to reduce the negative impact of noise on the model performance. By comparing with the FedAvg and APFL algorithms, SFA-DH has achieved significant advantages in performance. At the same time, SFA-DH is also superior to the existing DP-FL and ADP-FL algorithms in terms of privacy protection.

[0131] Although the embodiments of the present invention have been disclosed as above, it is not limited to the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to the specific details and the embodiments shown and described here.

Claims

1. A secure federated aggregation algorithm for solving data heterogeneity, characterized in that, It includes the following steps: Step 1: The server initializes the global model and distributes it to N clients. The N clients independently use their own local data to perform one round of iteration on the global model, obtain the model gradients after the first round of update, and upload them to the server; Step 2: The server distributes the model gradients after the first round of update to a subset of clients. The subset of clients independently use their own local data to perform iterative training on the model, and obtain the updated local models after the iterative training ends; Step 3: After the clients complete the iteration, they send the updated local gradients to the server. The server aggregates the local models and sends the obtained global model gradients to all clients for a new round of client local training until a global model that meets the convergence condition is obtained.

2. The secure federated aggregation algorithm for solving data heterogeneity according to claim 1, wherein The specific content of Step 2 includes: Step 1: Randomly select γ·N clients to generate a subset of clients Γ; Among them, γ is the proportion of local training; Step 2: The server distributes the model gradients θ1 after the first round of update to each member in the subset of clients Γ; Step 3: Each member in the subset of clients Γ performs iterative training. And when t≥2 for the client, during the local aggregation process of the client, aggregation is achieved by dynamically adjusting the weights of data elements; Step 4: Obtain the local initialization model of the i-th client according to the local aggregation weights; Step 5: The client selects δ of the data from the local dataset for training to obtain local model gradients; Step 6: Perform a layering operation on the local training model of the client, and divide the first h layers of the local training model into personalized layers, and the remaining layers are frozen layers; Step 7: If the layer of the local model of the client is a personalized layer, update the local gradient of the i-th client; If the layer of the local model of the client is a frozen layer, keep the local gradient of the client unchanged.

3. The secure federated aggregation algorithm for solving data heterogeneity according to claim 2, wherein The local aggregation weights in Step 3 satisfy: Wherein, is the aggregated weight after the t-th training update of the i-th client, ω i,t is the aggregated weight of the i-th client before the t-th training update, ρ is the local learning rate, is the gradient of the loss function L with respect to ω i,t , L(*) is the loss function, is the locally updated model gradient after noise addition by the i-th client in the (t - 1)-th training, D i,t,Γ is the dataset of the i-th client in the t-th training, θ t-1 is the global model gradient at the (t - 1)-th time.

4. The secure federated aggregation algorithm for solving data heterogeneity according to claim 3, characterized in that, Step 3 also includes clipping the local aggregation weights; In the formula, σ(ω) is the aggregated weight after clipping, and ω is the aggregated weight before clipping of client i in the t-th training.

5. The secure federated aggregation algorithm for solving data heterogeneity according to claim 4, wherein The local initialization model is: where is the local initialized updated model gradient of the i-th client in the t-th training, and ⊙ is the Hadamard product (element-wise product).

6. The secure federated aggregation algorithm for solving data heterogeneity according to claim 5, wherein The local model gradient is: where θ i,t is the local model update gradient of the i-th client before adding noise in the t-th training, μ is the global learning rate, is the gradient of the loss function L with respect to the parameter .

7. The secure federated aggregation algorithm for solving data heterogeneity according to claim 6, characterized in that, The gradient clipping threshold satisfies: where C i,t is the gradient clipping threshold of the i-th client in the t-th training, C i,t-1 is the gradient clipping threshold of the i-th client in the (t - 1)-th training, τ is the gradient clipping factor, θ i,t (x i ) is the local training model of the i-th client in the t-th training, θ i,t-1 (x i ) is the local training model of the i-th client in the t-th training, ||·||2 is the gradient l2 norm, β i,t-1 is the amount of noise of the i-th client in the (t - 1)-th training, and |Γ| is the number of clients participating in the training.

8. The secure federated aggregation algorithm for solving data heterogeneity according to claim 7, wherein, The noise amount satisfies: β i,t = η·β i,t-1 ; where β i,t is the noise volume of the i-th client in the t-th training, η is the attenuation rate, and η ∈ (0, 1).

9. The secure federated aggregation algorithm for solving data heterogeneity according to claim 8, wherein The update of the local gradient of the i-th client is: In the formula, is the local model gradient of the $i$-th client after adding noise in the $t$-th training, and $\beta$ t is the noise scale in the $t$-th training, and $C$ i,t is the gradient clipping threshold.

10. The secure federated aggregation algorithm for solving data heterogeneity according to claim 9, characterized in that, The server aggregates the local model as: where θ t represents the global model gradient at the t-th training.