Differential privacy federated learning method for minimizing noise mechanism and sharpness perception
By introducing adaptive noise mechanism and dynamic cropping threshold in differential privacy federated learning, combined with ASAM optimizer, the impact of noise on model performance is solved, and the effect of improving model performance is achieved while protecting privacy.
Patent Information
- Application Number
- CN202510366632.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-17
AI Technical Summary
The introduction of noise reduces the convergence speed and performance of the model in differential privacy federated learning, and how to maximize model performance while protecting client privacy is a challenge.
A differential privacy federated learning method with minimal sharpness awareness is adopted, through the introduction of an adaptive noise mechanism on the server side and dynamic adjustment of cropping thresholds, combined with the ASAM optimizer to optimize the local model on the client side to ensure a balance of privacy protection and model performance.
It effectively enhances privacy protection, and ensures the satisfaction of differential privacy by dynamically adjusting the noise intensity and cropping threshold, while improving the convergence speed and performance of the model through the ASAM optimizer.
Smart Images

Figure CN120163266A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of federated learning and privacy protection, and specifically to a differential privacy federated learning method based on a noise mechanism and sharpness-aware minimization. Background Art
[0002] With the continuous improvement of computer computing power, machine learning, as an analysis technology capable of processing massive data, has been widely applied in various fields of human society. However, traditional machine learning algorithms use large datasets stored in a central server to train models, usually requiring uploading the source data to a high-computing-power cloud server for centralized training. This approach brings the problem of uncontrollable data flow and may even lead to the risk of sensitive data leakage. Firstly, the security of user data cannot be effectively guaranteed, and there is a potential risk of privacy data leakage. Secondly, due to the existence of network security isolation and industry privacy issues, data barriers between different industries and departments result in "data islands" of data from all parties, and thus secure sharing cannot be achieved. In addition, the performance of machine learning models obtained by each party independently training its own data often cannot reach the global optimization;
[0003] To solve the above problems, Google proposed the federated learning technology, which can train a shared model between multiple terminal devices and a parameter server without transmitting local data to the central server. In this way, federated learning effectively guarantees the privacy security of users participating in the training, transfers the data storage and model training stages of machine learning to local users, and only exchanges model updates with the central server. This method has great potential in solving data privacy and distributed data processing. However, despite many advantages of federated learning, it also faces some challenges and limitations, among which the privacy protection problem is particularly prominent;
[0004] With the continuous improvement of people's awareness of the security of personal information data, the privacy protection problem has increasingly become the focus of people's attention, especially for big data applications and distributed machine learning systems. In this context, federated learning, as a forward-looking solution, is gradually attracting wide attention. A prominent advantage of federated learning is that it decentralizes the model training process to local devices where the data is located without the need to transmit the data to the central server. Specifically, federated learning allows multiple participants to jointly train a machine learning model without sharing the original data, thereby protecting the data of clients from being stolen by hidden adversaries. However, although federated learning avoids directly exposing data to third parties and has a natural protection effect on data privacy, there are still many potential risks of privacy leakage in practical applications;
[0005] Regarding the problem of parameter privacy protection, common techniques include homomorphic encryption, secure multi-party computation, and differential privacy. Under resource constraints, the privacy protection mechanism based on differential privacy has become the main choice for federated learning. This mechanism effectively protects client data privacy by adding noise that conforms to the definition of differential privacy to gradient information or model updates, and the communication and computational costs are relatively low. However, the introduction of noise inevitably reduces the convergence speed and performance of the model. Therefore, in order to maximize the model performance while protecting client privacy, methods to reduce the impact of noise on the model need to be explored, which may involve more refined noise addition strategies, optimization algorithms, or other innovative methods to minimize performance loss while protecting privacy;
[0006] In addition, the combination of federated learning and differential privacy technology has important application value in multiple fields, especially in scenarios where sensitive data needs to be protected. For example, in the field of medical and health, by training models locally in different hospitals and applying differential privacy, data such as medical images and disease prediction can be shared to avoid patient privacy leakage. In the financial field, banks and financial institutions can train credit scoring, anti-fraud models, etc. locally to protect users' transaction history and credit information. In the field of intelligent transportation and autonomous driving, by implementing local training in vehicles or transportation facilities and combining differential privacy protection, the intelligence of traffic scheduling and autonomous driving systems can be improved while ensuring the privacy of individual travel data. In smartphones and Internet of Things devices, users' behavioral data (such as health monitoring, application usage, environmental data, etc.) can be processed locally on the device side to ensure personalized services while protecting privacy. In addition, in the field of e-commerce and advertising recommendation, the combination of federated learning and differential privacy technology can optimize the recommendation system and improve the accuracy of personalized recommendations without revealing users' browsing and purchase behaviors. Through the application of these technologies, the differential privacy federated learning privacy protection scheme provides solutions for various industries that can both protect privacy and improve performance. Summary of the Invention
[0007] The purpose of the present invention is to provide a differential privacy federated learning method with a noise mechanism and sharpness-aware minimization to solve the problems proposed in the above background technology that the introduction of noise inevitably reduces the convergence speed and performance of the model, and how to maximize the model performance while protecting client privacy and explore ways to reduce the impact of noise on the model.
[0008] To achieve the above purpose, the present invention provides the following technical solution: A differential privacy federated learning method with a noise mechanism and sharpness-aware minimization, including the following steps:
[0009] S1: Start federated learning. The server initializes the global model parameters. The global model includes the linear transformation layer of the model as the initial layer to alleviate the heterogeneity of client data. Subsequently, the server randomly selects a fixed number of partial clients and sends the initial model parameters to these clients, and the clients train the local model based on this.
[0010] S2: The client receives the global model parameters and uses them as the initial parameters of the local model for training. In the local training stage, the client updates the local model through the linear transformation layer and the ASAM optimizer of the model. First, the client performs a linear transformation on the data through the linear transformation layer of the model. Then, the client calculates the local gradient, and then finds the maximum point of the loss function through gradient ascent. This point is within the adjustable region around the current model parameters. After determining the maximum point, the ASAM optimizer will perform gradient descent on the flat loss surface within this region to find the optimal solution. Finally, the client uploads the updated local model difference to the server for global aggregation.
[0011] S3: After the server receives the local model updates uploaded by the clients, it calculates the gradient norm of each update. Subsequently, the server adaptively adjusts the clipping threshold according to the gradient distribution of the clients by selecting a quantile decay method (linear decay, exponential decay, or cosine decay).
[0012] S4: To ensure that the global model meets the differential privacy requirements, the server first clips the local model parameters of each client, then combines the clipped gradient updates with noise, and finally the server obtains the model parameters after noise perturbation.
[0013] S5: The server aggregates the local models using the FedAVG aggregation rule to obtain a new global model and sends it to the clients.
[0014] S6: Repeat steps S1 to S5 until the given maximum number of communication rounds is reached.
[0015] Preferably, the specific steps of S2 are as follows:
[0016] S2.1: Linear transformation stage of data: When the data of each client is input into the local model, the data needs to be linearly transformed first. The transformation rule is as follows:
[0017] T k (x) = αx + β
[0018] Where: α is the scaling parameter used to adjust the amplitude of the input data and learn the best ratio in the global distribution.
[0019] β is the offset parameter used to adjust the center of the input data to make it closer to the mean of the global distribution.
[0020] x is the input data of client k, T k (x) is the transformed data;
[0021] S2.2. Gradient Ascent Phase: Each client first finds a point that maximizes the loss through gradient ascent. This point lies within an adjustable region centered on the current model parameters. The formula is as follows:
[0022]
[0023] where, T w is a normalization operator used to ensure the consistency of sharpness after parameter scaling. F k (w t,l ) is the local loss function of the k-th client. ρ is a predefined constant that controls the radius of the perturbation and determines the range of the region selected by ASAM during gradient ascent. ||.||2 represents the L2-norm;
[0024] S2.3. Gradient Descent Phase: After finding the maximization point, ASAM performs gradient descent on the flat loss surface within this region to find the optimal solution. The formula is as follows:
[0025]
[0026] where, w t,l (k) is the local model parameter of client k. F k (w t,l +∈(w t,l )); ξ) is the local model parameter of client k after perturbation. η is the learning rate, and ξ is the local dataset of client k.
[0027] Preferably, the specific steps of S3 are as follows:
[0028] S3.1. Calculate the update norm of the model: After the local client finishes training, it uploads the local model parameters for the server to calculate the update norm of the model. The formula is as follows:
[0029]
[0030] where, is the local model parameter of client k, and w t is the global model parameter at the t-th round;
[0031] S3.2. Select the quantile decay clipping strategy: The core of dynamic quantile clipping is the dynamic change rule of the quantile Q t The dynamic change rule of the quantile Q tThe adjustment strategy directly affects the dynamic range of the clipping threshold. In this paper, three quantile decay strategies are proposed, namely linear decay, exponential decay, and cosine decay, to flexibly control the intensity of gradient clipping;
[0032] (1) Linear decay is a simple and efficient strategy where the quantile Q t decreases with a fixed step size as the training round t progresses, decaying from the initial quantile Q0 to the final quantile Q T . The formula is as follows:
[0033]
[0034] where Q0 is the initial quantile, usually a relatively large value, and Q T is the final quantile, usually a relatively small value (0 < Q0 < Q T ≤ 1). Linear decay reduces the quantile by a fixed proportion in each round and is suitable for tasks that require a smooth adjustment of the clipping threshold;
[0035] (2) The exponential decay strategy uses an exponential function to adjust the quantile. It is characterized by a faster decay in the early stage and a tendency to level off in the later stage. The formula is as follows:
[0036] Q t = Q T +(Q0 - Q T ).e -αt , t ∈ {0, 1,..., T - 1}
[0037] where α is a parameter that controls the decay rate. Exponential decay rapidly degrades the quantile in the early stage and is suitable for scenarios where the initial gradients are large and the influence of outliers is significant in appropriate tasks;
[0038] (3) The cosine decay strategy smoothly adjusts the quantile through a cosine function, with slower changes in the initial and later stages and faster changes in the middle stage. The formula is as follows:
[0039]
[0040] Cosine decay is suitable for tasks that require a smooth transition, especially when more stable gradient clipping is needed in the later stage of training;
[0041] S3.3. Calculate the clipping threshold: The server calculates the quantile for the set of update norms of all selected clients. The formula for the clipping threshold is as follows:
[0042]
[0043] where, represents the quantile function, calculating the value corresponding to the quantile Q t in the set, and K tDenote the set of clients in the $t$-th round, $Q$ t is the quantile of the current round, and its dynamic change is determined by the quantile decay strategy defined later.
[0044] Preferably, the specific steps of S4 are as follows:
[0045] S4.1. To ensure that the global model meets the differential privacy requirements, the server first clips the local model parameters of each client, restricting the updates of all clients within the threshold $C$ t range. Specifically, for each client $k$, the clipping rule for its gradient update is defined as follows:
[0046]
[0047] S4.2. After clipping, DP noise needs to be added to the clipped parameters to ensure privacy. The formula is as follows:
[0048]
[0049] where $w$ t represents the model parameters before adding noise, and $w$ t, represents the model parameters after adding noise. represents Gaussian noise;
[0050] where $\sigma$ satisfies:
[0051]
[0052] where $\Delta f$ represents the sensitivity, which is used to measure the sensitivity of the algorithm to the change of a single sample in the dataset;
[0053] $\varepsilon$ represents the privacy budget, which measures the level of using noise;
[0054] $\delta$ represents the privacy failure probability.
[0055] Preferably, the specific steps of S5 are as follows:
[0056] The server uses the FedAVG aggregation rule to aggregate the local models to obtain a new global model. The formula is as follows:
[0057]
[0058] where $w$ t is the global model parameter of this round, and $w$ t+1 is the global model parameter of the next round, and $K$ s is the number of participating clients.
[0059] Compared with the prior art, the beneficial effects of the present invention are as follows: The differential privacy federated learning method with noise mechanism and sharpness perception minimization:
[0060] 1. Enhanced privacy protection: By introducing an adaptive noise mechanism and combining it with the adaptive adjustment of the clipping threshold, the present invention ensures that the requirements of differential privacy are effectively met during the federated learning process. In the initial stage of training, a larger clipping threshold is used to introduce stronger noise, protecting the privacy of user data. As the training progresses, the clipping threshold gradually decreases and the noise gradually reduces, ensuring the balance between privacy protection and model accuracy.
[0061] 2. Improved model performance: By combining a linear transformation layer and an ASAM optimizer, the present invention can effectively alleviate the problem of client data heterogeneity and improve the generalization ability of the local model. The ASAM optimizer enhances the stability of local training by flattening the loss surface, avoiding overfitting, and thus improving the performance of the global model.
[0062] 3. Flexible noise mechanism: The present invention proposes three different attenuation methods (linear attenuation, exponential attenuation, or cosine attenuation), providing a flexible solution that can be adjusted according to the characteristics of specific tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 is a schematic flow chart of the learning method of the present invention;
[0064] Figure 2 is a performance comparison chart of the method of the present invention and three traditional federated learning methods under different degrees of data Non-I.I.D in the MNIST dataset;
[0065] Figure 3 is a performance comparison chart of the method of the present invention and three traditional federated learning methods under different degrees of data Non-I.I.D in the EMNIST dataset;
[0066] Figure 4 is a performance comparison chart of the method of the present invention and three traditional federated learning methods under different degrees of data Non-I.I.D in the Fashion-MNIST dataset. DETAILED DESCRIPTION OF THE INVENTION
[0067] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0068] Please refer to Figures 1-4, the present invention provides a technical solution: a differential privacy federated learning method for noise mechanism and sharpness-aware minimization, aiming to improve the performance of the global model, such as Figure 1 as shown below:
[0069] S1. Global model initialization: At the beginning of training, the server initializes the global model parameters w T . The global model includes the linear transformation layer of the model as the initial layer to alleviate the heterogeneity of client data. Subsequently, the server randomly selects K s clients and sends the initial model parameters w 0 to these clients, and the clients train the local model based on this.
[0070] S2. Client local model training: After receiving the global model parameters w T , the client uses them as the initial parameters of the local model for training. In the local training stage, the client updates the local model through the linear transformation layer of the model and the ASAM optimizer. First, the client linearly transforms the data through the linear transformation layer of the model, and the formula is as follows:
[0071] T k (x) = αx + β
[0072] where: α is the scaling parameter used to adjust the amplitude of the input data and learn the optimal ratio in the global distribution; β is the offset parameter used to adjust the center of the input data to make it closer to the mean of the global distribution.
[0073] After that, the client calculates the local gradient and then finds the maximum point of the loss function through gradient ascent. This point is within the adjustable region around the current model parameters, and the formula is as follows:
[0074]
[0075] where, T w is the normalization operator used to ensure the consistency of sharpness after parameter scaling, F k (w t,l ) is the local loss function of the k-th client, ρ is a predefined constant that controls the radius of the perturbation and determines the range of the region selected by ASAM in gradient ascent, and ||.||2 represents the L2-norm.
[0076] After determining the maximum point, the ASAM optimizer will perform gradient descent on the flat loss surface within this region to find the optimal solution, and the formula is as follows:
[0077]
[0078] where, w t,l(k) are the local model parameters of client k, F k (w t,l +∈(w t,l )); ξ) are the perturbed local model parameters of client k, and η is the learning rate;
[0079] Finally, the client uploads the updated local model difference to the server for global aggregation;
[0080] S3. Dynamically adjust the clipping threshold: After the server receives the local model updates uploaded by the clients, it calculates the gradient norm of each update. The formula is as follows:
[0081]
[0082] Among them, are the local model parameters of client k, and w t is the global model parameter in the t-th round;
[0083] Subsequently, the server selects a quantile decay method (linear decay, exponential decay, or cosine decay) according to the gradient distribution of the clients to dynamically adjust the clipping threshold C t , and the formula is as follows:
[0084]
[0085] Among them, represents the quantile function, which calculates the value corresponding to the quantile Q t in the set, and K t represents the set of clients in the t-th round. Q t is the quantile in the current round, and its dynamic change is determined by the quantile decay strategy defined later;
[0086] S4. Differential privacy guarantee: To ensure that the global model w T meets the differential privacy requirements, the server first clips the local model parameters of each client. The formula is as follows:
[0087]
[0088] Then, it combines the clipped gradient update with noise. Finally, the server obtains the model parameters after noise perturbation. The formula is as follows:
[0089]
[0090] Among them, w t represents the model parameters before adding noise, and w t, represents the model parameters after adding noise. represents Gaussian noise;
[0091] S5. Global Model Update: The server aggregates the local models using the FedAVG aggregation rule to obtain a new global model and sends it to the clients. The formula is as follows:
[0092]
[0093] where w t is the global model parameter for this round, w t+1 is the global model parameter for the next round, and K s is the number of participating clients;
[0094] S6. Repeat S1 to S5 until the given maximum number of communication rounds is reached.
[0095] In the present invention, an adaptive noise mechanism is designed and the ASAM optimizer is introduced to form a collaborative optimization strategy, which is reflected in: the method of the present invention adopts an adaptive noise mechanism on the server side to dynamically adjust the intensity of noise introduction, and combines the ASAM optimizer on the client side to optimize the local model, generating a flatter local model for each client. In this way, it can effectively accelerate the convergence of the global model and improve the communication efficiency, while enhancing the overall performance of federated learning on the premise of ensuring privacy;
[0096] Example:
[0097] To intuitively demonstrate the effectiveness and superiority of the present method, the following experimental results are provided.
[0098] The ANSDPFL method proposed in the present invention is implemented using PyTorch and compared with the following 3 classic FL methods, namely:
[0099] DP-FedAvg (proposed in the paper "Learning differentially private recurrent language models"), DP-FedSAM (proposed in the paper "Make landscape flatter in differentially private federated learning") and Privatefl (proposed in the paper "{PrivateFL}: Accurate, differentially private federated learning via personalized data transformation");
[0100] The superiority of the ANSDPFL proposed in the present invention under non-IID data conditions is demonstrated through three groups of image classification experiments;
[0101] In this embodiment, three widely used public datasets are selected: MNIST, FashionMNIST, and EMNIST. A three-layer 3-layer DNN is used as our model. The method for setting the non-IID data environment in this embodiment is as follows: We randomly select N classes for each client uniformly. Then, we assign the training samples in class c to each client k, where the classes it selects contain c, and pk,c is a random number within the range (0.4, 0.6). Our default data distribution uses N = 2. We also evaluate different ε values, namely [2, 4, 6, 8]. Since ε depends on the number of training rounds, we use different numbers of training rounds for different ε values. In particular, for the considered ε values, we use [60, 80, 100, 150] training rounds for MNIST, [30, 60, 100, 150] training rounds for Fashion-MNIST, and [30, 50, 60, 100] training rounds for EMNIST. Each experiment is run 3 times, and the accuracy of the best average test global model in each experiment is reported.
[0102] Table 1 Performance comparison of the method of the present invention with three traditional federated learning methods under different privacy budgets
[0103]
[0104] The experimental results based on the three public datasets are shown in Table 1. Table 1 lists the experimental results of each method under different privacy budgets in the three groups of experiments. The ANSDPFL-Linear method means that linear decay is used in the decay method, the ANSDPFL-Exp method means that exponential decay is used in the decay method, and the ANSDPFL-Cosine method means that cosine decay is used in the decay method. The experimental results show that the proposed method is significantly better than the baseline method in all cases.
[0105] We also evaluated the impact of the degree of non-independent and identically distributed (Non-I.I.D.) of the client's local training data on the performance of our scheme. For this purpose, we conducted experiments on the MNIST, Fashion-MNIST, and EMNIST datasets and evaluated the performance of the model under the default privacy budget and parameter settings. In these experiments, we assigned different numbers of classes to each client. Specifically, the number of data classes for each client is 2, 4, 6, 8, or 10 classes, where the case where the client data is independent and identically distributed (I.I.D.) corresponds to the setting of 10 classes of data. In this way, we can examine the impact of the degree of data heterogeneity (i.e., Non-I.I.D.) on the model performance. Figure 2 、 Figure 3 、 Figure 4Shows our results, Figure 2 , Figure 3 , Figure 4 respectively represent the experimental results on the MNIST, EMNIST, and Fashion-MNIST datasets. The experimental results show that the method we proposed exhibits the best results in different data settings.
[0106] The above specific embodiments have further elaborated on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and do not limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are included within the protection scope of the present invention.
Claims
1. A differential privacy federated learning method with noise mechanism and sharpness perception minimization, characterized by: The following steps are involved: S1: Federated learning starts. The server initializes the global model parameters. The global model contains the linear transformation layer of the model as the initial layer to alleviate the heterogeneity of client data. Then, the server randomly selects a fixed number of clients and sends the initial model parameters to these clients. The clients train local models based on this. S2: The client receives the global model parameters and uses them as the initial parameters of the local model for training. In the local training phase, the client updates the local model through the linear transformation layer of the model and the ASAM optimizer. The client first performs a linear transformation on the data through the linear transformation layer of the model. After that, the client calculates the local gradient and then finds the maximization point of the loss function through gradient ascent. This point is located in the adjustable area around the current model parameters. After determining the maximization point, the ASAM optimizer performs gradient descent on the flat loss surface in the area to find the optimal solution. Finally, the client uploads the updated local model difference to the server for global aggregation. S3: After receiving the local model update uploaded by the client, the server calculates the gradient norm of each update. Then, the server selects a quantile decay method (linear decay, exponential decay, or cosine decay) according to the client's gradient distribution to adaptively adjust the clipping threshold; S4: To ensure that the global model meets the differential privacy requirements, the server first prunes the local model parameters of each client, then combines the pruned gradient update with the noise, and finally the server obtains the model parameters after noise perturbation; S5: The server aggregates the local model using the FedAVG aggregation rule to obtain a new global model and sends it to the client; S6: Repeat S1 to S5 until a given maximum number of communication rounds is reached.
2. A differential privacy federated learning method with noise mechanism and sharpness perception minimization according to claim 1, characterized in that: The specific steps of S2 are: S2.
1. Data linear transformation stage: When the data of each client is input into the local model, the data must first be linearly transformed. The transformation rules are as follows: T k (x)=αx+β Where: α is a scaling parameter used to adjust the amplitude of the input data and learn the best proportion in the global distribution; β is the offset parameter used to adjust the center of the input data to make it closer to the mean of the global distribution; x is the input data of client k, T k (x) is the transformed data; S2.2, Gradient Ascent Phase: Each client first finds a point where the loss is maximized through gradient ascent. This point is located in the adjustable region centered on the current model parameters. The formula is as follows: Among them, T w is a normalization operator used to ensure that the sharpness remains consistent after parameter scaling, F k (w t,l ) is the local loss function of the kth client, ρ is a predefined constant that controls the radius of the perturbation and determines the range of the region selected by ASAM in gradient ascent, and ||.||2 represents the L2-norm; S2.3, Gradient descent stage: After finding the maximization point, ASAM performs gradient descent on the flat loss surface in this area to find the optimal solution. The formula is as follows: Among them, w t,l (k) is the local model parameter of client k, F k (w t,l +∈(w t,l );ξ) is the local model parameter of client k after disturbance, η is the learning rate, and ξ is the local dataset of client k.
3. The differential privacy federated learning method with noise mechanism and sharpness perception minimization according to claim 1, characterized in that: The specific steps of S3 are: S3.
1. Calculate the update norm of the model: After the local client training is completed, the local model parameters need to be uploaded to the server to calculate the update norm of the model. The formula is as follows: in, is the local model parameter of client k, w t is the global model parameter of the tth round; S3.
2. Select the quantile attenuation pruning strategy: The core of dynamic quantile pruning is the quantile Q t Dynamic change rules, quantile Q t The adjustment strategy directly affects the dynamic range of the clipping threshold. This paper proposes three quantile attenuation strategies, namely linear attenuation, exponential attenuation and cosine attenuation, to flexibly control the intensity of gradient clipping. (1) Linear decay is a simple and efficient strategy. t As the training round t decreases with a fixed step size, it decays from the initial quantile Q0 to the final quantile Q T , the formula is as follows: Among them, Q0 is the initial quantile, usually a larger value, Q T is the final quantile, usually a small value (0 <Q0<Q T ≤1), linear decay reduces the quantile by a fixed proportion in each round, which is suitable for tasks that require smooth adjustment of the clipping threshold; (2) The exponential decay strategy uses an exponential function to adjust the quantiles. Its characteristics are that it decays faster in the early stage and tends to be stable in the later stage. The formula is as follows: Q t =Q T +(Q0-Q T ).e -αt ,t∈{0,1,...,T-1} Among them, α is a parameter that controls the decay speed. Exponential decay quickly degrades the quantile in the early stage, which is suitable for tasks with large initial gradients and significant impact of outliers. (3) The cosine decay strategy uses the cosine function to smoothly adjust the quantiles. The changes are slower in the early and late stages and faster in the middle stages. The formula is as follows: Cosine decay is suitable for tasks that require smooth transitions, especially when more stable gradient clipping is needed in the later stages of training; S3.
3. Calculate the clipping threshold: The server calculates the quantiles of the update norm set of all selected clients. The clipping threshold calculation formula is as follows: Among them, Q Qt (.) represents the function of dividing into and calculating the quantile Q in the set t The corresponding value, K t represents the set of clients in round t, Q t is the quantile of the current round, and its dynamic change is determined by the quantile decay strategy defined later.
4. The differential privacy federated learning method with noise mechanism and sharpness perception minimization according to claim 1, characterized in that: The specific steps of S4 are: S4.
1. To ensure that the global model meets the differential privacy requirements, the server first prunes the local model parameters of each client and limits the updates of all clients to a threshold C. t Specifically, for each client k, the clipping rule for its gradient update is defined as follows: S4.
2. After cropping, DP noise should be added to the cropped parameters to ensure privacy. The formula is as follows: Among them, w t represents the model parameters before adding noise, w t, represents the model parameters after adding noise, represents Gaussian noise; Where σ satisfies: Among them, Δf represents sensitivity, which is used to measure the sensitivity of the algorithm to changes in a single sample in the data set; ε represents the privacy budget measuring the level of noise used; δ represents the privacy failure probability.
5. The differential privacy federated learning method with noise mechanism and sharpness perception minimization according to claim 1, characterized in that: The specific steps of S5 are: The server aggregates the local model using the FedAVG aggregation rule to obtain a new global model. The formula is as follows: Among them, w t is the global model parameter of this round, w t+1 is the global model parameter for the next round, K s is the number of participating clients.
Citation Information
Cited By
Federal learning utility optimization system and method for resisting data heterogeneity
CN121094169A
Self-adaptive federated learning aggregation and privacy protection method based on sharpness perception minimization
CN121786880A