Regularization Constraint Adaptive Adjustment Method for Privacy-Preserving Heterogeneous Decentralized Learning

By adopting a regular constraint adaptive adjustment method in federated learning, combining stochastic gradient push and dynamic update of KL divergence, the challenges of data privacy protection and computing efficiency in federated learning are solved, and efficient optimization of non-independent and homogeneous data is achieved and model performance is improved.

CN119808896BActive Publication Date: 2025-05-23QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510296997.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-05-23
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

In federated learning, data privacy protection and computing efficiency face many challenges, especially in the negative impact of data heterogeneity and differential privacy noise on model performance.

Method used

The regular constraint adaptive adjustment method for privacy-saving heterogeneous decentralized learning is adopted, and the regularized parameters are dynamically updated and adjusted through the stochastic gradient push method, regularized optimization and adaptive parameter adjustment mechanism, combined with KL divergence, to achieve efficient optimization of non-independent and homogeneous data in the decentralized topology.

Benefits of technology

While protecting data privacy, this method reduces the negative impact of data heterogeneity and differential privacy noise on model performance, improves the generalization ability of the model, and is suitable for time-varying directed communication topology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119808896B_ABST
    Figure CN119808896B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of privacy protection, and more specifically, to a regularization constraint adaptive adjustment method for privacy-preserving heterogeneous decentralized learning. The method comprises: dividing the local model of each client into a shared model and a retained model, and performing regularization constraints on the shared model; the client uses the shared model and the local retained model after the previous round of aggregation to perform gradient descent updates based on the local data set; dynamically updates and adjusts the regularization parameters through KL divergence; differential privacy protection is performed on the shared model, and then the noisy shared model is broadcast to neighboring clients; the neighboring clients of client i receive the noisy shared model and aggregate it to obtain the local model of the next iteration round. While protecting data privacy, the present invention reduces the negative impact of data heterogeneity and differential privacy noise on model performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of privacy protection, and more specifically, to a regularization constraint adaptive adjustment method for privacy-preserving heterogeneous decentralized learning. Background Art

[0002] Federated Learning (FL) is an emerging distributed machine learning framework that aims to leverage multiple decentralized clients to collaboratively train global models without directly sharing the original data. By exchanging only local gradients or model updates, federated learning has shown great potential in data privacy protection, and has therefore been widely used in privacy-sensitive fields such as finance, healthcare, and the Internet of Things. However, in the actual application of federated learning, data privacy protection and computational efficiency still face many challenges. The first is the issue of data privacy. Although the client does not directly share data, sensitive information about the original data may be leaked in the exchanged gradients or model updates. However, in the actual application of federated learning, data privacy protection, communication efficiency, and model performance still face many challenges, and in-depth research and improvement are urgently needed.

[0003] First, in terms of data privacy, although federated learning avoids directly sharing original data by completing data processing locally on the client, sensitive information about the original data may still be leaked in the exchanged gradients or model updates. In recent years, differential privacy (DP) has been proposed as an effective privacy protection mechanism that limits potential information leakage by adding noise to gradients or model updates, making it difficult for attackers to infer the original data through the communication content. Nevertheless, how to strike a balance between privacy protection and model performance remains an important research topic.

[0004] Secondly, in terms of communication architecture, most existing federated learning frameworks rely on a centralized communication architecture, that is, all clients communicate and aggregate parameters through a central server. However, this architecture may be limited by high communication costs, single point failure risks, and system bottlenecks in actual deployment, which reduces its applicability in large-scale distributed systems. In contrast, decentralized communication architecture (Decentralized Federated Learning, DFL) has received widespread attention in recent years as a more robust and flexible solution. In a decentralized architecture, clients communicate and update models through adjacent nodes without relying on a central server, which significantly improves the fault tolerance and scalability of the system. However, this architecture still faces major challenges in terms of communication efficiency and global model consistency.

[0005] Finally, the data heterogeneity (Non-IID) problem is particularly prominent in federated learning. In actual scenarios, the client data is usually not independent and identically distributed, which may lead to significant differences in the learning process of the local model of each client, bringing model bias and performance inconsistency problems to federated learning.

[0006] Chinese patent document CN115510482A discloses a method for protecting the privacy of crowd-sensing data that integrates adaptive differential privacy and federated learning. The invention adopts a federated learning method, which does not require sharing of perception data. The data submitted by the perception user is stored and trained locally to form a local model. Secondly, differential privacy is used to add Gaussian noise to disturb the local model uploaded by the perception user to prevent information leakage. The processed local model is uploaded to the perception platform for aggregation. Finally, an adaptive differential privacy mechanism is used to adjust the trimming value in real time to reduce the errors caused by noise and parameter trimming and reduce the communication cost. However, this method still requires the participation of the central server and cannot be applied to directed and time-varying communication topologies. In addition, the real-time adjustment of the trimming value is used to constrain the size of the local model, which inevitably leads to trimming deviations.

[0007] In view of this, the present invention designs a regularization constraint adaptive adjustment method for privacy-preserving heterogeneous decentralized learning to achieve efficient optimization of non-independent and identically distributed (Non-IID) data in decentralized topology. Summary of the invention

[0008] The present invention aims to overcome at least one defect of the above-mentioned prior art and provide a regularization constraint adaptive adjustment method for privacy-preserving heterogeneous decentralized learning, aiming to balance data privacy protection and model performance, while protecting data privacy, reducing the negative impact of data heterogeneity and differential privacy noise on model performance. Specifically, this method achieves efficient optimization of non-independent and identically distributed (Non-IID) data in decentralized topology by combining stochastic gradient push method, regularization optimization and adaptive parameter adjustment mechanism.

[0009] The detailed technical scheme of the present invention is as follows:

[0010] A regularization constraint adaptive adjustment method for privacy-preserving heterogeneous decentralized learning, the method comprising:

[0011] S1. Each client Local Model Divide into shared models and retain the model , for the shared model Regularization constraints are applied to limit the update amplitude of model parameters;

[0012] S2, Client Use the shared model after the previous round of aggregation and locally retained models , based on local datasets Perform gradient descent updates;

[0013] S3, regularization parameters through KL divergence Perform dynamic update and adjustment;

[0014] S4. To protect the shared model Privacy, for shared models Perform differential privacy protection, and then broadcast the noisy shared model to neighbor clients to ensure privacy protection while completing decentralized communication;

[0015] S5. The neighboring clients of client i receive the noisy shared model and aggregate it to obtain the local model for the next iteration round.

[0016] Preferably, according to the present invention, step S1 is specifically as follows:

[0017] S11. Local Model The local model is processed by the deep neural network layer by layer. Divided into shared models and retain the model , where the shared model Used to exchange between neighbor clients and retain the model Keep it locally to maintain personalization features;

[0018] S12. Regularize the shared model to limit the model bias and reformulate the client's loss function as:

[0019] (1)

[0020] (2)

[0021] In formula (1) and formula (2), is the cross entropy loss function, is the regularization parameter, is the clipping threshold, and Represent the shared models of this round and the previous round respectively, represents the regularization term, Indicates The debiasing parameter in the round iteration, when t=0, = .

[0022] Preferably, according to the present invention, the rule of gradient descent update is as follows:

[0023] (3)

[0024] In formula (3), is the learning rate, is the loss function based on local data, represents the debiasing parameter, Represents a local dataset.

[0025] According to the preferred embodiment of the present invention, the KL divergence is used to measure the difference between the current shared model and the previous round of shared model distribution. The dynamic update adjustments are as follows:

[0026] (4)

[0027] In formula (4), is the step size parameter, and denote the shared models of round t and round t-1 respectively, represents the KL divergence, represents the updated regularization parameter of the t+1th round, represents the regularization parameter of this round.

[0028] Preferably, the differential privacy protection of the shared model refers to: achieving differential privacy by adding noise to the shared model, wherein the noise obeys a Gaussian distribution, as follows:

[0029] (5)

[0030] (6)

[0031] In formula (5) and formula (6), Represents the shared model parameters that need to be shared with neighboring clients. represents the clipping threshold, The expected value is 0 and the variance is Gaussian noise of , which meets the privacy budget of differential privacy Require, Representatives Shared model parameters after noise addition.

[0032] Preferably, the step S5 is as follows: the neighboring clients of client i receive the noisy shared model and aggregate it to obtain the local model of the next iteration round, specifically:

[0033] (7)

[0034] In formula (7), represents the local model of the t+1th round obtained by aggregation, represents the local model of round t, represents the weight from client j to client i in round t, Indicates the neighbor client updated in round t Scalar push weight of Indicates the neighbor client updated in the t+1th round Scalar push weight of ; From the client In the The debiasing parameter in the round iteration.

[0035] Compared with the prior art, the present invention has the following beneficial effects:

[0036] (1) The present invention adopts a random gradient push method, which eliminates the dependence on the central node, is suitable for time-varying directed communication topologies, and provides privacy protection for each client.

[0037] (2) The present invention designs an adaptive parameter optimization mechanism, which dynamically updates and adjusts the regularization parameter through KL divergence to improve the generalization ability of the model and provide convergence guarantee for non-convex optimization problems. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is a flow chart of the method of the present invention.

[0039] Figure 2 It is a result diagram of an experiment conducted on the SVHN dataset in an embodiment of the present invention. DETAILED DESCRIPTION

[0040] The present disclosure is further described below in conjunction with the accompanying drawings and embodiments.

[0041] Embodiment 1,

[0042] Ginseng Figure 1 This embodiment provides a regular constraint adaptive adjustment method for privacy-preserving heterogeneous decentralized learning, which is applied to a federated learning system. The federated learning system includes multiple clients, and each client has a local data set for image classification and recognition tasks; the method includes:

[0043] S1. Each client Local Model Divide into shared models and retain the model , for the shared model Regularization constraints are applied to limit the update amplitude of model parameters;

[0044] First, initialize the initial model parameters of each client , debiasing parameter .

[0045] In the present invention, the division body of the local model is a deep neural network, and the model parameters are deep neural network parameters; preferably, a convolutional neural network model CNN is used as the deep neural network of the local model.

[0046] S11. Local Model The deep neural network is layered and divided into a shared model and a reserved model. In this embodiment, the last layer in the model is set as the reserved model. , and the rest are set as shared models , where the shared model Used to exchange between neighbor clients and retain the model Keep it locally to maintain personalization.

[0047] Local Model Divided into shared models and retain the model , upload only , so we only need to Add noise, no need to Adding noise reduces the introduction of noise.

[0048] Local Model Divided into shared models and retain the model , upload only , can alleviate the impact of data heterogeneity, because Has local personalized knowledge, so retain ,make Preserve local, personalized knowledge.

[0049] S12. Regularize the shared model to limit the model bias. The client's loss function can be re-expressed as:

[0050]

[0051] in, is the cross entropy loss function, is the regularization parameter, is the clipping threshold, and Represent the shared models of this round and the previous round respectively, represents the regularization term, Indicates The debiasing parameter in the round iteration, when t=0, = .

[0052] S2, Client Use the shared model after the previous round of aggregation and locally retained models , based on local datasets Perform gradient descent update, the update rules are as follows:

[0053]

[0054] in is the learning rate, is the loss function based on local data, represents the debiasing parameter, Represents a local dataset.

[0055] S3, regularization parameters through KL divergence Make dynamic adjustments;

[0056] KL divergence is used to measure the difference between the current shared model and the previous round of shared model distribution. The KL divergence is used to regularize the parameter The dynamic update adjustments are as follows:

[0057]

[0058] in, is the step size parameter, and denote the shared models of round t and round t-1 respectively, represents the KL divergence, represents the updated regularization parameter of the t+1th round, represents the regularization parameter of this round.

[0059] S4. To protect the shared model To improve privacy, differential privacy is protected for the shared model. Specifically, differential privacy is achieved by adding noise to the shared model, and the noise follows a Gaussian distribution:

[0060]

[0061]

[0062] in, Represents the shared model parameters that need to be shared with neighboring clients. represents the clipping threshold, The expected value is 0 and the variance is Gaussian noise of , which meets the privacy budget of differential privacy Require, Representatives Shared model parameters after adding noise;

[0063] Then, the noisy shared model is broadcast to neighboring clients to ensure privacy protection while completing decentralized communication.

[0064] S5. The neighboring clients of client i receive the noisy shared model and aggregate it to obtain the local model of the next iteration round, which is:

[0065]

[0066] in, represents the local model of the t+1th round obtained by aggregation, represents the local model of round t, represents the weight from client j to client i in round t, Indicates the neighbor client updated in round t Scalar push weight of Indicates the neighbor client updated in the t+1th round Scalar push weight of ; From the client In the The debiasing parameter in the round iteration.

[0067] The convergence effect under differential privacy in decentralized federated learning is as follows:

[0068] in , represents the initial model of each client, L represents the smoothness coefficient of the loss function, and They represent the deviation between the gradient of a single client and the global gradient, and the variance generated by the gradient distribution of participating clients.

[0069] Experimental example

[0070] This scheme conducts comparative experiments on the SVHN Street View House Number Dataset. The SVHN dataset is a real-world image dataset for number recognition. It is derived from house numbers in Google Street View and contains more than 600,000 32×32 pixel color images. It has high complexity and diversity and is often used as a benchmark for deep learning and computer vision tasks. The comparative experimental results are shown in Figure 2. Figure 2As shown in the figure, the scheme of the present invention improves the accuracy by 3% to 6% compared with the benchmark scheme. Benchmark 1 is the DFedAvg algorithm, which does not use differential privacy technology. Benchmark 2 is the A(DP)2SGD algorithm, which is a decentralized FedAvg algorithm with DP. Benchmark 3 is the PPSGD algorithm, which improves the adaptability of the model through a personalized privacy protection mechanism and is combined with the differential privacy benchmark.

[0071] Embodiment 2,

[0072] In specific application scenarios, the SVHN street view house number dataset can be used for distributed learning, such as collaborative learning among multiple research institutions or companies to jointly train a digital recognition model for identifying and classifying house numbers in street view images. Each participant uses its own subset of data for local training and shares model updates through distributed learning without exchanging actual data. In this process, noise can be introduced through differential privacy technology to protect the data privacy of each participant and ensure that personal or company-level data will not be leaked, thereby improving the overall performance and security of the model. This approach can be used to develop intelligent traffic management systems, road sign recognition systems in autonomous driving, and real-time house number recognition functions in smart cities.

[0073] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solution of the present invention, and are not intended to limit the specific implementation methods of the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the claims of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A regularization constraint adaptive adjustment method for privacy-preserving heterogeneous decentralized learning, applied to a federated learning system, where the federated learning system contains multiple clients, and each client has a local dataset for image classification and recognition tasks; characterized in that: The method comprises: S1. Each client Local Model Divide into shared models and retain the model , for the shared model Regularization constraints are applied to limit the update range of model parameters, as follows: S11. Local Model The local model is processed by the deep neural network layer by layer. Divided into shared models and retain the model , where the shared model Used to exchange between neighbor clients and retain the model Keep it locally to maintain personalization features; S12. Regularize the shared model to limit the model bias and reformulate the client's loss function as: (1) (2) In formula (1) and formula (2), is the cross entropy loss function, is the regularization parameter, is the clipping threshold, and Represent the shared models of this round and the previous round respectively, represents the regularization term, Indicates The debiasing parameter in the round iteration, when t=0, = ; S2, Client Use the shared model after the previous round of aggregation and locally retained models , based on local datasets Perform gradient descent updates; S3, regularization parameters through KL divergence Perform dynamic update and adjustment; S4. Shared Model Perform differential privacy protection, and then broadcast the noisy shared model to neighbor clients to ensure privacy protection while completing decentralized communication; S5. The neighboring clients of client i receive the noisy shared model and aggregate it to obtain the local model for the next iteration round.

2. The method for adaptively adjusting regularization constraints for privacy-preserving heterogeneous decentralized learning according to claim 1, characterized in that: The gradient descent update rule is as follows: (3) In formula (3), is the learning rate, is the loss function based on local data, represents the debiasing parameter, Represents a local dataset.

3. The method for adaptively adjusting regularization constraints for privacy-preserving heterogeneous decentralized learning according to claim 1, characterized in that: The KL divergence is used to measure the difference between the current shared model and the previous round of shared model distribution. The dynamic update adjustments are as follows: (4) In formula (4), is the step size parameter, and denote the shared models of round t and round t-1 respectively, represents the KL divergence, represents the updated regularization parameter of the t+1th round, represents the regularization parameter of this round.

4. The method for adaptively adjusting regularization constraints for privacy-preserving heterogeneous decentralized learning according to claim 3, characterized in that: The differential privacy protection of the shared model refers to: achieving differential privacy by adding noise to the shared model, where the noise obeys Gaussian distribution, as follows: (5) (6) In formula (5) and formula (6), Represents the shared model parameters that need to be shared with neighboring clients. represents the clipping threshold, The expected value is 0 and the variance is Gaussian noise of , which meets the privacy budget of differential privacy Require, Representatives Shared model parameters after noise addition.

5. The method for adaptively adjusting regularization constraints for privacy-preserving heterogeneous decentralized learning according to claim 4, characterized in that: The step S5 is specifically as follows: the neighboring clients of client i receive the noisy shared model and aggregate it to obtain the local model of the next iteration round, specifically: (7) In formula (7), represents the local model of the t+1th round obtained by aggregation, represents the local model of round t, represents the weight from client j to client i in round t, Indicates the neighbor client updated in round t Scalar push weight of Indicates the neighbor client updated in round t+1 Scalar push weight of ; From the client In the The debiasing parameter in the round iteration.

Citation Information

Patent Citations

  • Crowd sensing data privacy protection fusing adaptive differential privacy and federated learning

    CN115510482A

  • Personalized federal learning identification method and system based on differential privacy

    CN117196012A

  • Heterogeneous federated learning-oriented adaptive differential privacy protection method

    CN119293861A