Adaptive Differential Privacy Protection Method for Heterogeneous Federated Learning

By adopting an adaptive differential privacy protection method based on regularization and momentum mechanisms in federated learning, differentiated noise is performed according to the importance of the model level, and the negative impact of traditional method noise on model performance is solved, achieving more flexible privacy protection and higher model performance.

CN119293861BActive Publication Date: 2025-07-01QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411845784.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-07-01
Estimated Expiration
2044-12-16

AI Technical Summary

Technical Problem

The traditional differential privacy federated learning method often has a negative impact on the performance of the model due to the introduction of noise, especially reducing the convergence speed and accuracy of the model.

Method used

Adaptive differential privacy federated learning method based on regularization and momentum mechanisms is adopted to introduce adaptive noise at different levels of the model, and differentiate noise addition is performed according to the importance of each layer, combining regularization and momentum mechanism to optimize the model.

Benefits of technology

This method effectively enhances the flexibility of privacy protection, reduces the negative impact of noise on model accuracy, optimizes model performance, and allows the model to learn data characteristics more accurately while ensuring user privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119293861B_ABST
    Figure CN119293861B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of federated learning, and more specifically, relates to an adaptive differential privacy protection method for heterogeneous federated learning. The method includes introducing adaptive noise at different levels of the model. The contributions of the various levels of the model to the overall learning effect are different. In order to minimize the damage to key features while adding noise, this paper performs differential noise addition on different parts based on the importance degree of the model levels, that is, less noise is applied to the more important levels, while more noise is applied to the secondary levels. The present invention solves the problem that traditional differential privacy federated learning methods usually have a negative impact on the performance of the model due to the introduction of noise, especially reducing the convergence speed and accuracy of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of federated learning, and more specifically, relates to an adaptive differential privacy protection method for heterogeneous federated learning. Background Art

[0002] Federated Learning (FL) is a distributed machine learning framework that enables multiple distributed clients to collaboratively train a shared model while keeping the data local, thus significantly reducing the risk of privacy leakage caused by data centralization. However, with the wide application of deep learning models in various tasks (such as signal processing, network modeling, and traffic analysis, etc.), research shows that even in the federated learning framework, the trained model may still leak users' sensitive information and be exposed to potential privacy threats. Therefore, Differential Privacy (DP) technology has become an important means to protect data privacy in federated learning.

[0003] Chinese Patent Document CN118674014A discloses a privacy-protected multi-level heterogeneous vertical federated learning method, which adopts a two-level data distribution: for data that is geographically dispersed and different entities within each region have different feature sets of local samples, the present invention allows effective model training in such an environment. It adopts model heterogeneity processing: for the dataset heterogeneity caused by different label distributions and model architectures between regions, the present invention conducts collaborative training through a specific mechanism. It enables parties with different data features in heterogeneous regions to collaboratively train a neural network model without sharing local data, and further uses differential privacy technology for privacy enhancement.

[0004] The core idea of differential privacy is to add noise to the model update, so that attackers cannot accurately infer the specific content of the original data from the analysis results, thereby effectively protecting user privacy. In the federated learning scenario, differential privacy introduces an appropriate amount of noise into the model parameters uploaded by the client. However, traditional differential privacy federated learning methods usually have a negative impact on the performance of the model due to the introduction of noise, especially reducing the convergence speed and accuracy of the model. Therefore, how to achieve a balance between privacy protection and model performance has become a key challenge in differential privacy federated learning. Summary of the Invention

[0005] The present invention aims to overcome at least one defect of the above-mentioned prior art, and provides an adaptive differential privacy federated learning method based on regularization and momentum mechanism to solve the problem that the traditional differential privacy federated learning method usually has a negative impact on the performance of the model due to the introduction of noise, especially reducing the convergence speed and accuracy of the model. Specifically, adaptive noise is introduced at different levels of the model. Since the contributions of different levels of the model to the overall learning effect are different, in order to minimize the damage to key features while adding noise, this paper performs differential noise addition on different parts based on the importance degree of the model levels, that is, less noise is applied to the more important levels, while more noise is applied to the less important levels. This adaptive noise addition strategy not only enhances the flexibility of privacy protection, but also effectively reduces the negative impact of noise on the model accuracy, thereby optimizing the model performance.

[0006] The present invention provides an apparatus for an adaptive differential privacy protection method based on heterogeneous federated learning.

[0007] The present invention also provides a computer-readable storage medium for implementing the above method.

[0008] The detailed technical solution of the present invention is as follows:

[0009] An adaptive differential privacy protection method for heterogeneous federated learning, the method comprising:

[0010] S1. Perform hierarchical processing on the deep neural network of the local model, set different loss functions for each layer and update the local model loss function for optimization to better capture the learning objectives of each layer;

[0011] S2. Aggregate the optimized deep neural network to obtain the initial global parameters of the central server, and the central server distributes the global parameters; the local model receives the global parameters, and then calculates the sample gradient in the current iteration round based on the local dataset;

[0012] S3. Apply the gradient descent algorithm, calculate the local gradient based on the sample gradient and update the local model parameters, and calculate the corresponding loss value of the local model;

[0013] S4. Allocate the privacy budget for each layer according to the calculated loss value, and perform noise addition processing on the local model parameters;

[0014] S5. Upload the updated local model parameters, that is, the noise-added model parameters, to the central server for parameter aggregation, and distribute the updated global model parameters to each client, and iterate and update until the iteration threshold.

[0015] Further, the S1 specifically includes:

[0016] S11. Perform hierarchical processing on the deep neural network of the local model, and divide the deep neural network of the local model into a convolutional layer and a fully connected layer , that is ;

[0017] S12. Set different loss functions according to the characteristics and requirements of different layers to optimize the model:

[0018] Add a piecewise regularization term to the convolutional layer to control the change of the convolutional layer parameters:

[0019] (1);

[0020] In formula (1), set the total number of global iterations as , the total number of local iterations as , in the th round of iteration, ; represents the client after the th local iteration in the tth global iteration, k ≤ , while represents the convolutional layer parameters of the (t - 1)th global iteration. This regularization term effectively reduces the negative impact of high noise on the feature extraction ability by restricting the update amplitude of the convolutional layer parameters, is the clipping threshold;

[0021] Add a momentum term to the fully connected layer to accelerate the convergence speed of the fully connected layer:

[0022] (2);

[0023] In formula (2), represents the fully connected layer parameters of the client after the th local iteration in the tth global iteration, while represents the fully connected layer parameters of the client after the th local iteration in the tth global iteration;

[0024] S13. Update the local model loss function :

[0025] (3);

[0026] In formula (3), , is the cross-entropy loss function; 、 are the coefficients for controlling the regularization mechanism and the momentum mechanism respectively, B represents the total number of samples in the local dataset, and bi represents the b-th sample in the local dataset of the i-th client.

[0027] Furthermore, the local model receives the global parameters and then calculates the sample gradients in the current iteration round based on the local dataset, specifically including:

[0028] In the local model, for the client , the local dataset is used to obtain the sample gradients in each iteration round. Let the total number of iterations be . In the -th global iteration, , the client The sample gradient in the -th global iteration is :

[0029] (4);

[0030] In formula (4), represents the model parameters of the client in the -th global iteration. represents the sample sampled by the client from its local dataset in the -th global iteration. Let and , that is:

[0031] (5);

[0032] (6);

[0033] In formulas (5) to (6), represents the model parameters of the convolutional layer of client i in the -th iteration and the -th local iteration. represents the model parameters of the global convolutional layer in the -th iteration. represents the model parameters of the fully connected layer of client i in the -th iteration and the -th local iteration. represents the model parameters of the fully connected layer of client i in the -th iteration and the -th local iteration;

[0034] The local gradient is calculated using the sample gradients as:

[0035] (7);

[0036] (8);

[0037] In formulas (7) to (8): and respectively represent the local gradients of the convolutional layer and the fully connected layer at the th round of iteration and the th local iteration of the client; represents the local dataset of the client in the total number of samples.

[0038] Furthermore, the S3 specifically includes:

[0039] (9);

[0040] (10);

[0041] In formulas (9) to (10): and respectively represent the parameters of the convolutional layer and the fully connected layer of the client at the th local iteration updated; and and respectively represent the learning rates of the convolutional layer and the fully connected layer during training;

[0042] Then, through the loss function calculate the loss values of the convolutional layer and the fully connected layer and .

[0043] Furthermore, the S4 specifically includes:

[0044] When the privacy budget allocated to the convolutional layer is , the noise allocated to the fully connected layer is satisfying differential privacy, where , is a coefficient for controlling the privacy budget allocation; is the privacy budget;

[0045] The main function of the convolutional layer is to extract features from the input data through convolutional operations and has universality. Therefore, in order for it to better learn global knowledge in the early stage of iteration, the present invention allocates a smaller privacy budget in the early stage, and the ratio of the privacy budget of the convolutional layer to the fully connected layer is 2:1, that is, when , ​If it is 1.5, the noise ratio between the convolutional layer and the fully connected layer is: = 1:2;

[0046] When According to the calculated loss value, allocate the privacy budget for each layer, that is:

[0047] (11);

[0048] Then, perform noise addition processing on the local model parameters to be uploaded:

[0049] , (12);

[0050] (13);

[0051] (14);

[0052] In formulas (12) to (14), and are the model difference parameters between the uploaded convolutional layer and the fully connected layer, and represent the parameters of the convolutional layer and the fully connected layer of the client at the th local iteration obtained by updating, is the clipping threshold, and are Gaussian noises with expectations of 0 and variances of and respectively; and represent the parameters of the convolutional layer and the fully connected layer of the client at the th local iteration obtained by updating.

[0053] In another aspect of the present invention, a device is further provided, including:

[0054] At least one processor; and

[0055] A memory storing a computer program running on the processor; when the computer program is executed by the at least one processor, the at least one processor is caused to execute the above-mentioned adaptive differential privacy protection method for heterogeneous federated learning.

[0056] In another aspect of the present invention, a machine-readable storage medium is further provided, which stores an executable computer program, and when the computer program is executed, the machine is caused to execute the above-mentioned adaptive differential privacy protection method for heterogeneous federated learning.

[0057] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0058] (1) The adaptive differential privacy protection method for heterogeneous federated learning provided by the present invention divides different levels in the model and applies noise according to the functions of each layer, thereby achieving more flexible privacy protection. This method effectively reduces the impact of noise on the overall performance of the model, enabling the model to more accurately learn data features while ensuring user privacy.

[0059] (2) The adaptive differential privacy protection method for heterogeneous federated learning provided by the present invention implements a regularization and momentum mechanism (RM) for the divided model levels to accelerate the convergence speed of the model and improve the convergence accuracy. Through theoretical analysis of the convergence, there is an optimal coefficient for the regularization and momentum mechanism, which can maximize the performance of the model and ensure stability and adaptability in different data environments; the significant advantages of this method in terms of convergence speed and final model accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 is a flowchart of the adaptive differential privacy protection method for heterogeneous federated learning according to the present invention.

[0061] Figure 2 is a schematic diagram of a comparative experiment in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] The present invention will be further described below in conjunction with the drawings and embodiments.

[0063] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0064] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0065] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0066] Embodiment 1

[0067] Refer Figure 1, this embodiment provides an adaptive differential privacy protection method for heterogeneous federated learning, which is an adaptive differential privacy federated learning method based on regularization and momentum mechanisms and is applied to a federated learning system. The federated learning system includes multiple clients and a central server, and each client has a local dataset for image classification and recognition tasks. The method includes:

[0068] S1. Perform hierarchical processing on the local model and set different loss functions:

[0069] Perform hierarchical processing on the deep neural network of the local model, set different loss functions for each layer according to the characteristics and requirements of different layers, and update the local model loss function for optimization to better capture the learning objectives of each level.

[0070] In the present invention, the processing entity of the local model is the deep neural network, and the parameters of this local model are the parameters of the deep neural network. Preferably, the deep neural network of the local model is a convolutional neural network model CNN.

[0071] S11. Perform hierarchical processing on the deep neural network of the local model, and divide the deep neural network into a convolutional layer and a fully connected layer , that is .

[0072] The main function of the convolutional layer is to extract features from the input data through convolutional operations, and generate new feature maps by sliding multiple filters (convolution kernels) on the input feature map. These feature maps capture the edges, textures, and other low-level features of the image, and provide rich information for subsequent layers.

[0073] The main function of the fully connected layer is to integrate the features extracted by the convolutional layer to generate the final output. The fully connected layer can learn higher-level abstract features by linearly combining the features, and map these features to specific output categories or regression results.

[0074] S12. Set different loss functions according to the characteristics and requirements of different layers to optimize the model:

[0075] In CNN, the functions of the convolutional layer and the fully connected layer are different, and their impacts on the model performance also vary. Due to its strong universality, the convolutional layer should be as close as possible to the global model during the local iteration process to effectively capture features and reduce local biases.

[0076] The intensity of the noise introduced by differential privacy is related to the sensitivity of the sample data. Usually, when introducing noise, the second norm of the gradient is used to measure the sensitivity of the sample data in the local dataset. Therefore, clipping techniques are usually needed to control the sensitivity of the sample data. Considering the impact of the clipping threshold on the update of the convolutional neural network model, a piecewise regularization term is added in the convolutional layer , so as to control the change of the convolutional layer parameters:

[0077] (1);

[0078] In formula (1), the total number of global iterations is set to , the total number of local iterations is , and in the th round of iteration, ; represents the convolutional layer parameters of client after the th local iteration in the t-th global iteration, while represents the convolutional layer parameters of the (t - 1)-th global iteration, and is the clipping threshold. This regularization term effectively reduces the negative impact of high noise on the feature extraction ability by restricting the update amplitude of the convolutional layer parameters.

[0079] The main function of the fully connected layer is to integrate the features extracted by the convolutional layer to generate the final output. Since the fully connected layer usually carries the learning of personalized knowledge, it is hoped to ensure its learning of local personalized knowledge as much as possible during the optimization process;

[0080] A momentum term is added in the fully connected layer to accelerate the convergence speed of the fully connected layer:

[0081] (2);

[0082] In formula (2), represents the fully connected layer parameters of client after the th local iteration in the t-th global iteration, while represents the fully connected layer parameters after the th local iteration in the t-th global iteration.

[0083] This method ensures that the fully connected layer can effectively utilize the previous learning results, accelerating the learning and adaptation of personalized features. By combining the above hierarchical noise addition and dynamic regularization strategies, the collaborative optimization of the convolutional layer and the fully connected layer can be achieved, thereby improving the overall performance and convergence speed of the model. This hierarchical method not only optimizes the learning ability of the local model but also effectively enhances the robustness of the global model, providing a more effective solution for distributed machine learning.

[0084] S13. Update the loss function :

[0085] (3);

[0086] In formula (3), , is the cross-entropy loss function; , are the coefficients controlling the regularization mechanism and the momentum mechanism respectively. B represents the total number of samples in the local dataset, and b represents the b-th sample in the local dataset.

[0087] S2. The local model receives the global parameters and calculates the sample gradients in the current iteration:

[0088] Aggregate the optimized deep neural network to obtain the initial global parameters of the central server. The central server distributes the global parameters, that is ; The local model receives the global parameters and then calculates the sample gradients in the current iteration based on the local dataset. Among them, the global parameters include the model parameters of the global convolutional layer and the model parameters of the global fully connected layer

[0089] In the local model, each client has a local dataset for the image classification and recognition task . In this embodiment, the client is the client, and the local dataset contains samples, which can be obtained from the CIFAR-10 dataset.

[0090] The CIFAR-10 dataset is a standard dataset widely used for image classification tasks. It was created by Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton from the University of Toronto, Canada. The dataset contains 10 classes of different color images, with 6000 images in each class, for a total of 60000 color images of 32x32 pixels. The image classes include airplanes, cars, birds, cats, deer, dogs, frogs, horses, ships, and trucks.

[0091] The CIFAR-10 dataset is divided into a training set and a test set. The training set contains 50,000 images, and the test set contains 10,000 images. The pixel values of each image range from 0 to 255, representing the color intensity in the RGB color space. Each image has a corresponding label, which is an integer between 0 and 9, used to represent the category it belongs to.

[0092] In specific application scenarios, the CIFAR-10 dataset can be used for distributed learning. For example, collaborative learning can be carried out among multiple research institutions or companies to jointly train an image classification model for identifying and classifying images of common objects. Each participant uses their own subset of data for local training and shares model updates through distributed learning without exchanging actual data. In this process, noise can be introduced through differential privacy techniques to protect the data privacy of each participant, ensuring that personal or company-level data will not be leaked, thereby improving the overall performance and security of the model. This method can be used to develop intelligent monitoring systems, object recognition systems in autonomous driving, and real-time image classification functions in augmented reality devices.

[0093] For the client , use the local dataset to obtain the sample gradient in each iteration round. Set the total number of iterations to . In the th iteration round ( ), the client 's sample gradient in the th iteration round is :

[0094] (4);

[0095] In formula (4), represents the model parameters of the client in the th iteration round, represents the sample sampled by the client from its local dataset in the th iteration round. Let and , that is:

[0096] (5);

[0097] (6);

[0098] In formulas (5) - (6), represents the th iteration, and the The model parameters of the convolutional layer of client i in the current local iteration represent the model parameters of the global convolutional layer at the th iteration, represent the th iteration of the model parameters of the fully connected layer of client i in the current local iteration, represent the th iteration of the model parameters of the fully connected layer of client i in the current local iteration;

[0099] The local gradient calculated using the sample gradient is:

[0100] (7);

[0101] (8);

[0102] In formulas (7) to (8): and respectively represent the local gradients of the convolutional layer and the fully connected layer of client in the th round of iteration and the th local iteration; represents the total number of samples in the local dataset of client ; represents the th sample in the local dataset .

[0103] S3. Apply the gradient descent algorithm and calculate the loss value:

[0104] Apply the gradient descent algorithm to update the local model parameters based on the local gradient calculated from the sample gradient, and calculate the corresponding loss value (loss) of the local model.

[0105] (9);

[0106] (10);

[0107] In formulas (9) to (10): and respectively represent the parameters of the convolutional layer and the fully connected layer of client updated at the th local iteration, and respectively represent the learning rates of the convolutional layer and the fully connected layer during training;

[0108] Then, through the loss function Calculate their respective loss values and .

[0109] S4. Add noise to the local model parameters based on the loss value:

[0110] Allocate the privacy budget for each layer according to the calculated loss value, and add noise to the local model parameters, that is, the parameters of the deep neural network.

[0111] Lemma 1: For the mechanism , where is the input data set, is a -dimensional function, and its , the noise is Gaussian noise with an expectation of 0 and a variance of , then the mechanism satisfies differential privacy;

[0112] If there is a mechanism , where is also a -dimensional function, and for , where correspond to the Gaussian noise added to each dimension respectively, then the mechanism also satisfies differential privacy, provided that the following conditions are met: ; where is the privacy budget, is the failure probability of the client , is the variance of the Gaussian noise added to each dimension.

[0113] It can be seen from Lemma 1 that when the privacy budget allocated to the convolutional layer is , the noise (privacy budget) allocated to the fully connected layer is satisfies differential privacy, where is a coefficient that controls the privacy budget allocation.

[0114] Since the main role of the convolutional layer is to extract features from the input data through convolutional operations and has strong universality, in order for it to better learn global knowledge in the early stage of iteration, the present invention allocates a smaller privacy budget in the early stage, and the ratio of the privacy budget of the convolutional layer to the fully connected layer is 2:1, that is, when , is 1.5, then the noise ratio of the convolutional layer to the fully connected layer is: = 1:2;

[0115] When When, allocate the privacy budget for each layer according to the calculated loss value, that is:

[0116] (11);

[0117] Then, perform noise addition processing on the local model parameters to be uploaded:

[0118] , (12);

[0119] (13);

[0120] (14);

[0121] In formulas (12) to (14), and are the difference parameters of the uploaded convolutional layer and fully connected layer model respectively, is the clipping threshold, and are Gaussian noises with an expectation of 0 and variances of and respectively.

[0122] S5. Loop and iterate until the end:

[0123] Upload the updated local model parameters, that is, the locally model parameters with noise addition, to the central server for parameter aggregation, and send the aggregated new global model parameters to each client, and iterate and update like this until the iteration threshold.

[0124] (15);

[0125] (16);

[0126] In formulas (15) to formula (16), and are the parameters of the convolutional layer and fully connected layer in the global model respectively, and are the learning rates of the convolutional layer and fully connected layer during global update respectively, and send the updated global model parameters to each client.

[0127] This embodiment not only has good privacy, but also shows good convergence effect:

[0128] Privacy: There exist constants and , such that given the number of iterations , for any , when At this time, each client satisfies - Differential privacy.

[0129] : is the privacy budget, which reflects the strength of privacy protection. The smaller the value, the stronger the privacy protection.

[0130] : is the leakage probability allowed in differential privacy, usually a very small value.

[0131] : is the number of iterations of the algorithm.

[0132] : is the sample sampling rate.

[0133] Convergence:

[0134] (17);

[0135] In formula (17), represents the gradient deviation caused by clipping, where , , are constants, representing the inherent difference, Lipschitz coefficient, and bounded variance in federated learning respectively;

[0136] (18);

[0137] (19);

[0138] (20);

[0139] This embodiment conducts experiments on the CIFAR10 dataset and obtains the following experimental results:

[0140] First, set , , , , , , , .

[0141] When not considering the clipping deviation, it can be seen from the second term of the convergence result that when a given is given, there is an optimal such that the upper bound of the convergence of is minimized. Let , corresponding to the optimal value of the local iteration number . Set , and set ​= 1, minimize Obtain .

[0142] The results of the comparative experiments are as follows Figure 2 shown. The solution of the present invention improves the accuracy by 9% to 15% compared with the benchmark solution. The hierarchical solution refers to the model method in this solution that does not adopt the regularization mechanism and the momentum mechanism. Benchmark 1 is the FedDPA algorithm, which is based on the personalized screening of dynamic Fisher information and combines an adaptive constraint technique to alleviate the convergence difficulty problem caused by data non-independent and identically distributed and differential privacy clipping operations. Benchmark 2 is the DP-FedAvg algorithm, which is a FedAvg algorithm with DP. Benchmark 3 is the PPSGD algorithm, which improves the model adaptability through a personalized privacy protection mechanism while ensuring privacy security.

[0143] Embodiment 2

[0144] This embodiment provides a device for implementing an adaptive differential privacy protection method for heterogeneous federated learning. The device includes:

[0145] At least one processor; and

[0146] A memory that stores a computer program. When the computer program is executed by the at least one processor, the at least one processor executes the adaptive differential privacy protection method for heterogeneous federated learning as described above.

[0147] In this embodiment, the electronic device includes but is not limited to: personal computers, server computers, workstations, desktop computers, laptop computers, notebook computers, mobile computing devices, smart phones, tablet computers, cellular phones, personal digital assistants (PDAs), handheld devices, messaging devices, wearable computing devices, consumer electronic devices, etc.

[0148] Embodiment 3

[0149] This embodiment also provides a computer-readable storage medium that stores an executable computer program. When the computer program is executed, the machine executes the adaptive differential privacy protection method for heterogeneous federated learning as described above.

[0150] Specifically, a system or device equipped with a readable storage medium can be provided. On the readable storage medium, software program code for implementing the functions of any one of the above embodiments is stored, and the computer or processor of the system or device reads and executes the computer program stored in the readable storage medium.

[0151] In this case, the program code read from the readable medium itself can implement the functions of any one of the above embodiments. Therefore, the machine-readable code and the readable storage medium storing the machine-readable code constitute a part of this specification.

[0152] Examples of the readable storage medium include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD-RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code can be downloaded from a server computer or a cloud via a communication network.

[0153] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0154] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more flows and / or Figure 1 blocks specified in the one or more blocks.

[0155] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more flows and / or Figure 1 blocks specified in the one or more blocks.

[0156] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide means for implementing the functions specified in Figure 1One or more processes and / or boxes Figure 1 Steps of the functions specified in one or more boxes.

[0157] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the technical solutions of the present invention, rather than limitations on the specific implementation manners of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the claims of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. An adaptive differential privacy protection method for heterogeneous federated learning, characterized in that: The method comprises: S1. Perform layered processing on the deep neural network, set different loss functions for each layer and update the local model loss function for optimization; S2. Aggregate the optimized deep neural network to obtain the initial global parameters of the central server, and the central server sends the global parameters; the local model receives the global parameters, and then calculates the sample gradient in the current iteration round based on the local data set; in the local model, each client has a local data set for image classification and recognition tasks , the local data set contains B samples; S3, applying the gradient descent algorithm, updating the local model parameters based on the local gradient obtained by the sample gradient calculation, and calculating the corresponding loss value of the local model; S4. Allocate the privacy budget of each layer according to the calculated loss value and perform noise processing on the local model parameters; S5, uploading the updated local model parameters, i.e., the noise-processed local model parameters, to the central server, performing parameter aggregation, and sending the updated global model parameters to each client, and iterating and updating until the iteration threshold is reached; The S1 specifically includes: S11. Perform layered processing on the local model’s deep neural network, dividing the deep neural network into convolutional layers and fully connected layers ,Right now ; The convolutional layer generates a new feature map by sliding multiple filters on the input feature map; S12. Set different loss functions according to the characteristics and requirements of different layers to optimize the model: Add a piecewise regularization term to the convolutional layer : (1); In formula (1), the total number of global iterations is set to , the total number of local iterations is , in In the round iteration, ; Representing the client In the tth global iteration, The convolutional layer parameters after local iterations, k≤ ,and represents the convolutional layer parameters of the t-1th global iteration, is the trimming threshold; Adding a momentum term to the fully connected layer : (2); In formula (2), Representing the client In the tth global iteration, The fully connected layer parameters after the local iteration, and Represents the client In the tth global iteration, The fully connected layer parameters after the local iteration; S13. Update local model loss function : (3); In formula (3), , is the cross entropy loss function; , are the coefficients controlling the regularization mechanism and momentum mechanism, B represents the total number of samples in the local dataset, and b i represents the bth sample in the local dataset of the i-th client; The S4 specifically includes: When the privacy budget allocated to the convolutional layer is When , the noise assigned to the fully connected layer is satisfy Differential privacy, where , is a coefficient that controls the allocation of privacy budget; Budget for privacy; The ratio of the privacy budget of the convolutional layer to the fully connected layer is 2:1, that is, when hour, If is 1.5, the noise ratio of the convolutional layer to the fully connected layer is: =1:2; when When the calculated loss value is used, the privacy budget at each level is allocated, i.e.: (11); Then add noise to the model parameters to be uploaded: , (12); (13); (14); In formulas (12)~(14), and They are the model difference parameters of the uploaded convolutional layer and the fully connected layer, respectively. They are the crop thresholds of the convolutional layer and the fully connected layer, respectively. and The corresponding expectation of the convolution layer and the fully connected layer is 0, and the variance is and Gaussian noise, and Indicates the updated The parameters of the client's convolutional layer and fully connected layer during the local iteration.

2. The adaptive differential privacy protection method for heterogeneous federated learning according to claim 1, characterized in that: The local model receives the global parameters and then calculates the sample gradient in the current iteration round based on the local data set, specifically including: In the local model, for client i, the local dataset is used Get the sample gradient in each iteration round and set the total number of iterations to , in In the round iteration, , Client In the The sample gradient in the round iteration is : (4); In formula (4), Represents the client In the Model parameters in round global iteration, Represents the client In the The global iteration starts from its local dataset The sample is sampled from and ,Right now: (5); (6); In formulas (5)~(6), Representative The first iteration The model parameters of the convolutional layer of client i in the local iteration, Representative The model parameters of the global convolutional layer at the iteration, Representative The first iteration The model parameters of the fully connected layer of client i in the local iteration, Representative The first iteration Model parameters of the fully connected layer of client i at the local iteration; The local gradient is calculated using the sample gradient: (7); (8); In formulas (7)~(8), and Respectively represent the client In the In the round iteration Local gradients of convolutional and fully connected layers at the local iteration; Represents the client Local dataset The total number of samples in .

3. The adaptive differential privacy protection method for heterogeneous federated learning according to claim 2, characterized in that: The S3 specifically includes: (9); (10); In formulas (9)~(10), and Respectively represent the updated The client side of the local iteration The convolutional layer and fully connected layer parameters, and Respectively represent the learning rates of the convolutional layer and the fully connected layer during training; Then, through the loss function Calculate the respective loss values and .

4. A device based on an adaptive differential privacy protection method for heterogeneous federated learning, characterized in that: The device comprises: at least one processor; and a memory having stored thereon a computer program running on the processor; Wherein, when the computer program is executed by the processor, the steps of the adaptive differential privacy protection method for heterogeneous federated learning as described in any one of claims 1 to 3 are implemented.

5. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 3 are implemented.

Citation Information

Patent Citations

  • Multi-level heterogeneous longitudinal federated learning method for privacy protection

    CN118674014A

  • An image classification method based on differential privacy and hierarchical correlation propagation

    CN109034228A

  • Channel-based hierarchical structure differential privacy federal learning method and device

    CN117473548A