A personalized federated learning method and device based on privacy protection

By calculating the client frequency sparseness for K-Means clustering, the personalized collaborative training problem of Non-IID data distribution in federated learning is solved, and privacy protection and model accuracy are improved.

CN115329885BActive Publication Date: 2025-07-29ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211014315.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-23
Publication Date
2025-07-29
Estimated Expiration
2042-08-23

AI Technical Summary

Technical Problem

When facing the distribution of Non-IID data, existing federated learning methods are difficult to achieve personalized collaborative training while ensuring privacy protection, and may lead to overfitting or bias.

Method used

K-Means clustering is performed by calculating the frequency sparseness of the client, forming K clusters, and model parameters aggregating clients in the same cluster to realize personalized federated learning and protect user privacy.

Benefits of technology

It realizes that on the premise of protecting user privacy, the collaborative training effect of federated learning is improved, overfitting and bias are reduced, and model accuracy is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0003811903040000033
    Figure BDA0003811903040000033
  • Figure BDA0003811903040000041
    Figure BDA0003811903040000041
  • Figure BDA0003811903040000067
    Figure BDA0003811903040000067
Patent Text Reader

Abstract

The present invention discloses a personalized federated learning method and device based on privacy protection. Each client conducts training in each round, obtains the trained local model matrix parameters and uploads them to the corresponding personalized model; according to the change amount of the trained local model matrix parameters in this round, the threshold of each client is calculated, and then the frequency sparsity of each client is calculated; according to the frequency sparsity set, the clients are clustered; the local model parameters uploaded by the clients in the same cluster are averaged, and the average value is used as the new global model matrix parameters and sent to the clients in the same cluster; repeat until the global model converges, and complete the training of personalized federated learning. The present invention always protects user privacy during the calculation of frequency sparsity, clusters the clients according to the similarity of sparsity to form K clusters, and performs aggregation operations on the client distributions of these K clusters to achieve the effects of collaborative training and personalized federated learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of federated learning, and particularly relates to a personalized federated learning method and device based on privacy protection. Background Art

[0002] Nowadays, deep learning has received great attention in both academia and industry. Since the performance of deep learning has been greatly improved compared with traditional algorithms, deep learning is widely used in various fields, such as machine translation, image recognition, autonomous driving, natural language processing, etc. Deep learning is changing the way people live. Its success depends on powerful computers and the availability of large amounts of data. However, learning systems that require all data to be input into a learning model running on a central server bring serious privacy problems. Moreover, with the rise of the Internet of Things and edge computing, big data often does not stick to a single whole, but is distributed in many aspects. How to safely and effectively update and share models among multiple locations is a new challenge faced by various computing methods. Due to considerations of data privacy and security, data owners cannot directly share data to jointly train deep learning models. People have begun to seek a method. To solve the data silo problem and user privacy, federated learning has emerged as a highly promising solution. Now federated learning (FL) has become a popular distributed machine learning paradigm. Federated learning has been proposed and attracted great attention because it can collaboratively train a shared global model in a decentralized data environment.

[0003] A large number of distributed clients participate in the learning process together by uploading the gradients of their local models (or model weights) to the server through multiple iterations without sharing the original data between clients. At the beginning of the FL task, the server initializes the global model. In each learning iteration, the server distributes the current global model matrix parameters to the selected clients. Each selected client continues to independently train the received model using its local data by following a predefined learning protocol. At the end of each learning iteration, the server collects and aggregates the updates from the clients using a gradient aggregation rule (such as FedAvg). This mechanism for protecting user privacy has been widely applied in the practical deployment of FL in recent years, such as loan status prediction, health condition assessment, and next word prediction. Although FL has been proven effective in generating better federal models, it may not be the best solution for each client because the data distribution of the clients may be Non-IID. To better adapt to its unique data distribution, personalized operations need to be considered. Existing research has noticed the problem of data heterogeneity and proposed some personalized methods to solve this problem. These include schemes such as using fine-tuning of the federal model to achieve personalization, multi-task learning, and knowledge extraction. Although these methods can promote personalization to a certain extent, they have a significant drawback, that is, the personalization process is limited to a single device, which may bring some bias or overfitting problems because the data in the device is extremely limited. At the same time, due to relevant privacy laws, the personalized federated learning framework should fully consider the privacy protection of its clients when designing. Therefore, when facing Non-IID data distribution, how to better ensure the main performance and security of federated learning has become the focus of attention. Summary of the Invention

[0004] The purpose of the present invention is to provide a privacy-preserving personalized federated learning method and device in view of the deficiencies of the prior art.

[0005] The purpose of the present invention is achieved by the following technical solutions: A privacy-preserving personalized federated learning method includes the following steps:

[0006] (1) Initialize the federated learning training environment;

[0007] (2) The server sets a corresponding personalized model for each client in the cloud, and each personalized model distributes its own global model matrix parameters to the corresponding client to start federated learning training;

[0008] (3) The clients participating in the training perform the t-th round of training, obtain the trained local model matrix parameters and upload them to the corresponding personalized model; according to the change amount of the trained local model matrix parameters in this round, calculate the threshold α of each client i,t, and update the matrix [(wcf-matrix) for each client i,t ]; According to each client's matrix [(wcf-matrix) i,t ], calculate the frequency sparsity ε of each client i,t , and obtain the frequency sparsity set {ε 1,t ,ε 2,t ,…,ε i,t ,…,ε k,t};

[0009] (4) According to the frequency sparsity set {ε 1,t ,ε 2,t ,…,ε i,t ,…,ε k,t}Perform K-Means clustering operation and set the frequency sparsity set {ε 1,t ,ε 2,t ,…,ε i,t ,…,ε k,t} is divided into K clusters; and the frequency sparsity ε in the same cluster is i,t The represented clients are grouped into the same cluster;

[0010] (5) Calculate the average of the local model parameters uploaded by the clients of the same cluster, and send the average value as the new global model matrix parameter to the clients of the same cluster;

[0011] (6) Repeat steps (3) to (5) until the global model converges and the training of the personalized federated learning model is completed.

[0012] Furthermore, the step (1) is specifically as follows: setting the overall training round E, local data D, and the overall number of clients k participating in federated learning.

[0013] Furthermore, the step (2) specifically includes the following sub-steps:

[0014] (2.1) The server provides each client p i Set up the corresponding personalized model N i , and initialize each personalized model N i Get the initialized global model matrix parameters i=1,2,…i,…k;The initialized global model matrix parameters The matrix size is W×H;

[0015] (2.2) Each personalized model N i The global model matrix parameters will be initialized Send it to the corresponding client p i , start federated learning training.

[0016] Furthermore, step (3) specifically includes the following sub-steps:

[0017] (3.1) The client p participating in the training i does not share data and performs local model training on the globally issued model weights locally:

[0018] For the t-th round of training, the trained local model matrix parameters are obtained and uploaded to the corresponding personalized model N i ;

[0019] (3.2) Calculate the threshold α of each client p after the t-th round of local model training i , and the calculation formula is as follows: i,t

[0020]

[0021] where α i,t represents the threshold of each client p after the t-th round of local model training; i represents the sub-parameter in the u-th row and v-th column of the local model matrix parameter ; represents the sub-parameter in the u-th row and v-th column of the initialized global model matrix parameter ;

[0022] (3.3) If then update the sub-parameter in the u-th row and v-th column of the matrix (wcf-matrix) i,t :

[0023] If then update the sub-parameter in the u-th row and v-th column of the matrix [(wcf-matrix) i,t :

[0024] where, [(wcf-matrix) i,t u,v represents the sub-parameter in the u-th row and v-th column of the matrix [(wcf-matrix) i,t ;

[0025] Repeat the above steps to update the entire matrix [(wcf-matrix) i,t ;

[0026] (3.4) Calculate the frequency sparsity ε of each client from the matrix [(wcf-matrix) i,t of each client i,t ​​​, and obtain a frequency sparsity set {ε 1,t , ε 2,t , …, ε i,t , …, ε k,t}, and the calculation formula of the frequency sparsity ε i,t is as follows:

[0027]

[0028] Furthermore, the step (4) specifically includes the following sub-steps:

[0029] (4.1) Perform K-Means clustering operation according to the frequency sparsity set {ε 1,t , ε 2,t , …, ε i,t , …, ε k,t}, set the number of clusters to K, and the specific K-Means clustering operation is as follows:

[0030] Randomly select K frequency sparsities from the frequency sparsity set {ε 1,t , ε 2,t , …, ε i,t , …, ε k,t} as the initial centroids; calculate the distance from each frequency sparsity ε i,t to each centroid, and divide the frequency sparsity ε i,t into the cluster corresponding to the nearest centroid; calculate the mean value of all frequency sparsities ε i,t within each cluster, and use this mean value to update the centroid of the cluster. Repeat the steps until the maximum number of iterations is reached, and finally form K clusters: S1, S2…S K ;

[0031] (4.2) Divide the clients represented by the frequency sparsities ε i,t within the same cluster into the same cluster.

[0032] The present invention also provides a privacy-preserving personalized federated learning device, including one or more processors for implementing the above privacy-preserving personalized federated learning method.

[0033] The present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it is used to implement the above privacy-preserving personalized federated learning method.

[0034] The beneficial effects of the present invention are as follows: Compared with traditional federated learning, this method calculates the change amount of the matrix parameters of the locally trained model in each round, and the frequency sparsity of each client. Since the frequency sparsity belongs to a statistic and does not involve the local data of the client, user privacy is always protected during this process. The clients are clustered based on the similarity of the frequency sparsity to form K clusters, and an aggregation operation is performed on the client distributions of these K clusters to achieve collaborative training, thereby achieving the effect of personalized federated learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 Schematic diagram of the local and cloud devices of the present invention;

[0036] Figure 2 Schematic diagram of the local and cloud devices of the present invention;

[0037] Figure 3 Schematic diagram of a personalized federated learning device based on privacy protection provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0039] The present invention attempts to collaboratively train clients with similar distributions using a privacy protection method. Compared with traditional federated learning, this method calculates the sparsity of the frequency by collecting the change frequency of the weights of the locally trained models uploaded by the clients through the server. Since the frequency sparsity belongs to a statistic and does not involve the local data of the client, user privacy is always protected during this process. The clients are clustered based on the similarity of the sparsity to form four clusters, and an aggregation operation is performed on the client distributions of these four clusters to achieve collaborative training, thereby achieving the effect of personalized federated learning.

[0040] Embodiment 1

[0041] As Figure 1 and Figure 2 shown, the present invention provides a personalized federated learning method based on privacy protection, including the following steps:

[0042] (1) Obtain local data D:

[0043] In this embodiment, the MNIST dataset and the ImageNet dataset are used as local data; the MNIST dataset contains a total of 60,000 grayscale images, each image is 28*28 in size, and is divided into 10 categories; 50,000 of them are taken as local data. The ImageNet dataset has a total of 1000 categories, each category contains 1000 samples, each picture is an RGB color image, and each sample is 224*224 in size; 30% of the pictures are randomly selected from each category as local data.

[0044] (2) Initialize the federated learning training environment: Set the overall number of training rounds E, local data D, and the total number of devices k participating in federated learning.

[0045] (3) The server sets corresponding personalized models for each client in the cloud, and each personalized model sends its own global model matrix parameters to the corresponding client, and starts federated learning training.

[0046] Step (3) includes the following sub-steps:

[0047] (3.1) The server sets the corresponding personalized model N i for each client p i , and initializes each personalized model N i to obtain the initialized global model matrix parameters i = 1, 2,... i,... k; the initialized global model matrix parameters are: Its matrix size is W×H, where, is the sub-parameter of the u-th row and v-th column in the initialized global model matrix parameter ; u = 1, 2,... u,... H, v = 1, 2,... v,... W; ;

[0048] (3.2) Each personalized model N i sends the initialized global model matrix parameters to the corresponding client p i , and starts federated learning training.

[0049] (4) The participating clients perform the t-th round of training, obtain the trained local model matrix parameters and upload them to the corresponding personalized model; according to the change amount of the trained local model matrix parameters in this round, calculate the threshold α i,t for each client, and update the matrix [(wcf-matrix) i,t of each client; according to the matrix [(wcf-matrix) i,t of each client, calculate the frequency sparsity ε i,t, obtain the frequency sparsity set {ε 1,t , ε 2,t , …, ε i,t , …, ε k,t}.

[0050] The said step (4) includes the following sub-steps:

[0051] (4.1) The client p participating in the training i does not share data and conducts the t-th round of training, locally trains the globally issued model weights to obtain the trained local model matrix parameters and uploads them to the corresponding personalized model N i ; The said trained local model matrix parameters are: Wherein, is the sub-parameter of the u-th row and v-th column in the local model parameter ;

[0052] (4.2) According to the change amount of the locally trained model matrix parameters in this round, calculate the threshold α i of each client p i,t after the t-th round of local model training. The calculation formula of the threshold α i,t is as follows:

[0053]

[0054] Wherein, α i,t represents the threshold of each client p i after the t-th round of local model training;

[0055] (4.3) If then update the sub-parameter of the u-th row and v-th column in the matrix [(wcf-matrix) i,t :

[0056] If then update the sub-parameter of the u-th row and v-th column in the matrix [(wcf-matrix) i,t :

[0057] Wherein, [(wcf-matrix) i,t u,v represents the sub-parameter of the u-th row and v-th column in the matrix [(wcf-matrix) i,t ;

[0058] Repeat the above steps to update the entire matrix [(wcf-matrix) i,t , the said matrix [(wcf-matrix)​i,t For

[0059] (4.4) Calculate the frequency sparsity ε of each client according to the matrix of each client [(wcf - matrix) i,t , and obtain the set of frequency sparsities {ε i,t , ε 1,t , …, ε 2,t , …, ε i,t , …, ε k,t}, and the calculation formula of the frequency sparsity ε i,t is as follows:

[0060]

[0061] (5) Perform K - Means clustering operation according to the set of frequency sparsities {ε 1,t , ε 2,t , …, ε i,t , …, ε k,t}, divide the frequency sparsities in the set of frequency sparsities {ε 1,t , ε 2,t , …, ε i,t , …, ε k,t} into K clusters; and divide the clients represented by the frequency sparsities within the same cluster into the same cluster. i,t

[0062] Step (5) includes the following sub - steps:

[0063] (5.1) Perform K - Means clustering operation according to the set of frequency sparsities {ε 1,t , ε 2,t , …, ε i,t , …, ε k,t}, set the number of clusters to K, and the specific K - Means clustering operation is as follows:

[0064] Randomly select K frequency sparsities from the set of frequency sparsities {ε 1,t , ε 2,t , …, ε i,t , …, ε k,t} as the initial centroids; calculate the distance from each frequency sparsity ε i,t to each centroid, and divide the frequency sparsity ε i,t into the cluster corresponding to the nearest centroid; calculate the mean value of all frequency sparsities ε i,t within each cluster, and use this mean value to update the centroid of the cluster. Repeat the steps until the maximum number of iterations is reached, and finally form K clusters: S1, S2…S K ; ​

[0065] (5.2) Divide the clients represented by the frequency sparsity ε within the same cluster into the same cluster. i,t

[0066] (6) Average the local model parameters uploaded by the clients in the same cluster, and use the average value as the new global model matrix parameter to be sent to the clients in the same cluster;

[0067] (7) Repeat steps (4) - (6) until the global model converges, and complete the training of the personalized federated learning model.

[0068] Apply the method of the present invention to the image classification task for testing. The convolutional neural networks (CNNs) are trained on MNIST and ImageNet respectively. The MNIST dataset contains 70,000 grayscale handwritten digit images; the ImageNet dataset contains 1,000 classes of RGB color images. The CNNs in the experiment are LeNet and VGG16 models respectively, and the cross-entropy loss function is used. The client local update uses the training method of mini-batch stochastic gradient descent (mini-batch SGD). In terms of data distribution, the data in the clients is non-independent and identically distributed. In this simulation experiment, for the MNIST dataset, set its batch size to 50, learning rate to 0.01, E = 100, k = 50, N = 50, K = 4. For the ImageNet dataset, set its batch size to 64, learning rate to 0.001, E = 100, k = 50, N = 50, K = 4.

[0069] The experiment tests the model accuracy of the above-trained personalized federated learning model and the success rate of the gradient reversal attack against privacy leakage, and verifies the usability of our method. Experiments are carried out on two datasets to illustrate the effectiveness of privacy protection and the model utility.

[0070] Table 1 shows the global model utility results of different clusters. Each client is clustered into different clusters and the global model aggregation is performed separately to achieve personalization, and the success rate against the gradient reversal attack is 0%, and the privacy protection effect exists.

[0071] Table 1: Model accuracy of the personalized federated learning model

[0072]

[0073] Corresponding to the foregoing embodiments of the privacy protection-based personalized federated learning method, the present invention also provides embodiments of a privacy protection-based personalized federated learning device.

[0074] SeeFigure 3 In an embodiment of the present invention, a personalized federated learning device based on privacy protection includes one or more processors for implementing the personalized federated learning method based on privacy protection in the above embodiment.

[0075] The embodiment of the personalized federated learning device based on privacy protection of the present invention can be applied to any device with data processing capabilities. Such a device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented by software, or by hardware, or by a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From a hardware perspective, as Figure 3 shown, it is a hardware structure diagram of any device with data processing capabilities where the personalized federated learning device based on privacy protection of the present invention is located. In addition to Figure 3 the processor, memory, network interface, and non-volatile memory shown, usually according to the actual functions of the device with data processing capabilities where the device in the embodiment is located, other hardware may also be included, which will not be elaborated here.

[0076] For the implementation process of the functions and roles of each unit in the above device, please refer to the implementation process of the corresponding steps in the above method for details, which will not be elaborated here.

[0077] For the device embodiment, since it basically corresponds to the method embodiment, relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0078] The embodiment of the present invention also provides a computer-readable storage medium with a program stored thereon. When the program is executed by a processor, it implements the personalized federated learning method based on privacy protection in the above embodiment.

[0079] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or a memory. The computer-readable storage medium may also be any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit of any device with data processing capabilities and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or is to be output.

[0080] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A personalized federated learning method based on privacy protection, characterized in that, Including the following steps: (1) Initialize the federated learning training environment; (2) The server sets corresponding personalized models for each client in the cloud. Each personalized model sends its own global model matrix parameters to the corresponding client, and starts federated learning training; (3) The clients participating in the training perform the t-th round of training, obtain the trained local model matrix parameters, and upload them to the corresponding personalized model; Calculate the threshold α for each client according to the change amount of the local model matrix parameters trained in this round i,t , and update the matrix of each client [(wcf - matrix) i,t ; According to the matrix of each client [(wcf - matrix) i,t , calculate the frequency sparsity ε i,t , and obtain the frequency sparsity set {ε 1,t , ε 2,t , …, ε i,t , …, ε k,t}; The step (3) specifically includes the following sub-steps: (3.1) The participating client p i Does not share data and performs local model training on the global model weights sent down locally: For the t-th round of training, obtain the trained local model matrix parameters and upload them to the corresponding personalized model N i ; (3.2) Calculate the threshold α of each client p after the t-th round of local model training i as follows: it The calculation formula is as follows: Among them, α i,t represents the threshold of each client p i after the t-th round of local model training; represents the sub-parameter of the u-th row and v-th column in the local model matrix parameter ; represents the sub-parameter of the u-th row and v-th column in the initialized global model matrix parameter ; (3.3) If then update the sub-parameters in the u-th row and v-th column of the matrix [(wcf-matrix) i,t : If then update the sub-parameters in the u-th row and v-th column of the matrix [(wcf-matrix) i,t : where, [(wcf-matrix) i,t u,v represents the sub-parameter in the u-th row and v-th column of the matrix [(wcf-matrix) i,t ;​ Repeat the above steps to update the entire matrix [(wcf-matrix) i,t ; (3.4) From the matrix of each client [(wcf - matrix) i,t , calculate the frequency sparsity ε of each client i,t , and obtain the frequency sparsity set {ε 1,t , ε 2,t , …, ε i,t , …, ε k,t}, and the calculation formula of the frequency sparsity ε i,t is as follows: (4) Perform K-Means clustering operations according to the frequency sparsity set {ε 1,t , ε 2,t , …, ε i,t , …, ε k,t}, and divide the frequency sparsities in the frequency sparsity set {ε 1,t , ε 2,t , …, ε i,t , …, ε k,t} into K clusters; and divide the clients represented by the frequency sparsities ε i,t within the same cluster into the same cluster; (5) Average the local model parameters uploaded by the clients in the same cluster, and send the average value as the new global model matrix parameters to the clients in the same cluster; (6) Repeat step (3) - step (5) until the global model converges, and complete the training of the personalized federated learning model.

2. The personalized federated learning method based on privacy protection according to claim 1, wherein The step (1) is specifically: Set the overall number of training rounds E, local data D, and the total number of clients participating in federated learning k.

3. A personalized federated learning method based on privacy protection according to claim 2, characterized in that, The step (2) specifically includes the following sub-steps: (2.1) The server sets up a corresponding personalized model N for each client p in the cloud i and initializes each personalized model N i to obtain the initialized global model matrix parameters i where i = 1, 2, … i, … k; the matrix size of the initialized global model matrix parameters is W × H; ​ (2.2) Each personalized model N i will initialize the global model matrix parameters and send them to the corresponding client p i , and start federated learning training.

4. The personalized federated learning method based on privacy protection according to claim 1, wherein The step (4) specifically includes the following sub-steps: (4.1) Perform K-Means clustering operation according to the frequency sparsity set {ε 1,t , ε 2,t , …, ε i,t , …, ε k,t}, set the number of clusters to K, and the specific K-Means clustering operation is as follows: Randomly select K frequency sparsities from the frequency sparsity set {ε 1,t , ε 2,t , …, ε i,t , …, ε k,t} as the initial centroids; Calculate the frequency sparsity ε for each one i,t the distances to each centroid, and classify the frequency sparsity ε i,t into the cluster corresponding to the centroid with the closest distance; calculate the mean of all frequency sparsities ε within each cluster i,t and use this mean to update the centroid of the cluster. Repeat the steps until the maximum number of iterations is reached, and finally form K clusters: S1, S2…S K ; (4.2) The clients represented by the frequency sparsity ε within the same cluster i,t are divided into the same cluster.

5. A personalized federated learning device based on privacy protection, characterized in that, Comprising one or more processors for implementing the privacy-preserving personalized federated learning method according to any one of claims 1-4.

6. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it is used to implement the privacy-preserving personalized federated learning method according to any one of claims 1-4.