Personalized federal learning method based on multi-granularity calculation unit

Through the personalized federated learning method of multi-grained computing units, soft clustering and sparse weight activation training is used to solve the problems of slow model convergence and high computational overhead in federated learning, and more efficient personalized model adaptability and computing efficiency are achieved.

CN120258092APending Publication Date: 2025-07-04CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510340430.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In federated learning, non-independent and homogeneous client data leads to slow convergence of the global model, hard clustering cannot capture the diversity of client data distribution, and personalized federated learning has high computational overhead.

Method used

Using a personalized federated learning method of multi-grained computing units, the cluster dependence of the client is calculated through soft clustering, dynamically fused the cluster model and local model, and using sparse weights to activate training to reduce computational overhead.

Benefits of technology

The implementation of personalized federated learning is accelerated, the adaptability and computing efficiency of the model is improved, and the training time and computing cost are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258092A_ABST
    Figure CN120258092A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of federated learning technology application, and particularly relates to a personalized federated learning method based on a multi-granularity calculation unit, which comprises the following steps: taking a global client as a calculation unit at a coarse granularity level, calculating the membership degree of each cluster by each client, and then voting to select the cluster to which the client belongs, initially aggregating the clients with similar data distribution conditions together; a client cluster is used as a computing unit on a medium granularity level, and a cluster model and a client model are dynamically integrated; a single client is used as a computing unit on a fine-grained level, and implementation of personalized federated learning is accelerated by adjusting computing participation degrees of different layers of a neural network in a federated process, adopting a sparse weight activation training mode and executing sparse convolution in forward and backward propagation. According to the method provided by the invention, from the perspective of a multi-granularity calculation unit, under a data heterogeneous scene, the performance of a client model is effectively improved, and the calculation overhead of federal training is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of application of federated learning technology, and particularly relates to a personalized federated learning method based on multi-granularity computing units. Background Art

[0002] Since the concept of federated learning was proposed, it has quickly attracted wide attention and has been verified in multiple practical application scenarios. Federated learning trains models in a distributed environment, effectively protecting data privacy and avoiding problems caused by data centralization. However, due to the high heterogeneity of client data and insufficient model personalization, a single global model is difficult to meet the diverse needs of all clients. To this end, researchers have proposed various personalized federated learning methods to address these challenges. Currently, personalized federated learning has made significant progress and practical applications in fields such as healthcare, finance, and infrastructure services.

[0003] The core idea of personalized federated learning is to capture the personalized information of each client and conduct research based on its heterogeneous data distribution, so as to obtain a high-quality personalized model. Currently, researchers divide personalized federated learning into two categories: global model personalization and personalized model learning. Global model personalization aims to improve the performance of the globally shared model trained federally on heterogeneous data. This process is divided into two stages. First, a shared global model is trained, and then additional training is performed on local data to achieve the personalized goal. Personalized model learning aims to provide a personalized solution by modifying the aggregation process of the non-personalized model to directly construct a personalized model.

[0004] Based on the current research situation of personalized federated learning, there are still the following deficiencies in how to make the model better adapt to the distribution of various heterogeneous data:

[0005] 1. The non-independent and identically distributed client data causes the global model to converge slowly. In a federated learning environment, the institutions participating in federated learning usually have different data distributions. Therefore, training a model suitable for each data distribution has become a major challenge in federated learning.

[0006] 2. Hard clustering cannot capture the diversity of client data distributions. For the federated learning method using hard clustering, a single cluster model cannot enable each participant to well adapt to its multi-distributed data. Secondly, hard clustering cannot effectively utilize the similarities between different clusters, and the data and model parameters between clusters cannot be shared, easily leading to the phenomenon of information islands.

[0007] 3. Personalized federated learning is usually accompanied by higher computational overhead. To meet the personalized needs of each client, the federated learning system may need to train an independent model for each client, which significantly prolongs the training time. In addition, to optimize the model performance and improve its performance on each client, the system often needs to perform multiple iterations and adjustments, further increasing the computational cost. Summary of the Invention

[0008] To solve the above technical problems, the present invention provides a personalized federated learning method based on multi-granularity computing units, including:

[0009] S1. Construct a personalized federated learning system including N clients and S cluster centers, and a cluster center model with initialized parameters is provided in the server.

[0010] S2. The client downloads all cluster center models from the server and calculates the dependence of each cluster model, and at the same time fuses each cluster model and uses it as the local model.

[0011] S3. The client trains the weight parameters of the adaptive aggregation module based on local data, and at the same time adjusts the computational participation of different layers of the neural network in the federated process.

[0012] S4. The client dynamically fuses the local model and the cluster center model through the adaptive aggregation module to obtain an update of the basic layer weights.

[0013] S5. The client adopts a sparse weight activation training method, and accelerates the implementation of personalized federated learning by performing sparse convolution in the forward and backward propagation.

[0014] S6. The client uploads all personalized models to the server for global aggregation.

[0015] S7. Determine whether the federated learning iteration threshold is reached. If so, enter step S8; otherwise, return to step S2.

[0016] S8. Using the idea of multi-granularity computing units, perform federated learning on the computational scale of the client from coarse to fine granularity, and finally output S cluster center models and N personalized models.

[0017] Advantages of the present invention:

[0018] By improving the parameters and training process of each client that needs to be uploaded to federated learning, the present invention proposes a personalized federated learning method based on multi-granularity computing units. First, at the coarse-grained level, iterative federated soft clustering is used to calculate the cluster dependence of each client to achieve soft clustering of clients. At the medium-grained level, for the high-dependence clusters of each client, the cluster model C k,s and the local model θk , to realize the personalization of the client model. At the fine-grained level, a parameter hierarchy acceleration method is used to reduce the participation of the deep network, and at the same time, a sparse weight activation training method is used to reduce the training overhead of the client and accelerate the implementation of the personalized federated task. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a flowchart of a personalized federated learning method based on multi-granularity computing units according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0021] A personalized federated learning method based on parameter stratification, which groups through the similarity of the update directions of the basic layer parameters to accelerate the convergence speed of the global model. At the same time, the personalized layer parameters are used locally to alleviate the local distribution difference problem, and finally, the personalized model customization of the local client is realized. As Figure 1 shown, the specific steps are as follows:

[0022] S1. Construct a personalized federated learning system including N clients and S cluster centers, and a cluster center model with initialized parameters is set in the server.

[0023] S2. The client downloads all the cluster center models from the server and calculates the dependence of each cluster model, and at the same time fuses each cluster model and uses it as the local model.

[0024] S3. The client trains the weight parameters of the adaptive aggregation module based on the local data, and at the same time adjusts the computing participation of different layers of the neural network in the federated process.

[0025] S4. The client dynamically fuses the local model and the cluster center model through the adaptive aggregation module to obtain the update of the basic layer weights.

[0026] S5. The client adopts a sparse weight activation training method, and by performing sparse convolution in the forward and backward propagations, it accelerates the implementation of personalized federated learning.

[0027] S6. The client uploads all the personalized models to the server for global aggregation.

[0028] S7. Determine whether the federated learning iteration threshold is reached. If so, go to step S8; otherwise, return to step S2.

[0029] S8. Using the idea of multi-granularity computing units, the computing scale of the client is federated from coarse to fine granularity for learning, and finally S cluster center models and N personalized models are output.

[0030] Further, the specific process of the client calculating the cluster model dependency in step S2 includes:

[0031] S21. Assume that the client k ∈ [N] has a local dataset D k , containing |D k | = n k data samples, where n ks data samples are sampled from P s . Usually, u ks is unknown in advance, and the soft clustering algorithm will try to estimate its value during the learning iteration. Here we define as the membership degree of client k to cluster S.

[0032] S22. Next, the client updates the dependency by traversing each data point of itself and selecting the best cluster center. A data point may also be a fusion of multiple data distributions, or there may be an intersection between multiple cluster centers. The contribution of each data point to the cluster center is no longer a simple binary choice (belonging or not belonging), but a continuous score. By accumulating the data scores of each client for each cluster, the belonging probability of the data point in different clusters can be more accurately reflected, which is beneficial to dealing with the overlap and uncertainty of the data distribution. Therefore, in order to better handle this uncertainty of the soft clustering method, the formula is as follows:

[0033]

[0034] where score kj is the evaluation score of a single data point on the cluster center model, represents the cumulative score of client k on cluster center s in the t-th round.

[0035] S23. Assume that the client k ∈ [N] with the local dataset D k has |D k | = n k data points, where n ks data points are sampled from the distribution P s . Therefore, the true risk of the client can be written as the average of the cluster risks:

[0036]

[0037] S24. Secondly, once the server receives the cluster dependencies Next, it will calculate the aggregation weights The calculation formula is defined as follows:

[0038]

[0039] That is to say, when the client has a high dependence on the cluster s, a higher aggregation weight v will be assigned, and vice versa. The introduction of the smoothing parameter σ avoids the situation where the cluster dependence of some clusters is 0. When the training starts, the distribution center may not show a high dependence on any cluster model, that is then the cluster will equally aggregate all client updates. Otherwise, if σ is extremely small, then the weight aggregation of other clients is not affected by it.

[0040] Furthermore, step S3 is an adaptive aggregation step, and the specific process includes:

[0041] S31. At the t-th iteration, the server sends the old global model θ t-1 to client k, and θ t-1 overwrites the old local model to perform local model training, that is In the adaptive aggregation module AAM, we aggregate the cluster model and the client model element-wise instead of using the overwrite method. The aggregation form is as follows:

[0042]

[0043] where ⊙ is the Hadamard product, W k is the aggregation weight of local client k, and C k is the fused cluster model to which client k belongs after soft clustering.

[0044] S32. Since each layer in the personalized model contributes differently: the shallow layer pays more attention to local feature extraction, while the deep layer is used to extract global feature information. Therefore, the client hopes for most of the information in the lower layers of the global model. To reduce the computational overhead, we introduce a hyperparameter p to control the aggregation participation of different neural network layers by applying it to p higher layers and overwriting the parameters in the lower layers for local initialization:

[0045]

[0046] where |θ i | is the number of layers in , and has the same shape as the lower layers in . The elements in are constants 1. The weight has the same shape as the remaining p higher layers.

[0047] S33. Initialize the value of each element of to 1 at the beginning and learn new based on the old in each iteration. To further reduce the computational overhead, we randomly sample s% of the samples of D i at the t-th iteration and denote it as D s,t,k . The client trains

[0048]

[0049] using a gradient-based learning method. Here, η is the learning rate for weight learning. We freeze the other trainable parameters in AAM, including the cluster model and the local model. After local initialization, client i performs local model training.

[0050] S34. After steps S32 and S33, the weight gradient can be obtained: where represents the local loss of client k We can regard the update of W k as an update of

[0051]

[0052] in AAM. The gradient is scaled element-wise at the t-th iteration. Different from local model training (or fine-tuning) that only focuses on local data, the entire update process can perceive the general information in the global model.

[0053] Furthermore, in step S5, a sparse weight activation training method is adopted, and the specific process is as follows:

[0054] S51. Use sparse weights in forward propagation, sparse weights and activations in backward propagation, while the gradients remain dense. Introduce sparsity by retaining the sp proportion of elements with the highest magnitudes in the tensor, that is, the Top-K elements, and setting the remaining elements to zero.

[0055] Top K (X) = {x ∈ X ∣ x ≥ x (K)}

[0056] S52. In a neural network trained using stochastic gradient descent, where the function of the l-th layer that maps the input activation a l-1 to a l is f l :

[0057] al = f l (a l-1 , θ l )

[0058] S53. During the backpropagation process, the gradient δ of the loss of the l-th layer with respect to its output activation al . This gradient is used to calculate the gradients of the loss with respect to its input activation δ al-1 and the weights δ wl , respectively, using the functions G l and H l . Thus, the backpropagation of the l-th layer can be defined as:

[0059] δa l-1 = G l (δa l , θ l )

[0060] δθ l = H l (δa l , a l-1 )

[0061] S54. Although the positions of the Top-K weights mostly overlap among clients, sending only the number of weights corresponding to the sparsity level may be too strict and hinder the natural variation of the weights. Therefore, we introduce a parameter, called the mask ratio rmask, which represents the additional number of weights that the selected clients will choose to send to the server after local training. Let sp be the sparsity level of local training. After the local training process, each selected client sparsifies their model by retaining only the top (1 - sp + rmask) weights, while setting the remaining weights to zero.

[0062] Train personalized models adapted to all data distributions at different computational unit granularities, and at the same time accelerate the computational process of personalized federation. Use personalized methods to obtain models that are more adapted to the local data distribution. The clear problem definition is:

[0063]

[0064] S81. Among them, the dataset D owned by the client device = (D1, D2,..., D n ). Each client participating in the federated training has a different data distribution. Initially, the client needs to provide its own dataset to participate in the federated learning, and different data distributions will also make the optimization directions different. The basic cluster model C = (C1, C2,... C s)。The server first randomly defines S data distribution models and distributes them to all clients. Each client iteratively evaluates the membership degrees of each domain to achieve soft clustering. Sparsify the weight hyperparameter T. In the forward propagation, the K weight elements with the largest absolute values in the weight matrix are selected through Top-K, and the remaining weight elements are set to zero. In the backward propagation, the K activation values with the largest absolute values in the activation vector are selected through Top-K, and the remaining activation values are set to zero.

[0065] S82.S cluster models Through continuous iterative training of the S initial models randomly distributed by the server, the server finally outputs the cluster model C that satisfies the S data distributions. * 。After convergence, the cluster model is beneficial to improving the model performance of clients with lower data volume and can quickly cluster newly added clients. N personalized models After the client adaptively aggregates with the cluster model with high membership degree, N personalized client models θ are finally output to adapt to the data distribution of each client.

[0066] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A personalized federated learning method based on multi-granularity computing units, characterized in that, Including: S1. Construct a personalized federated learning system including N clients and S cluster centers, and a cluster center model after parameter initialization is set in the system server; S2. The clients download all the cluster center models from the server and calculate the dependencies of each cluster center model. At the same time, a cluster model suitable for their own data distribution is fused as the local model; S3. The clients train the weight parameters of the adaptive aggregation module based on local data, and at the same time adjust the computing participation of different layers of the neural network in the federated process; The neural network includes: a personalized layer and a basic layer; S4. The clients dynamically fuse the local model and the cluster center model through the adaptive aggregation module to obtain an update of the basic layer weights; S6. The clients train the local model in a sparse weight activation manner, and accelerate the implementation of personalized federated learning by performing sparse convolution in the forward and backward propagations; S7. The clients upload the local model to the server for global aggregation; S8. Judge whether the federated learning iteration threshold is reached. If so, go to step S9, otherwise return to step S2; S9. Using the idea of multi-granularity computing units, perform federated learning on the computing scale of the clients from coarse to fine granularity, and finally output S cluster center models and N personalized models.

2. The personalized federated learning method based on multi-granularity computing units according to claim 1, wherein The clients download all the cluster center models from the server and calculate the dependencies of each cluster model, including: S21. To calculate the membership degrees of each client to each cluster, it is necessary to define some parameter descriptions: Define that client k ∈ N has a local dataset D k , which contains |D k | = n k data samples, where n ks data samples are sampled from P s (S cluster data distributions), and u ks (the membership degree of client k to cluster S) is unknown in advance, and the soft clustering algorithm will try to estimate its value during the learning iteration; S22. The client k traverses the dataset D k and calculates the score score of each data point on each cluster center model ks and accumulates the data scores of each client for each cluster; Among them, v t ks represents the cumulative score of client k in the t-th round on cluster center s; S23. When the clients calculate the cumulative scores of each data point in the current round, calculate their own cluster dependencies: Among them, u ks represents the cluster dependence degree of client k on cluster S, and v k is the total score of client k.

3. A personalized federated learning method based on a multi-granularity computing unit according to claim 1, characterized in that, Fusing a cluster model suitable for their own data distribution, including: Among them, C k represents a cluster model suitable for its own data distribution, and C s represents the model of the s-th cluster, and u ks represents the cluster dependence degree of client k on cluster S.

4. A personalized federated learning method based on multi-granularity computing units according to claim 1, characterized in that, The clients train the weight parameters of the adaptive aggregation module based on local data, and at the same time adjust the computing participation of different layers of the neural network in the federated process, including: S31. At the t-th iteration, the server sends the old global model θ t-1 to client k, and θ t-1 overwrites the old local model to perform local model training. Meanwhile, in the Adaptive Aggregation Module AAM, the cluster model and the client model are aggregated element-wise, and the aggregation form is as follows: where, ⊙ represents the Hadamard product, and W k represents the aggregation weight of the local client k, and C k represents the fused cluster model to which the client k belongs after soft clustering, represents the model parameters of the client k at the round t, represents the model parameters of the client k in the previous round, represents the fused cluster model to which the client k belonged in the previous round; S32. Since each layer in the clients' deep neural network contributes differently, the personalized layer pays more attention to local feature extraction, while the basic layer is used to extract global feature information. Therefore, the clients hope for most of the information in the personalized layer of the global model. To reduce the computational overhead, a hyperparameter p is introduced to control the aggregation participation of different neural network layers. Apply it to p basic layers and overwrite the parameters in the personalized layer for local initialization: Among them, θ i represents the number of layers in and the shape of is the same as that of the lower layer in The elements in are the constant 1, and the weights have the same shape as the remaining p higher layers.

5. A personalized federated learning method based on a multi-granularity computing unit according to claim 1, characterized in that, The clients dynamically fuse the local model and the cluster center model through the adaptive aggregation module to obtain an update of the basic layer weights, including: At the beginning, initialize the value of each element of to 1, and in each iteration, learn the new based on the old To further reduce the computational overhead, randomly sample s% of the samples of D i at the t-th iteration and denote it as D s,t,k , and the client trains Among them, η represents the learning rate of weight learning, represents the aggregated weight of the p-th layer network of the local client k, represents the gradient of the aggregated weight, represents the model parameters when the client k is at round t, represents the sampling of s% of the local data by the client k at the t-th round, and L represents the loss function; S42: Freeze other trainable parameters in the AAM, including the cluster model and the local model, and perform federated learning after local initialization; S42. To obtain the weight gradient, define: where represents the local loss of client k, and the aggregated weight W of local client k will be updated k is regarded as an update in AAM Among them, the gradient is scaled element-wise in the t-th iteration.

6. The personalized federated learning method based on multi-granularity computing units according to claim 1, wherein The clients adopt a sparse weight activation training method, and accelerate the implementation of personalized federated learning by performing sparse convolution in the forward and backward propagations, including: S51. Use sparse weights in the forward propagation of the clients' neural network training, use sparse weights and activation in the backward propagation, while keeping the gradients dense, retain the sp proportion of elements with the highest magnitudes in the tensor, that is, the Top-K elements, and set the remaining elements to zero to introduce sparsity; Top K X = {x ∈ X | x ≥ x K} Among them, Top K X represents a set composed of the top K elements with the highest amplitudes selected from the tensor X. X represents the input tensor, and x K represents the K-th largest amplitude element in X, and x represents any element in X; S52. In a neural network trained using stochastic gradient descent, where the l-th layer maps the input activation value a l-1 to a l using the function f l to gradually extract features and transmit information for the achievement of the target task: a l = f l a l-1 , θ l Among them, a l represents the output activation value of the l-th layer, and a l-1 represents the output activation value of the (l - 1)-th layer, θ l represents the parameters of the l-th layer, and f l represents the mapping function of the l-th layer; S53. During the backpropagation process, the gradient δ of the loss of the l-th layer with respect to its output activation al , which is used to calculate the gradient of the loss with respect to its input activation δa l-1 and weights δ wl , are calculated using functions G l and H l respectively. Thus, the backpropagation of the l-th layer can be defined as: δa l-1 = G l δa l , θ l δθ l = H l δa l , a l-1 Among them, δa l-1 represents the gradient of the loss L with respect to the input activation value a l-1 at the (l - 1)-th layer, G l represents the function for calculating δa l-1 ; δa l represents the gradient of the loss L with respect to the output activation value a l at the l-th layer; δθ l represents the gradient of the loss L with respect to the parameter θ l at the l-th layer, H l represents the function for calculating δθ l ; Introduce a mask ratio \(r_{mask}\), which represents the additional weight amount that the selected clients will choose to send to the server after local training, so that the clients can still retain the features of the local data as much as possible even when using sparse training; at the same time, let \(s_p\) be the sparsity level of local training. After the local training process, each selected client sparsifies their model by only retaining the top \((1 - s_p + r_{mask})\) weights, while setting the remaining weights to zero.

7. A personalized federated learning method based on a multi-granularity computing unit according to claim 1, characterized in that, Using the idea of multi-granularity computing units, the computing scale of the clients is federated from coarse to fine-grained, and finally \(S\) cluster center models and \(N\) personalized models are output, including: S81. Each client participating in the federated training has a different data distribution. Initially, the clients need to provide their own datasets to participate in the federated learning, and different data distributions will also make the optimization directions different; the server will first randomly define \(S\) data distribution models and send them to all clients. Each client iteratively evaluates the membership degrees of each domain to achieve soft clustering, sparsifies the weight hyperparameter \(T\). In the forward propagation, the top \(K\) weight elements with the largest absolute values in the weight matrix are selected through Top-K, and the remaining weight elements are set to zero. In the backward propagation, the top \(K\) activation values with the largest absolute values in the activation vector are selected through Top-K, and the remaining activation values are set to zero; The server continuously iteratively trains through S initially issued models, and finally outputs the cluster model C that meets S data distributions. * After the client adaptively aggregates with the cluster model with high membership degree, N personalized client models θ are finally output. * to adapt to the data distribution of each client.

Citation Information

Cited By

  • Personalized federal learning method based on dynamic modal routing

    CN121168580A