Personalized federal learning method based on parameter decoupling

By decoupling the local model into two parts in federated learning, and using the difference coefficient and dynamic weighting strategy, the poor model performance and knowledge forgetting caused by the data distribution of end-side equipment are solved, and the stability and generalization capabilities of the personalized model are improved.

CN120494135APending Publication Date: 2025-08-15HEBEI UNIV OF TECH
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510606924.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In federated learning, the data distribution differences of end-side devices lead to poor model performance, and existing methods are difficult to effectively take into account individual differences and model generalization capabilities, and there are problems of global knowledge forgetting and personalized model interference.

Method used

The parameter decoupling method is adopted to decouple the local model of the end-side device into global shared model parameters and local personalized model parameters. By calculating the difference coefficient, the personalized model is guided to update it, and dynamic weighted fusion is carried out to avoid global knowledge forgetting and reduce personalized model interference.

Benefits of technology

It improves the performance and stability of the personalized model, enhances the generalization ability of the overall model, effectively solves the personalized learning problem in non-independent and same-distributed data environments, and improves training efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494135A_ABST
    Figure CN120494135A_ABST
Patent Text Reader

Abstract

The invention discloses a personalized federated learning method based on parameter decoupling. The personalized federated learning method comprises the following steps: S1, constructing a federated learning framework including model parameter decoupling and personalized cooperative training; s2, in a model parameter decoupling module, a local model of the end side equipment is decoupled into global shared model parameters and local personalized model parameters; s3, calculating a difference coefficient between the personalized model parameters and the global shared model parameters, wherein the difference coefficient is used for guiding updating of the personalized model; s4, in the personalized collaborative training module, the end-side equipment performs local model training based on the difference coefficient, and performs dynamic weighting according to the similarity with personalized models of other equipment to realize selective fusion of knowledge; s5, the cloud server aggregates the global shared model parameters uploaded by the end side devices, and finally a complete personalized model is formed. According to the method, the performance of the personalized model is improved, the collaborative effect of personalized training is enhanced, and the generalization ability of the whole model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of federated learning technology, and in particular to a personalized federated learning method based on parameter decoupling. Background Art

[0002] In the federated learning framework, participating end-side devices independently collect and process local data sets without sharing them with end-side devices or cloud servers. This can lead to significant differences in the data distribution of different end-side devices, which in turn causes the data to exhibit non-independent and identically distributed characteristics. However, the distribution and amount of data collected by end-side devices may vary significantly, resulting in poor overall performance of the model. The article [Filip Hanzely, Peter Richtarik. Federated Learning of a Mixture of Global and Local Models. [J], Computing Research Repository, 2021.] studies traditional federated learning and points out that in certain specific scenarios, the single global model trained by federated learning deviates significantly from the user's typical usage scenario. This phenomenon highlights the problems faced by federated learning in balancing individual differences and model generalization capabilities. Therefore, how to optimize the global model to adapt to the specific needs of each participant has become an important research direction.

[0003] Multi-task learning is a strategy for improving model performance by simultaneously learning multiple related tasks. Its core concept is to share knowledge learned across tasks and leverage similarities between tasks to improve model generalization. By jointly optimizing the objectives of different tasks, the model can learn more general feature representations from limited data. Federated learning enables multi-task learning across multiple clients, allowing each device to participate in learning different tasks while protecting privacy. Through the federated learning framework, different devices can jointly train a global model while simultaneously performing personalized optimization for their local tasks. This combination not only improves the efficiency of cross-device knowledge sharing but also addresses data privacy concerns, enabling the application of multi-task learning in decentralized data environments. The paper [Smith V, Chiang CK, Sanjabi M, et al. Federated multi-task learning [J]. Advances in neural information processing systems, 2017, 30.] proposes the MOHCA framework, which aims to handle multiple related tasks while ensuring data privacy and security. By distributing the training process for each task across different devices, task-specific model parameters are designed, and the aggregation mechanism of federated learning is leveraged to share useful knowledge between devices. It can effectively promote cross-device knowledge sharing while maintaining personalized learning, thereby improving the performance of multi-task learning. By performing task learning in parallel on different devices, the problem of data heterogeneity is solved. The article [Li T, Hu S, Beirami A, et al. Ditto: Fair and robust federated learning through personalization [C] / / International conference on machine learning. PMLR, 2021: 6357-6368.] proposes a personalized federated learning method called Ditto, which aims to solve the fairness and robustness problems in data heterogeneity networks. The core idea is to train a personalized model for each client through a multi-task learning framework while maintaining consistency with the global model. The specific steps include: the client optimizes a regularized objective function locally, and limits the deviation of the local model by introducing the regularization term of the global model; the server aggregates client updates to maintain global consistency.

[0004] Meta-learning is a learning method that enables models to quickly adapt to new tasks. Its core concept is to extract transferable knowledge or strategies by training on multiple tasks, helping the model learn effectively even with small amounts of data. The paper [Yihan J, Jakub K, JKR, Sreeram K, et al. Improving Federated Learning Personalization Via Model-Agnostic Meta Learning[J], ICLR 2020, 2020.] proposes a personalized federated learning method based on model-agnostic meta-learning (MAML). Its core idea is to use a meta-learning framework to enable each device to quickly adapt to the characteristics of its local data based on a globally shared model. The detailed explanation is as follows: by training a global model that can adapt to different task initializations, each client can quickly fine-tune the model on a small amount of local data, thereby achieving personalized learning results. This method effectively improves the personalization capabilities between devices in federated learning, especially in non-IID data environments, and can better cope with data heterogeneity and uneven distribution. The article [Fallah A, Mokhtari A, Ozdaglar A. Personalized federated learning with theoretical guarantees: A model-agnostic meta-learning approach [J]. Advances in neural information processing systems, 2020, 33: 3557-3568.] proposes the Per-FedAvg algorithm. This method designs an initial shared model, enabling each client to quickly adjust the model based on local data using a small number of gradient descent steps to achieve personalized adaptation. The core idea is to combine federated learning with model-agnostic meta-learning (MAML). After the client updates the model locally, the server aggregates the global initial parameters. The training process involves the client calculating the local gradient and performing a one-step adjustment, followed by uploading the updated parameters to the server for integration.

[0005] Transfer learning is a technique that accelerates the learning of new tasks by borrowing previously learned knowledge. Its core concept is that the knowledge learned by the model on one task can be effectively applied to other related tasks. By migrating the features, model parameters or representations of existing tasks, transfer learning helps the model overcome problems such as insufficient data or task differences, and improves learning efficiency and accuracy. Transfer learning can help the model transfer knowledge between different devices, especially when the data distribution and task characteristics of each device are different. Through transfer learning, federated learning enables each device to share useful knowledge learned from other devices, so that in a Non-IID (Non-Independent and Identically Distributed) data environment, each device can quickly adapt and optimize its local model, thereby improving overall learning effects and task adaptability. The QuPeD algorithm proposed in the article [Ozkara K, Singh N, Data D, et al. Quped: Quantized personalization via distillation with applications to federated learning [J]. Advances in Neural Information Processing Systems, 2021, 34: 3622-3634.] is a quantitative and personalized federated learning method designed to solve the problem of data and resource heterogeneity. Traditional federated learning usually trains a global model, while QuPeD allows clients to learn compressed personalized models with different quantization parameters and structures. The algorithm achieves collaboration between clients through knowledge distillation and learns the quantization model by relaxing the optimization problem, in which the quantization value is also optimized. QuPeD uses alternating proximal gradient updates to solve the compressed personalization problem and analyzes its convergence properties. The article [Wang K, Mathews R, Kiddon C, et al. Federated evaluation of on-device personalization [J]. arXiv preprint arXiv:1910.10252, 2019.] proposes a new evaluation framework that leverages the aggregation mechanism of federated learning combined with personalization strategies to enable devices to adaptively adjust based on local data while transmitting training results back to the cloud for aggregation. This framework emphasizes how to personalize devices in practical applications and improves personalized learning through cross-device collaborative optimization. While these methods all improve the accuracy of local personalized models, they ignore the interference between different on-device personalized models and the forgetting of important global knowledge, resulting in bottlenecks in improving personalized model learning. Summary of the Invention

[0006] The purpose of the present invention is to provide a personalized federated learning method based on parameter decoupling, which includes a federated dual contribution evaluation module and a federated contribution benefit balance algorithm. Among them, the local model of the end-side device is decoupled into two parts: global shared model parameters and local personalized model parameters. During the local training process, each device updates the personalized model parameters separately, and at the same time, calculates the difference between it and the global shared model, generates a difference coefficient, and avoids the loss of global shared knowledge. The end-side device adaptively adjusts the weight according to the similarity between the local personalized model and the personalized model of other devices, reduces the interference of irrelevant personalized models on local training, and improves the stability and training efficiency of the personalized model. Finally, the cloud server aggregates the global shared model parameters uploaded by the end-side devices, updates the global model, and synchronizes it to each device, effectively improving the performance of the personalized model, avoiding the forgetting of global knowledge, and at the same time enhancing the synergistic effect of personalized training and improving the generalization ability of the overall model.

[0007] The technical solution adopted by the present invention is:

[0008] A personalized federated learning method based on parameter decoupling includes the following steps:

[0009] S1: Build a federated learning framework that includes model parameter decoupling and personalized collaborative training;

[0010] S2: In the model parameter decoupling module, the local model of the end-side device is decoupled into two parts: global shared model parameters and local personalized model parameters;

[0011] S3: Calculate the difference coefficient between the personalized model parameters and the global shared model parameters to guide the update of the personalized model and prevent the global shared knowledge from being forgotten;

[0012] S4: In the personalized collaborative training module, the end device trains a local model based on the difference coefficient and dynamically weights it based on the similarity with the personalized models of other devices to achieve selective knowledge fusion.

[0013] S5: The cloud server aggregates the global shared model parameters uploaded by each end-side device, updates the global model and synchronizes it with each device. Each device combines the locally stored personalized parameters to form a complete personalized model.

[0014] Furthermore, in step S2, in the model parameter decoupling module, the local model of the end-side device is decoupled into two parts: a global shared model parameter and a local personalized model parameter, including the following steps:

[0015] 1-1) Assume that there are N end-side devices in the federated learning system, and the model parameter of the i-th device is θi , then parameter decoupling can be expressed as:

[0016] θ i =θ g +Δθ i

[0017] Among them, the parameters are decoupled. g represents the global shared model parameters, Δθ i Represents the personalized model parameters of the i-th device. This decoupling approach allows each device to maintain its own personalized characteristics while participating in federated learning;

[0018] Furthermore, in step S3, the difference coefficient between the personalized model parameters and the global shared model parameters is calculated to guide the update of the personalized model and prevent the global shared knowledge from being forgotten, including the following steps:

[0019] 2-1) To avoid forgetting global shared knowledge during personalized training, a difference coefficient calculation mechanism is introduced. The difference coefficient is used to quantify the difference between the personalized model parameters and the global shared model parameters and guide the update of the personalized model. For the i-th device, its difference coefficient α i The calculation formula is:

[0020]

[0021] Where |·|2 represents the L2 norm and ε is a small constant used to avoid the denominator being zero. i The value range of is [0,1]. The smaller the value, the smaller the difference between the personalized model and the global model, and the larger the value, the greater the difference.

[0022] 2-2) In each round of federated learning, the end-side device updates the model parameters based on local data. To balance personalized training and global knowledge preservation, this paper proposes a personalized model update method based on the difference coefficient. The local training objective function of the i-th device is defined as:

[0023]

[0024] Among them, L i is the loss function on device i, D i is the local dataset of device i, λ is the regularization coefficient, which controls the balance between the personalized model and the global model. The second term is a regularization term based on the difference coefficient, which is used to prevent the personalized model from deviating too much from the global model, thereby avoiding forgetting the global shared knowledge. i It plays the role of adaptive adjustment here: when the personalized model is significantly different from the global model, α iWhen the difference is small, the influence of the regularization term is enhanced, which makes the personalized model closer to the global model. i Small, weakening the influence of the regularization term, allowing the personalized model to adapt more freely to the local data distribution. Use stochastic gradient descent to update the parameters:

[0025]

[0026] Where η is the learning rate and t represents the number of local iterations;

[0027] 2-3) After completing local training, the client device needs to extract the personalized parameters Δθ i , for the next round of learning, the update of personalized parameters is as follows:

[0028]

[0029] This updating method enables personalized parameters to capture device-specific knowledge while avoiding excessive deviation from the global model through the constraint of the coefficient of variation;

[0030] Furthermore, in step S4, in the personalized collaborative training module, the end-side device performs local model training based on the difference coefficient and dynamically weights it according to the similarity with the personalized models of other devices to achieve selective knowledge fusion, including the following steps:

[0031] 3-1) Assume that in a round of federated learning, there are M devices (M≤N) participating in the training. For device i, it is necessary to calculate the similarity between its personalized model and the personalized models of other devices. The similarity calculation formula is:

[0032]

[0033] Where, · represents the vector inner product, s i,j The value range of is [-1,1]. The larger the value, the more similar the personalized models of the two devices are.

[0034] 3-2) Based on similarity, dynamic weighting is performed on the personalized models of other devices. The weight calculation formula is:

[0035]

[0036] Where τ is a temperature parameter that controls the smoothness of the weight distribution. A smaller τ will make the weight distribution more concentrated, while a larger τ will make the weight distribution more uniform.

[0037] 3-3) Using the calculated weights, device i can selectively obtain valuable knowledge from other devices and integrate it into its own personalized model:

[0038]

[0039] Where β∈[0,1] is the fusion coefficient, which controls the balance between the knowledge of the device itself and the knowledge of other devices. A larger β value means more absorption of the knowledge of other devices, while a smaller β value means more retention of the device's own knowledge.

[0040] Furthermore, in step S5, the cloud server aggregates the global shared model parameters uploaded by each end-side device, updates the global model and synchronizes it with each device. Each device forms a complete personalized model in combination with the locally stored personalized parameters:

[0041] 4-1) After each round of federated learning, participating devices complete model training locally and upload the updated model parameters to the cloud for aggregation. Unlike traditional federated learning, which directly aggregates the entire model, in this method, the cloud server only aggregates the globally shared portion, while the personalized portion remains locally on each device. After receiving the model parameters uploaded by each device, the cloud server extracts the globally shared portion and performs a weighted average:

[0042]

[0043] Among them, D i Represents the data volume of device i, which serves as the aggregation weight. This aggregation method takes into account the data distribution of each device, making devices with large data volumes contribute more to the global model.

[0044] 4-2) After the aggregation is completed, the cloud server will update the global model After receiving the new global model, each device combines the personalized parameters stored locally. Form a complete personalized model:

[0045]

[0046] This update method enables devices to acquire global shared knowledge while maintaining personalized features, achieving a balance between knowledge sharing and personalization capabilities.

[0047] The beneficial effects of adopting the above technical solution are:

[0048] The purpose of the present invention is to provide a personalized federated learning method based on parameter decoupling, which includes a federated dual contribution evaluation module and a federated contribution benefit balance algorithm. Among them, the local model of the end-side device is decoupled into two parts: global shared model parameters and local personalized model parameters. During the local training process, each device updates the personalized model parameters separately, and at the same time, calculates the difference between it and the global shared model, generates a difference coefficient, and avoids the loss of global shared knowledge. The end-side device adaptively adjusts the weight according to the similarity between the local personalized model and the personalized model of other devices, reduces the interference of irrelevant personalized models on local training, and improves the stability and training efficiency of the personalized model. Finally, the cloud server aggregates the global shared model parameters uploaded by the end-side devices, updates the global model, and synchronizes it to each device, effectively improving the performance of the personalized model, avoiding the forgetting of global knowledge, and at the same time enhancing the synergistic effect of personalized training and improving the generalization ability of the overall model.

[0049] The proposed personalized method was applied to the public datasets MNIST, EMNIST, and CIFAR-10. Experimental analysis showed that compared to existing methods, the accuracy improved by 0.81%, 2.13%, and 2.33%, respectively. Furthermore, as the degree of non-IID increases, the proposed method demonstrates stronger robustness, with minimal performance degradation. Furthermore, it is able to achieve the target accuracy in fewer communication rounds, further validating the method's ability to optimize personalized learning and global aggregation in non-IID environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 Framework diagram of personalized federated learning method based on parameter decoupling;

[0051] Figure 2 Accuracy curves of different sub-end devices under α=0.5;

[0052] Figure 3 Accuracy curves of different sub-end devices when α=0.3;

[0053] Figure 4 Accuracy curves of different sub-end devices under α=0.1;

[0054] Figure 5 Accuracy curves of each method under IID;

[0055] Figure 6 Accuracy curves of various methods under α=0.5;

[0056] Figure 7 Accuracy curves of various methods under α=0.3;

[0057] Figure 8 Accuracy curves of various methods under α=0.1. DETAILED DESCRIPTION

[0058] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0059] This paper takes federated learning personalization as the background and federated learning architecture as the carrier, and proposes a personalized federated learning method based on parameter decoupling. Figure 1 As shown, the following steps are included:

[0060] S1: Build a federated learning framework that includes model parameter decoupling and personalized collaborative training;

[0061] S2: In the model parameter decoupling module, the local model of the end-side device is decoupled into two parts: global shared model parameters and local personalized model parameters;

[0062] 1-1) Assume that there are N end-side devices in the federated learning system, and the model parameter of the i-th device is θ i , then parameter decoupling can be expressed as:

[0063] θ i =θ g +Δθ i ;

[0064] Among them, the parameters are decoupled. g represents the global shared model parameters, Δθ i Represents the personalized model parameters of the i-th device. This decoupling approach allows each device to maintain its own personalized characteristics while participating in federated learning;

[0065] S3: Calculate the difference coefficient between the personalized model parameters and the global shared model parameters to guide the update of the personalized model and prevent the global shared knowledge from being forgotten;

[0066] 2-1) To avoid forgetting global shared knowledge during personalized training, a difference coefficient calculation mechanism is introduced. The difference coefficient is used to quantify the difference between the personalized model parameters and the global shared model parameters and guide the update of the personalized model. For the i-th device, its difference coefficient α i The calculation formula is:

[0067]

[0068] Where |·|2 represents the L2 norm and ε is a small constant used to avoid the denominator being zero. i The value range of is [0,1]. The smaller the value, the smaller the difference between the personalized model and the global model, and the larger the value, the greater the difference.

[0069] 2-2) In each round of federated learning, the end-side device updates the model parameters based on local data. To balance personalized training and global knowledge preservation, this paper proposes a personalized model update method based on the difference coefficient. The local training objective function of the i-th device is defined as:

[0070]

[0071] Among them, L i is the loss function on device i, D i is the local dataset of device i, λ is the regularization coefficient, which controls the balance between the personalized model and the global model. The second term is a regularization term based on the difference coefficient, which is used to prevent the personalized model from deviating too much from the global model, thereby avoiding forgetting the global shared knowledge. i It plays the role of adaptive adjustment here: when the personalized model is significantly different from the global model, α i When the difference is small, the influence of the regularization term is enhanced, which makes the personalized model closer to the global model. i Small, weakening the influence of the regularization term, allowing the personalized model to adapt more freely to the local data distribution. Use stochastic gradient descent to update the parameters:

[0072]

[0073] Where η is the learning rate and t represents the number of local iterations;

[0074] 2-3) After completing local training, the client device needs to extract the personalized parameters Δθ i , for the next round of learning, the update of personalized parameters is as follows:

[0075]

[0076] This updating method enables personalized parameters to capture device-specific knowledge while avoiding excessive deviation from the global model through the constraint of the coefficient of variation;

[0077] S4: In the personalized collaborative training module, the end device trains a local model based on the difference coefficient and dynamically weights it based on the similarity with the personalized models of other devices to achieve selective knowledge fusion.

[0078] 3-1) Assume that in a round of federated learning, there are M devices (M≤N) participating in the training. For device i, it is necessary to calculate the similarity between its personalized model and the personalized models of other devices. The similarity calculation formula is:

[0079]

[0080] Where, · represents the vector inner product, si,j The value range of is [-1,1]. The larger the value, the more similar the personalized models of the two devices are.

[0081] 3-2) Based on similarity, dynamic weighting is performed on the personalized models of other devices. The weight calculation formula is:

[0082]

[0083] Where τ is a temperature parameter that controls the smoothness of the weight distribution. A smaller τ will make the weight distribution more concentrated, while a larger τ will make the weight distribution more uniform.

[0084] 3-3) Using the calculated weights, device i can selectively obtain valuable knowledge from other devices and integrate it into its own personalized model:

[0085]

[0086] Where β∈[0,1] is the fusion coefficient, which controls the balance between the knowledge of the device itself and the knowledge of other devices. A larger β value means more absorption of the knowledge of other devices, while a smaller β value means more retention of the device's own knowledge.

[0087] S5: The cloud server aggregates the global shared model parameters uploaded by each end-side device, updates the global model and synchronizes it with each device. Each device combines the locally stored personalized parameters to form a complete personalized model.

[0088] 4-1) After each round of federated learning, participating devices complete model training locally and upload the updated model parameters to the cloud for aggregation. Unlike traditional federated learning, which directly aggregates the entire model, in this method, the cloud server only aggregates the globally shared portion, while the personalized portion remains locally on each device. After receiving the model parameters uploaded by each device, the cloud server extracts the globally shared portion and performs a weighted average:

[0089]

[0090] Among them, D i Represents the data volume of device i, which serves as the aggregation weight. This aggregation method takes into account the data distribution of each device, making devices with large data volumes contribute more to the global model.

[0091] 4-2) After the aggregation is completed, the cloud server will update the global model After receiving the new global model, each device combines the personalized parameters stored locally. Form a complete personalized model:

[0092]

[0093] This update method enables the device to acquire global shared knowledge while maintaining personalized characteristics, achieving a balance between knowledge sharing and personalization capabilities.

[0094] Based on the above steps, the present invention effectively solves the problem of poor performance of personalized models of end-side devices under non-independent and identically distributed data, and proposes a personalized federated learning method based on parameter decoupling, which includes a personalized collaborative training algorithm containing a parameter decoupling module and dynamic weighting. Among them, the local model of the end-side device is decoupled into two parts: global shared model parameters and local personalized model parameters. During the local training process, each device updates the personalized model parameters separately, and at the same time, calculates the difference between it and the global shared model and generates a difference coefficient to avoid losing global shared knowledge. The end-side device adaptively adjusts the weight based on the similarity between the local personalized model and the personalized models of other devices, reduces the interference of irrelevant personalized models on local training, and improves the stability and training efficiency of the personalized model. Finally, the cloud server aggregates the global shared model parameters uploaded by the end-side devices, updates the global model, and synchronizes it to each device, effectively improving the performance of the personalized model, avoiding the forgetting of global knowledge, and at the same time enhancing the collaborative effect of personalized training and improving the generalization ability of the overall model.

[0095] This invention is based on experimental verification of a personalized federated learning method based on parameter decoupling:

[0096] 1. Test environment

[0097] The operating system is Windows 11 64-bit, the processor is Intel Core i5-12600KF, the memory is 32G, and the graphics card is NVIDIA RTX 4070. The experimental software is developed using Python 3.11.3, Torch 2.1.0, and Torchvision 0.16.0.

[0098] 2. Test verification

[0099] Experimental results and analysis on three public datasets: MNIST, EMNIST, and CIFAR-10

[0100] (1) Dataset description

[0101] Dataset 1: The MNIST dataset is a handwritten digit recognition dataset. Its input is a 28×28 image, and its output is a class label between 0 and 9. It includes 60,000 training samples and 10,000 test samples.

[0102] Dataset 2: The EMNIST dataset is an extended version of the MNIST handwriting dataset. It contains a wider set of characters, including uppercase letters A-Z, lowercase letters A-Z, and numbers 0-9. It also inputs 28×28 images and outputs 62 class labels.

[0103] Dataset 3: The CIFAR-10 dataset is an image dataset for object recognition. It has 10 classes, each class has 6,000 32×32 color images, including 50,000 training samples and 10,000 test samples.

[0104] Table 1 Experimental dataset information

[0105] Dataset Number of training samples Number of test samples Feature Label quantity MNIST 60000 10000 784 10 EMNIST 697932 116323 784 62 CIFAR-10 50000 10000 1024 10

[0106] (2) Model parameters

[0107] The specific parameters of the CNN model used in this section are shown in Table 2.

[0108] Table 2 Model training related information

[0109] describe set up Activation Function ReLU activation function Optimizer SGD optimizer Learning rate 0.01 momentum 0.5 Number of local updates 5

[0110] (3) Analysis of comparative experimental results

[0111] To verify the accuracy and convergence of the PD-PFL (Personalized federated learning method based on parameter decoupling) method in this paper, we compared it with four other models and set different Non-IID levels (α = 0.5, α = 0.3, α = 0.1). Each dataset was divided into three different levels of Non-IID, and the accuracy of each child was used to show the effect of collaborative training, such as Figure 2 、 Figure 3 、 Figure 4 shown.

[0112] Across three datasets and at three levels of Non-IID partitioning, the PD-PFL method achieved generally higher accuracy on each child-end than the FedAvg, Per-FedAvg, FedBN, and Scaffold methods, demonstrating stronger generalization and personalized adaptability. In particular, at α = 0.1, the strongest Non-IID partitioning level, PD-PFL significantly improved model performance on the vast majority of child-ends, demonstrating its robust personalized learning capabilities even in highly heterogeneous environments.

[0113] In addition, the standard deviation of the model accuracy of each method across 30 sub-clients was compared to measure the volatility of model performance. Experimental results show that the PD-PFL method achieves significantly lower accuracy variance than the comparison methods, indicating more consistent training results across different sub-clients. This stability stems from the parameter decoupling mechanism and dynamic weighting strategy proposed in this paper. The former ensures the preservation of globally shared knowledge during personalized training, while the latter promotes collaboration between sub-clients, mitigates the impact of non-IID data on training results, and achieves collaborative sub-client personalized training. In contrast, the FedAvg and FedBN methods exhibit significant fluctuations in personalized model accuracy when using non-IID data, with some sub-clients even achieving significantly lower-than-average accuracy, indicating their difficulty adapting to significant differences in data distribution across devices. While Per-FedAvg achieves higher accuracy on individual sub-clients, it still exhibits significant overall fluctuation. Scaffold improves stability in some scenarios, but its overall collaboration is still inferior to PD-PFL.

[0114] The accuracy of independent and identically distributed (IID) in the experimental results is shown in Table 3. A detailed analysis and comparison is conducted by comparing the accuracy of IID with different degrees of Non-IID, as well as the convergence speed under different degrees of Non-IID.

[0115] Table 3 Comparison of accuracy of various methods under IID

[0116] method MNIST EMNIST CIFAR-10 FedAvg 0.9623 0.8532 0.6458 Per-FedAvg 0.9682 0.8617 0.6653 FedBN 0.9753 0.8597 0.6715 Scaffold 0.9786 0.8628 0.6762 PD-PFL 0.9814 0.8725 0.6864

[0117] From Table 3 and Figure 5 As can be seen from the figure, under the IID data environment, the PD-PFL method proposed in this paper achieved the best performance on all three datasets. Compared with FedAvg, PD-PFL improved the accuracy by 1.91%, 1.93%, and 4.06% on the MNIST, EMNIST, and CIFAR-10 datasets, respectively. This shows that even when the data is evenly distributed, PD-PFL's parameter decoupling mechanism and personalized collaborative training strategy can effectively improve model performance. Especially on the more complex CIFAR-10 dataset, PD-PFL showed more significant advantages over other personalized federated learning methods (such as Per-FedAvg and FedBN), which proves the effectiveness of this method in handling complex tasks.

[0118] In order to verify the robustness of PD-PFL in different degrees of Non-IID data environments, we continue to compare the performance of various methods under three different Non-IID levels of α = 0.5, α = 0.3 and α = 0.1.

[0119] Table 4 Comparison of the accuracy of various methods under α=0.5

[0120]

[0121]

[0122] From Table 4 and Figure 6 As can be seen, under the non-IID data setting of α = 0.5, the proposed PD-PFL method achieves the best performance on all three datasets. On the MNIST dataset, PD-PFL achieves an accuracy of 87.58%, a 4.91 percentage point improvement over the baseline method FedAvg. On the EMNIST dataset, PD-PFL achieves an accuracy of 75.93%, a 4.68 percentage point improvement over FedAvg. On the more complex CIFAR-10 dataset, the performance gap is even more significant, with PD-PFL achieving an accuracy of 66.42%, a 4.99 percentage point improvement over FedAvg. Notably, PD-PFL maintains a significant advantage even when compared to other personalized federated learning methods, such as outperforming Scaffold by 1.37 percentage points on the MNIST dataset and FedBN by 1.76 percentage points on the EMNIST dataset.

[0123] Table 5 Comparison of the accuracy of various methods under α=0.3

[0124] method MNIST EMNIST CIFAR-10 FedAvg 0.6824 0.5758 0.5039 Per-FedAvg 0.7336 0.6205 0.5426 FedBN 0.7487 0.6307 0.5635 Scaffold 0.7446 0.6425 0.5553 PD-PFL 0.7674 0.6721 0.5938

[0125] From Table 5 and Figure 7 As can be seen, the PD-PFL method's advantage further expands in the more severe non-IID setting (α = 0.3). On the MNIST dataset, PD-PFL achieves an accuracy of 76.74%, an 8.50 percentage point improvement over the baseline method FedAvg. On the EMNIST dataset, PD-PFL achieves an accuracy of 67.21%, a 9.63 percentage point improvement over FedAvg. On the CIFAR-10 dataset, PD-PFL achieves an accuracy of 59.38%, a significant 8.99 percentage point improvement over FedAvg. PD-PFL also maintains its lead compared to other personalized federated learning methods, outperforming the closest FedBN method by 1.87 percentage points on the MNIST dataset and Scaffold by 2.96 percentage points on the EMNIST dataset. These results strongly demonstrate that PD-PFL's parameter decoupling mechanism can more effectively alleviate the non-IID problem when data distribution diversity increases.

[0126] Table 6 Comparison of accuracy of various methods under α=0.1

[0127] method MNIST EMNIST CIFAR-10 FedAvg 0.5703 0.4515 0.3646 Per-FedAvg 0.6436 0.5653 0.4487 FedBN 0.6675 0.5786 0.4528 Scaffold 0.6751 0.5684 0.4615 PD-PFL 0.6982 0.6253 0.5176

[0128] Several important observations emerge from the experimental results: As the degree of Non-IID increases, the performance of all methods decreases to varying degrees. This is expected, as increasing differences in data distribution make model optimization more difficult. PD-PFL performs best under all Non-IID settings, particularly in the highly Non-IID setting (α = 0.1), where it significantly outperforms FedAvg by 10.50%, 11.38%, and 13.30% on the MNIST, EMNIST, and CIFAR-10 datasets, respectively. PD-PFL is the least sensitive to the degree of Non-IID, with performance decreasing significantly less from α = 0.5 to α = 0.1 than the other methods. For example, on the CIFAR-10 dataset, the accuracy of FedAvg, Per-FedAvg, FedBN, and Scaffold dropped by 12.97%, 7.95%, 8.18%, and 8.00%, respectively, while PD-PFL only dropped by 4.66%. This fully demonstrates the robustness of PD-PFL when dealing with highly non-IID data.

[0129] As the value of α decreases (i.e., the degree of Non-IID increases), the performance of the PD-PFL method decreases less, indicating that it has a significant advantage in maintaining the performance of the personalized model on the end. This result verifies that the parameter decoupling mechanism and personalized co-training strategy proposed in this chapter can effectively alleviate the negative impact of the Non-IID data distribution. In particular, in the highly Non-IID environment with α = 0.1, PD-PFL still improves the accuracy by 5.61% on the CIFAR-10 dataset compared to the closest Scaffold method. This significant advantage fully demonstrates the effectiveness of PD-PFL in handling complex Non-IID scenarios.

[0130] In addition to model accuracy, convergence speed is also an important metric for evaluating the practicality of federated learning methods. Tables 7, 8, and 9 show the number of training rounds required for each method to achieve average accuracy at three levels of non-IID.

[0131] Table 7 Training rounds required for each method when α=0.5

[0132] method MNIST (80%) EMNIST (70%) CIFAR-10 (60%) FedAvg 32 34 67 Per-FedAvg 28 25 55 FedBN 25 22 51 Scaffold 27 28 58 PD-PFL 21 20 44

[0133] Table 8 Training rounds required for each method when α=0.3

[0134] method MNIST (65%) EMNIST (55%) CIFAR-10 (50%) FedAvg 30 37 90 Per-FedAvg 23 25 50 FedBN 19 16 37 Scaffold 21 22 45 PD-PFL 17 14 33

[0135] Table 9: Training rounds required for each method under α=0.1

[0136] method MNIST (55%) EMNIST (40%) CIFAR-10 (40%) FedAvg 61 53 Not reached Per-FedAvg 40 33 35 FedBN 38 31 41 Scaffold 34 28 49 PD-PFL 27 23 31

[0137] Convergence speed comparisons reveal the following conclusions: PD-PFL exhibits the fastest convergence speed under various non-IID settings. Under α = 0.5, PD-PFL reduces training rounds by 34.4%, 41.1%, and 35.8% compared to FedAvg on the MNIST, EMNIST, and CIFAR-10 datasets, respectively. PD-PFL's convergence advantage becomes more pronounced with increasing non-IID levels. Under highly non-IID conditions (α = 0.1), PD-PFL reaches the target accuracy more than 2.3 times faster than FedAvg, whereas on the CIFAR-10 dataset, FedAvg fails to converge to the target accuracy. PD-PFL's convergence advantage is particularly pronounced on complex datasets. On the CIFAR-10 dataset, PD-PFL demonstrates a significant convergence speed advantage over other personalized federated learning methods, such as Per-FedAvg and FedBN. Comparing the convergence performance of PD-PFL and Scaffold reveals that PD-PFL not only improves accuracy but also exhibits a significant advantage in convergence speed. This demonstrates that the parameter decoupling mechanism and personalized collaborative training strategy proposed in this chapter not only improve the final performance of the model but also accelerate the model training process.

[0138] Comparative experiments reveal the following: The PD-PFL method achieves optimal performance in all experimental scenarios, regardless of IID or non-IID data environments. This demonstrates the effectiveness of its parameter decoupling and personalized collaborative training strategies. As the degree of non-IID increases, PD-PFL demonstrates greater robustness than other methods, with minimal accuracy degradation, demonstrating its ability to effectively address data distribution variations in real-world applications. The PD-PFL method exhibits significant advantages in convergence speed, particularly in highly non-IID environments, significantly reducing the number of training rounds required to achieve target performance and improving training efficiency. This demonstrates the effectiveness of personalized federated learning methods based on parameter decoupling in addressing the problem of on-device personalized model training in non-IID data environments. Through model parameter decoupling, variance coefficient calculation, and a dynamic weighting strategy, PD-PFL achieves a balance between global knowledge sharing and local personalized needs, effectively improving the performance of personalized models while ensuring model generalization.

[0139] To address the poor performance of personalized models on end-devices under non-IID data, a personalized federated learning method based on parameter decoupling is proposed. This method decouples the local model of the end-device into global shared model parameters and local personalized model parameters. By calculating the difference between the two, a difference coefficient is obtained and used as a reference to avoid forgetting the global shared knowledge during personalized training. In addition, a dynamic weighting strategy is proposed to mitigate the influence of other end-device personalized model parameters. By adjusting the influence of personalized models between different devices and reorganizing the new global shared model parameters, the effectiveness of personalized collaborative training is further improved.

[0140] The above examples of the present invention are described in detail, but the content is only a preferred embodiment of the present invention and should not be considered to limit the scope of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.

Claims

1. A personalized federated learning method based on parameter decoupling, characterized by: Enhancing the learning effect of the personalized model on the end device includes the following steps: S1: Build a federated learning framework that includes model parameter decoupling and personalized collaborative training; S2: In the model parameter decoupling module, the local model of the end-side device is decoupled into two parts: global shared model parameters and local personalized model parameters; S3: Calculate the difference coefficient between the personalized model parameters and the global shared model parameters to guide the update of the personalized model and prevent the global shared knowledge from being forgotten; S4: In the personalized collaborative training module, the end device trains a local model based on the difference coefficient and dynamically weights it based on the similarity with the personalized models of other devices to achieve selective knowledge fusion. S5: The cloud server aggregates the global shared model parameters uploaded by each end-side device, updates the global model and synchronizes it with each device. Each device combines the locally stored personalized parameters to form a complete personalized model.

2. The personalized federated learning method based on parameter decoupling according to claim 1, characterized in that: In step S2, in the model parameter decoupling module, the local model of the end-side device is decoupled into two parts: a global shared model parameter and a local personalized model parameter, including the following steps: 1-1) Assume that there are N end-side devices in the federated learning system, and the model parameter of the i-th device is θ i , then parameter decoupling can be expressed as: i i =θ g +Δθ i ; Among them, parameter decoupling processing, θ g represents the global shared model parameters, Δθ i Represents the personalized model parameters of the i-th device, so that each device maintains its own personalized characteristics while participating in federated learning.

3. The personalized federated learning method based on parameter decoupling according to claim 1, characterized in that: In step S3, the difference coefficient between the personalized model parameters and the global shared model parameters is calculated to guide the update of the personalized model and prevent the global shared knowledge from being forgotten, including the following steps: 2-1) To avoid forgetting global shared knowledge during personalized training, a difference coefficient calculation mechanism is introduced. The difference coefficient is used to quantify the difference between the personalized model parameters and the global shared model parameters and guide the update of the personalized model. For the i-th device, its difference coefficient α i The calculation formula is: Among them, |·|2 represents the L2 norm, ε is a small constant used to avoid the denominator being zero, and the difference coefficient α i The value range of is [0,1]. The smaller the value, the smaller the difference between the personalized model and the global model, and the larger the value, the greater the difference. 2-2) In each round of federated learning, the end-side device updates the model parameters based on local data. A personalized model update method based on the variance coefficient balances personalized training and global knowledge preservation. The local training objective function of the i-th device is defined as: Among them, L i is the loss function on device i, D i is the local dataset of device i, λ is the regularization coefficient, which controls the balance between the personalized model and the global model. The second term is the regularization term based on the difference coefficient, which is used to prevent the personalized model from deviating too much from the global model, thereby avoiding forgetting the global shared knowledge. The difference coefficient α i It plays the role of adaptive adjustment here: when the personalized model is significantly different from the global model, α i When the difference is small, the influence of the regularization term is enhanced, which makes the personalized model closer to the global model. i Small, weakening the impact of the regularization term, allowing the personalized model to adapt more freely to the local data distribution, and using stochastic gradient descent to update the parameters: Where η is the learning rate and t represents the number of local iterations; 2-3) After completing local training, the client device needs to extract the personalized parameters Δθ i , for the next round of learning, the update of personalized parameters is as follows: This updating approach enables personalized parameters to capture device-specific knowledge while avoiding excessive deviation from the global model through the constraint of the coefficient of variation.

4. The personalized federated learning method based on parameter decoupling according to claim 1, characterized in that: In step S4, in the personalized collaborative training module, the end-side device performs local model training based on the difference coefficient and dynamically weights it according to the similarity with the personalized models of other devices to achieve selective knowledge fusion, including the following steps: 3-1) Assume that in a certain round of federated learning, there are M devices (M≤N) participating in the training. For device i, it is necessary to calculate the similarity between its personalized model and the personalized models of other devices. The similarity calculation formula is: Where, · represents the vector inner product, s i,j The value range of is [-1,1]. The larger the value, the more similar the personalized models of the two devices are. 3-2) Based on similarity, the personalized models of other devices are dynamically weighted. The weight calculation formula is: Among them, τ is the temperature parameter that controls the smoothness of the weight distribution. A smaller τ will make the weight distribution more concentrated, and a larger τ will make the weight distribution more uniform. 3-3) Using the calculated weights, device i can selectively obtain valuable knowledge from other devices and integrate it into its own personalized model: Among them, β∈[0,1] is the fusion coefficient, which controls the balance between its own knowledge and the knowledge of other devices. A larger β value means more absorption of knowledge from other devices, and a smaller β value means more retention of its own knowledge.

5. The personalized federated learning method based on parameter decoupling according to claim 1, characterized in that: In step S5, the cloud server aggregates the global shared model parameters uploaded by each end-side device, updates the global model and synchronizes it to each device. Each device forms a complete personalized model based on the locally stored personalized parameters: 4-1) After each round of federated learning, the participating devices complete model training locally and upload the updated model parameters to the cloud for aggregation. Unlike traditional federated learning, which directly aggregates the entire model, in this method, the cloud server only aggregates the globally shared part, while the personalized part is retained locally on each device. After receiving the model parameters uploaded by each device, the cloud server extracts the globally shared part and performs a weighted average: Among them, D i represents the data volume of device i, which serves as the aggregation weight and takes into account the data distribution of each device, so that devices with large data volumes contribute more to the global model; 4-2) After the aggregation is completed, the cloud server will update the global model Send it to all end-side devices. After receiving the new global model, each device combines it with the personalized parameters saved locally. Form a complete personalized model: This allows devices to acquire global shared knowledge while maintaining personalized characteristics.

Citation Information

Cited By

  • Dynamic federal mutual learning method and system for balancing personalization and generalization

    CN121031721A

  • A dynamic federated mutual learning method and system balancing personalization and generalization

    CN121031721B

  • Federal transfer learning driven multi-scene communication parameter optimization system and method

    CN121037870A

  • Federal learning-based human body activity identification method and system for Internet of Things equipment

    CN121167329A

  • Personalized federal learning method and system oriented to non-independent identically distributed data

    CN121766394A