A Group Personalized Federated Learning Method
By establishing a personalized federated learning method of group collaboration in wireless edge networks, using gradient descent and attention messaging mechanisms, the problems of group-level data heterogeneity and device availability fluctuations are solved, and more efficient model training and generalization are achieved.
Patent Information
- Application Number
- CN202311239858.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-25
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-09-25
AI Technical Summary
Existing personalized federated learning programs are not effective in the face of group-level data heterogeneity and device availability fluctuations, especially in the case of unstable device availability, making it difficult to effectively train global models.
By establishing a three-layer wireless edge network, using groups composed of cloud nodes, edge nodes and devices, personalized federated learning optimization problems are split into gradient descent and near-end operators, and an attention messaging mechanism is introduced to promote pairwise collaboration between groups and train personalized group models.
It improves the generalization and training speed of the model, reduces the impact of device availability fluctuations on model training, and improves the performance of personalized model in wireless edge network scenarios.
Smart Images

Figure CN117313834B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of wireless communication and relates to a group personalized federated learning method. Background Art
[0002] Federated learning (FL) is a very promising distributed machine learning paradigm that can train a global model for all participating devices without directly accessing their private data. However, deploying FL at the edge of wireless networks faces some technical challenges. Among them, there are two very critical challenges: data heterogeneity and device availability.
[0003] Data heterogeneity stems from the variations in device attributes, preferences, and data patterns. Currently, personalized federated learning (PFL) as an FL solution for data heterogeneity has been extensively studied, and there are various PFL solutions that can effectively address the data heterogeneity challenge. However, the current PFL solutions are mainly designed for individual-level heterogeneity, which makes them less effective for group-level heterogeneity. The device availability problem is caused by factors such as connection loss, phone calls, or low battery levels, which may lead to unstable participation of devices during training. However, there is currently no PFL solution that can adapt to the fluctuations in device availability. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a group personalized federated learning method.
[0005] To achieve the above purpose, the present invention provides the following technical solutions:
[0006] A group personalized federated learning method, which includes the following steps:
[0007] S1: Establish a three-layer wireless edge network composed of a cloud node, edge nodes, and devices, where the devices connected to the same edge node form a group;
[0008] S2: Establish a personalized federated learning optimization problem and split it into gradient descent and proximal operator;
[0009] S3: The cloud node collects the latest personalized group models from the edge nodes and executes an attention message passing mechanism equivalent to gradient descent to obtain the intermediate model of the group, and then returns the intermediate model to the corresponding edge node;
[0010] S4: The edge node cooperates with the available clients in the group to solve the proximal operator to obtain the personalized group model;
[0011] S5: Repeat steps S3 and S4 until the convergence condition is reached.
[0012] Optionally, S1 specifically includes the following steps:
[0013] Establish a wireless edge network model including a cloud node, a group of M edge nodes and a group of N devices ; where the edge nodes serve a specific group of devices i.e., group m, where the cloud node is connected to the edge nodes through a backhaul link, and the edge nodes communicate with the devices through a wireless channel; use a binary random variable X m,n ∈ {0, 1} to represent the availability status of device n in group m, where X m,n = 1 indicates available, and X m,n = 0 indicates unavailable; define the probability of device n being available as Use to represent the dataset of device n in group m, and the dataset of group m can be The data distribution of the group dataset is non-IID.
[0014] Optionally, S2 specifically includes the following steps:
[0015] Construct the following personalized federated learning optimization problem P1:
[0016]
[0017] where the optimization variables are all personalized group models λ > 0 is a weighting coefficient; in the objective function, represents the sum of the training losses of all personalized group models , where f n (w m ) represents the loss function of the personalized group model w m on the local dataset ; represents the difference between personalized group models, where φ > 0 is a hyperparameter; where the optimization problem P1 can be split into a gradient descent for optimizing the dissimilarity penalty term and a proximal operator for optimizing the group loss; the optimization problem P1 can be solved iteratively by alternately executing gradient descent and the proximal operator; use to represent the number of iterations, where is the index of the T-th iteration;
[0018] One-step gradient descent for optimizing the dissimilarity penalty term in the t-th iteration:
[0019]
[0020] where, Denote the intermediate model of all groups at the t-th iteration, Denote the personalized group model of all groups at the (t - 1)-th iteration, α t Denote the step size of gradient descent;
[0021] The proximal operator for optimizing the group loss at the t-th iteration:
[0022]
[0023] where, is the personalized group model obtained at the t-th iteration, ||·|| F represents the Frobenius norm;
[0024] Rewrite the proximal operator as a series of sub-problems P2:
[0025]
[0026] where, is the intermediate model of group m at the t-th iteration, is the personalized model of group m obtained at the t-th iteration; Problem P2 can be solved by the edge service cooperating with the available devices in group m.
[0027] Optionally, the S3 specifically includes the following steps:
[0028] S31: The cloud node collects the personalized group models, as follows:
[0029] The cloud node collects the personalized group models from all edge nodes through the feedback network collect the personalized group models
[0030] S32: The cloud node uses the attention message passing mechanism to obtain the intermediate model of the group, as follows:
[0031] One-step gradient descent for minimizing the dissimilarity penalty term can be rewritten as the following attention message mechanism:
[0032]
[0033] where, the coefficient satisfies and is given by the following formula:
[0034]
[0035] where represents the similarity between the models of group m and group i at the (t - 1)-th iteration; the more similar the models of group m and group i are, the coefficient The larger; the coefficient The increase of will promote the pairwise cooperation between group m and group i;
[0036] S33: The cloud node sends the intermediate model of the group to the corresponding edge node, the content is as follows:
[0037] The cloud node sends the latest intermediate model of the group to the corresponding edge node through the feedback network
[0038] Optionally, the S4 specifically includes the following steps:
[0039] S41: The edge node broadcasts the intermediate model of the group, the content is as follows:
[0040] The edge node sends the latest intermediate model to the available devices in the served group m through wireless broadcast;
[0041] S42: The available devices train the model locally, the content is as follows:
[0042] 1) Edge service broadcasts the latest intermediate model of the group to group m
[0043] 2) The available device n in group m, with the intermediate model as the starting point, performs K steps of stochastic gradient descent with a step size of η t The index of each step is k ∈ {1, 2,..., K}; The expression of each step of stochastic gradient descent is as follows:
[0044]
[0045] Among them, is the stochastic gradient operator, represents the local model of client n in group m after k steps of stochastic gradient descent in t iterations,
[0046] 3) The available device n in group m transmits the locally trained model to the edge node m through the wireless channel;
[0047] S43: The edge node aggregates the models uploaded by the available devices and updates the personalized group model, the process is as follows:
[0048] 1) The edge node m receives the models uploaded by the available devices in group m
[0049] 2) The edge node m aggregates the models uploaded by the available devices and updates the personalized group model:
[0050]
[0051] 3) The edge node stores the latest personalized group model locally.
[0052] Optionally, the S5 is specifically:
[0053] Repeat S3 and S4. Each time it is executed, the iteration count t is incremented by 1; if the iteration count t = T, stop the iteration and output
[0054] The beneficial effects of the present invention are as follows:
[0055] 1. The present invention forms a group of devices connected to the same edge node. Within the group, a personalized group model is trained through the federated averaging mechanism, and between groups, the attention message passing mechanism is used to promote paired collaboration of data-similar groups.
[0056] 2. Through the attention message passing mechanism, the present invention realizes paired collaboration between groups with similar models, thereby training the group personalized models simultaneously, improving the generalization of the models and accelerating the training speed.
[0057] 3. The present invention trains a personalized group model for each group through federated averaging to mitigate the impact of device availability fluctuations on model training.
[0058] Other advantages, objectives, and features of the present invention will to some extent be elaborated in the subsequent description, and to some extent, will be obvious to those skilled in the art based on the examination and research of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:
[0060] Figure 1 is the flowchart of the method proposed by the present invention;
[0061] Figure 2 is the system architecture diagram of the method proposed by the present invention;
[0062] Figure 3 is the average test accuracy of the method proposed by the present invention and the comparative method at different device availability rates when the number of fixed devices, device group division, and device data are group non-IID;
[0063] Figure 4The average test accuracy of the method proposed in the present invention and the comparative method under different device availability probabilities when the number of fixed devices, device group division, and device data are independently and identically distributed (hereinafter referred to as IID). Detailed implementation manners
[0064] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0065] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and cannot be understood as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged or reduced, which do not represent the dimensions of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0066] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, it is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and cannot be understood as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific situations.
[0067] The present invention provides a group personalized federated learning method, which is applicable to group non-IID data sets. The flow of the method is as Figure 1 shown;
[0068] Furthermore, Figure 1 The content of step S1 is to establish a wireless edge network system model, a device disconnection model, and a group data model as Figure 2 shown, specifically including the following content:
[0069] Establish a wireless edge network model including a cloud node, a group of M edge nodes and a group of N devices where the edge nodes Serve a specific group of devices That is, group m, where The cloud node is connected to the edge node through a backhaul link, and the edge node communicates with the device through a wireless channel. Use a binary random variable X m,n ∈ {0, 1} to represent the availability status of device n in group m, where X m,n = 1 indicates available, and X m,n = 0 indicates unavailable. Define the probability that device n is available as Use To represent the dataset of device n in group m. The dataset of group m can be The data distribution of the group dataset is non-IID;
[0070] Furthermore, Figure 1 The content of step S2 is to construct a personalized federated learning optimization problem, and the process is as follows:
[0071] 1) Construct the following personalized federated learning optimization problem P1:
[0072]
[0073] Among them, the optimization variable is all personalized group models λ > 0 is the weighting coefficient. In the objective function, Represents the sum of the training losses of all personalized group models Of which f n (w m ) represents the loss function of the personalized group model w m On the local dataset ; Represents the difference between personalized group models, where φ > 0 is a hyperparameter. Among them, the optimization problem P1 can be split into a gradient descent for optimizing the dissimilarity penalty term and a proximal operator for optimizing the group loss. The optimization problem P1 can be solved iteratively by alternately executing gradient descent and the proximal operator. Use To represent the number of iterations, where Is the index of the Tth iteration;
[0074] 2) One-step gradient descent for optimizing the dissimilarity penalty term in the tth iteration:
[0075]
[0076] Among them, Represents the intermediate model of all groups in the tth iteration, Represents the personalized group models of all groups in the (t - 1)th iteration, α trepresents the step size of gradient descent;
[0077] 3) The proximal operator of the optimized group loss in the t-th iteration:
[0078]
[0079] where is the personalized group model obtained in the t-th iteration, ||·|| F represents the Frobenius norm;
[0080] Rewrite the proximal operator as a series of sub-problems P2:
[0081]
[0082] where is the intermediate model of group m in the t-th iteration, is the personalized model of group m obtained in the t-th iteration. The problem P2 can be solved by the edge service cooperating with group m;
[0083] Furthermore, Figure 1 The step S3 specifically includes the following sub-steps:
[0084] C1. The cloud node collects the personalized group model, as follows:
[0085] The cloud node collects the personalized group model from all edge nodes through the feedback network collect the personalized group model
[0086] C2. The cloud node uses the attention message passing mechanism to aggregate the intermediate model, as follows:
[0087] One-step gradient descent for minimizing the dissimilarity penalty term can be rewritten as the following attention message mechanism:
[0088]
[0089] where the coefficient satisfies and is given by the following formula:
[0090]
[0091] where represents the similarity between the models of group m and group i in the (t - 1)-th iteration. The more similar the models of group m and group i are, the coefficient is larger. The increase in the coefficient will promote the pairwise cooperation between group m and group i;
[0092] C3. The cloud node sends the intermediate model of the group to the corresponding edge node, the content is as follows:
[0093] The cloud node sends the obtained intermediate model of the group
[0094] Furthermore, Figure 1 Step S4 includes the following sub-steps:
[0095] D1. The edge node broadcasts the intermediate model of the group, the content is as follows:
[0096] The edge node sends the latest intermediate model of the group to the available devices in the served group m through wireless broadcast;
[0097] D2. The available devices locally train the model, the content is as follows:
[0098] 1) Edge service broadcasts the latest intermediate model to group m
[0099] 2) The available device n in group m, with the intermediate model as the starting point, executes K steps of stochastic gradient descent with a step size of η t The index of each step is k ∈ {1, 2,..., K}. The expression of each step of stochastic gradient descent is as follows:
[0100]
[0101] Where, is the stochastic gradient operator, represents the local model of client n in group m after k stochastic gradients in t iterations,
[0102] 3) The available device n in group m transmits the locally trained model to the edge node m through the wireless channel;
[0103] D3. The edge node aggregates the models uploaded by the available devices and updates the personalized group model, the process is as follows:
[0104] 1) The edge node m receives the models uploaded by the available devices in group m
[0105] 2) The edge node m aggregates the models uploaded by the available devices and updates the personalized group model:
[0106]
[0107] 3) The edge node stores the latest personalized group model locally;
[0108] Furthermore, the content of step S5 is to determine whether the convergence condition is reached, and the process is as follows:
[0109] 1) Repeat steps S3 and S4. Each time it is executed, the iteration count t is incremented by 1;
[0110] 2) If the iteration count t = T, stop the iteration and output
[0111] Furthermore, the present invention conducts simulation analysis in a wireless edge network by using the proposed method;
[0112] As Figures 3 to 4 shown, the present invention gives the relationship between the average test accuracy of the FL system trained using the CIFAR-10 dataset and the iteration count under different schemes and different client availability probabilities. Among them, the individual personalized federated learning scheme is a PFL scheme based on the attention message transmission mechanism for data heterogeneity at the individual level, and the traditional federated learning scheme is the federated average scheme for data homogeneity at the individual level.
[0113] Figure 3 shows the average test accuracy of the proposed group personalized method under different device availability probabilities when the number of devices, device group division, and device data are IID. From Figure 3 the following can be seen:
[0114] 1) When all devices are available, i.e., ρ m,n = 1, the performance of all schemes is similar;
[0115] 2) When the device connection probability ρ m,n decreases, the performance of the individual personalized federated learning scheme drops significantly, but the proposed method of the present invention is hardly affected, and its performance is always close to that of the traditional federated learning scheme, proving that the proposed group personalized method of the present invention is still an effective FL scheme when the device data is IID;
[0116] Figure 4 shows the average test accuracy of the proposed group personalized method under different device availability probabilities when the number of devices, device group division, and group data are group non-IID. From Figure 3 the following can be seen:
[0117] 1) The performance of the individual personalized federated learning scheme is better than that of the traditional federated average scheme, proving the effectiveness of the PFL scheme on non-IID datasets;
[0118] 2) The method proposed in the present invention for group-level heterogeneity is superior to other solutions, including individual personalized federated learning solutions and traditional federated averaging solutions, demonstrating the superiority of the proposed group personalized method on group non-IID data;
[0119] 3) For the individual personalized federated learning solution, when the device availability probability ρ m,n decreases, the performance drops significantly, while the method proposed in the present invention is hardly affected, demonstrating the robustness of the proposed group personalized method when some devices are unavailable;
[0120] Comprehensively Figure 3 and Figure 4 analyzing, the group personalized method proposed in the present invention can effectively improve the performance of personalized models in wireless edge network scenarios.
[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A group personalized federated learning method, characterized in that: The method includes the following steps: S1: Establish a three - layer wireless edge network composed of cloud nodes, edge nodes, and devices. Devices connected to the same edge node form a group, which specifically includes the following: Build a wireless edge network model including a cloud node, a set of M edge nodes and a set of N devices ; among them, the edge nodes serve a specific set of devices c m , that is, group m, where the cloud node is connected to the edge nodes through a fronthaul link, and the edge nodes communicate with the devices through a wireless channel; use a binary random variable X m,n ∈ {0, 1} to represent the availability status of device n in group m, where X m,n = 1 means available, and X m,n = 0 means unavailable; define the probability that device n is available as Use to represent the dataset of device n in group m, and the dataset of group m is The data distribution of the group dataset is non-IID; S2: Establish a personalized federated learning optimization problem and split it into gradient descent and proximal operator, which specifically includes the following steps: Construct the following population - personalized federated learning optimization problem P1: Among them, the optimization variables are all personalized group models λ > 0 is the weighting coefficient; in the objective function, represents the sum of the training losses of all personalized group models where f n (w m ) represents the loss function of the personalized group model w m on the local dataset ; represents the difference between personalized group models, where φ > 0 is a hyperparameter; among them, the optimization problem P1 is split into a gradient descent for optimizing the dissimilarity penalty term and a proximal operator for optimizing the group loss; the optimization problem P1 is iteratively solved by alternately executing the gradient descent and the proximal operator; use to represent the number of iterations, where is the index of the T-th iteration; One - step gradient descent of the optimization dissimilarity penalty term at the t - th iteration: Among them, represents the intermediate model of all groups in the t-th iteration, represents the personalized group model of all groups in the (t - 1)-th iteration, α t represents the step size of gradient descent; Proximal operator of the optimization group loss at the t - th iteration: Among them, is the personalized group model obtained in the t-th iteration, and ‖·‖ F represents the Frobenius norm; Rewrite the proximal operator as a series of sub - problems P2: Among them, is the intermediate model of group m after t iterations, is the personalized model of group m obtained in the t-th iteration; problem P2 is solved by the edge service in cooperation with the available devices in group m; S3: The cloud node collects the latest personalized group models from the edge nodes and executes an attention message - passing mechanism equivalent to gradient descent to obtain the intermediate model of the group, and then returns the intermediate model to the corresponding edge nodes, which specifically includes the following steps: S31: The cloud node collects the personalized group models, as follows: The cloud node collects personalized group models from all edge nodes through the backhaul network S32: The cloud node uses the attention message - passing mechanism to obtain the intermediate model of the group, as follows: One - step gradient descent for minimizing the dissimilarity penalty term is rewritten as the following attention message mechanism: Among them, the coefficient satisfies and is given by the following formula: Among them represents the similarity between the models of group m and group i in the (t - 1)-th iteration; the more similar the models of group m and group i are, the coefficient is larger; the increase of the coefficient will promote the pairwise cooperation between group m and group i; S33: The cloud node sends the group intermediate model to the corresponding edge nodes, as follows: The cloud node will send the obtained group intermediate model to the corresponding edge node through the feedback network S4: The edge node cooperates with the available clients in the group to solve the proximal operator to obtain the personalized group model, which specifically includes the following steps: S41: The edge node broadcasts the intermediate model of the group, as follows: Edge node Send the latest intermediate model To the available devices in the served group m via wireless broadcast; S42: The available devices locally train the model, as follows: 1) Edge service Broadcast the latest intermediate model to group m 2) The available device n in group m starts from the intermediate model and performs K steps of stochastic gradient descent with a step size of η t The index of each step is k ∈ {1, 2, …, K}; the expression for each step of stochastic gradient descent is as follows: Among them, is the stochastic gradient operator, represents the local model of customer n in group m after k stochastic gradients in t iterations, 3) The available device n in group m transmits the locally trained model through a wireless channel to the edge node m; S43: The edge node aggregates the models uploaded by the available devices and updates the personalized group model. The process is as follows: 1) The edge node m receives the models uploaded by the available devices in group m 2) The edge node m aggregates the models uploaded by the available devices and updates the personalized group model: 3) The edge node stores the latest personalized group model locally; S5: Repeat steps S3 and S4 until the convergence condition is reached.
2. The group personalized federated learning method according to claim 1, wherein: The specific content of S5 is: Repeat S3 and S4. Each time they are executed, the iteration count t is incremented by 1. If the iteration count t = T, stop the iteration and output
Citation Information
Patent Citations
Cloud edge network communication optimization method and system based on distributed federated learning
CN115277689A
Federal learning aggregation optimization system and method for power data sharing
CN115358487A