Personalized Federated Learning Method and System Based on Parameter Similarity
By calculating the position similarity of key parameters of the client model on the server side and building a similar client set, the weighted average aggregation of key parameters of the personalized model in federated learning is achieved, which solves the problem of poor aggregation effect of personalized model in the existing technology, and improves the reliability and accuracy of the model.
Patent Information
- Application Number
- CN202411424047.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-12
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-10-12
AI Technical Summary
The existing federated learning methods have poor aggregation of personalized models under non-IID data distribution, resulting in a lack of reliability, credibility and accuracy of personalized models.
By calculating the position similarity of key parameters of the client model on the server side, building a similar client set, and calculating a personalized aggregation matrix based on the cosine similarity of the similar client model, the weighted average aggregation of key parameters is achieved.
Accurate personalized aggregation of key parameters is realized, so that the personalized model can better adapt to local data distribution and improve the reliability and accuracy of the model.
Smart Images

Figure CN119294559B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a personalized federated learning method and system based on parameter similarity. Background Art
[0002] Federated learning (FL) is a distributed machine learning framework that allows multiple participants to collaborate in training a model without directly sharing data, effectively protecting user privacy and data security. It has been widely applied in fields such as the Internet of Things, mobile devices, and healthcare.
[0003] As a key research in federated learning, FedAvg can perform multiple local iterations before each client uploads model parameters, greatly improving the efficiency of the entire FL process compared to traditional distributed model training. However, a key problem with the FedAvg method is that it assumes that the data distributions of all clients are independently and identically distributed (IID), and the global model should be suitable for all clients. However, in real-world application scenarios, data is usually non-IID. In this case, the global model may not be able to adapt well to the specific situation of each client. Considering that in most cases, different clients do not need to jointly use a global model, and only need each client to have a personalized model that can perform well under its own data distribution. Therefore, some research works have studied personalized federated learning methods under non-IID data distributions.
[0004] To preserve the expressive effect of the model parameters of each client, FedALA proposes to use learnable weights to dynamically weighted aggregate the model parameter values and personalized model parameter values for some layers of the network; to extract important knowledge from existing models to stitch out personalized models, PFedHR uses clustering to generate stitching models that are similar to the performance of each client model on the public dataset, and uses the stitching models to guide the subsequent training of each client model. However, the implementation complexity of this method is relatively high; to achieve the aggregation of personalized models, FedGraph determines the similarity between different client models by analyzing the expressive ability of model parameters. It constructs a graph matrix to represent the weights of personalized aggregation for each client, and constructs an optimization target using the similarity information of each client model to further obtain the personalized aggregation weights of the clients on the server, so that clients with high similarity will obtain larger aggregation weights, thus realizing the generation of personalized models. However, it does not distinguish between the aggregation of key parameters and non-key parameters, resulting in inaccurate and unreasonable aggregation of key parameters, making the personalized model lack reliability, credibility, and accuracy. To better aggregate key parameters, FedCAC determines the set of key parameters according to the sensitivity of the parameters and further determines similar clients. For key parameters, it aggregates them using similar clients, and for non-key parameters, it uses conventional weighted average aggregation. However, this method does not consider the differences between clients in the aggregation of key parameters, so its personalized effect is still not satisfactory. Summary of the Invention
[0005] The object of the present invention is to provide a personalized federated learning method and system based on parameter similarity to solve the problems existing in the above-mentioned prior art.
[0006] The personalized federated learning method based on parameter similarity in the present invention includes the following steps:
[0007] S1. Each client collects data locally and constructs a local dataset. Each client independently trains a local model using its dataset locally, and uses a proportional function related to the number of iterations to weighted accumulate the parameter criticality value of each iteration;
[0008] S2. Set the proportion of key parameters to obtain a binary key parameter mask matrix;
[0009] S3. Each client uploads the trained model parameters and the binary mask matrix to the server;
[0010] S4. The server calculates the l-based 1The overlap rate of the norms is used to obtain the similarity matrix of the positions of the key parameters, where each element represents the similarity degree between the two models at the corresponding position; based on the similarity matrix of the key parameter positions, the server selects several clients that are most similar to the parameter distribution of the client model to construct a set of similar clients;
[0011] S5. The server calculates the personalized aggregation matrix according to the cosine similarity between the key parameters of the similar client models, and performs weighted averaging of the key parameters of the similar client models of each client to guide the personalized aggregation of the key parameters;
[0012] S6. The server downloads the personalized model parameters of each aggregated client to the corresponding client, and repeats steps S1 to S5 until a predetermined number of communication rounds is reached.
[0013] The personalized federated learning system based on parameter similarity described in the present invention includes a server and a plurality of clients communicatively connected to the server, and uses the method for federated learning.
[0014] The advantages of the personalized federated learning method and system based on parameter similarity described in the present invention are as follows:
[0015] 1. By the training situation of each client model in each round of training, the key parameter distribution is determined, and similar clients are determined according to the key parameter distribution similarity, so that the similar clients have similar training degrees, and the aggregation of key parameters is realized. According to the expression similarity of the key parameters of each similar client, the aggregation ratio of the key parameters is further determined to realize the precise personalized aggregation of the key parameters.
[0016] 2. The importance value of each parameter is determined according to the number of iterations in a single round, and the position of the key parameter is further determined. This key parameter determination principle enables the client model to assign a smaller weight to the sensitivity value of the parameter at a single iteration time point when the model has a large deviation from the data distribution in the early stage of a single round of training, and a higher weight when the client model fits well in the later stage of training, improving the accuracy of determining the position of the key parameter.
[0017] 3. Based on the set of similar clients of each client, an optimization objective for the personalized aggregation weight of the key parameters is constructed according to the similarity of the key parameters of the client and the data volume ratio, realizing the personalized aggregation of the key parameters based on the parameter expression effect, so that the personalized aggregation model better fits the local data distribution. Description of the Drawings
[0018] Figure 1 is a schematic diagram of the training process of the model described in the present invention.
[0019] Figure 2It is a schematic diagram of personalized aggregation of key parameters of the model described in the present invention. Detailed implementation manners
[0020] Federated learning has a wide range of application scenarios, especially in fields with high requirements for data privacy, security, and distributed computing. In the field of healthcare, by training models locally in each hospital and only sharing model parameters instead of data, patient privacy is protected while the generalization ability of the model is improved, such as disease prediction, medical image analysis, etc. In the field of financial services, banks and financial institutions are also restricted by data privacy and regulations and cannot freely share customer information. Federated learning can help different financial institutions collaborate to jointly improve the performance of credit scoring, fraud detection, and risk management models without exposing customer data. In the field of autonomous driving, autonomous vehicles need to learn driving behaviors from a large amount of sensor data. The data collected locally by each vehicle is highly private. Through federated learning, automobile manufacturers can collaboratively train models without sharing specific data, improving the accuracy of vehicle perception and decision-making. In the field of education, educational institutions can use federated learning to build personalized learning systems and improve the learning effects of students. Without directly sharing student data, the model can use data from different regions and schools to build personalized teaching models.
[0021] As Figure 1 and Figure 2 shown, the personalized federated learning method based on parameter similarity described in the present invention includes the following steps:
[0022] S1. Each client collects data locally and constructs a local dataset. Each client independently trains a local model using its dataset locally and cumulatively weights the parameter criticality values of each iteration using a proportional function regarding the number of iterations. The specific steps are as follows:
[0023] S1.1. Each client uses the local dataset to train the model and records the gradient of each iteration and the parameter amplitude before the iteration.
[0024] S1.2. Each client calculates the sensitivity value of each parameter according to the gradient and parameter amplitude before this iteration. The parameter sensitivity matrix is:
[0025]
[0026] where ⊙ represents element-wise multiplication, t represents the round when the server completes the aggregation operation, k represents the number of iterations in each round, i represents the client number, θ represents the model parameter, and each element in the parameter sensitivity matrix reflects the training situation of the corresponding parameter before the k-th iteration.
[0027] S1.3. Calculate the weighted coefficient of the sensitivity value obtained in each iteration according to the number of training iterations.
[0028] S1.4. Cumulate the parameter sensitivity matrix according to the weighted coefficient to obtain the parameter criticality matrix. Since in the early stage of training after each model deployment, the deviation of each client model on the local dataset is relatively serious, while the fitting situation is better in the later stage, the sensitivity value in the later stage can better reflect the importance of the parameters. The parameter criticality matrix obtained by cumulative with the increasing weighted coefficient is as follows:
[0029]
[0030] This matrix reflects the training situation of each parameter of each client model in the current round.
[0031] S2. Set the hyperparameter τ = 0.5 to represent the proportion of critical parameters. According to Take the parameters with the largest importance value in the top τ and mark them as 1, and mark the rest as 0 to obtain the binary critical parameter mask matrix.
[0032] S3. Each client uploads the trained model parameters. And the binary mask matrix indicating the positions of the critical parameters. To the server.
[0033] S4. The server calculates the overlap rate based on the l 1 norm between all client critical mask matrices to obtain the similarity matrix of critical parameter positions, where each element represents the similarity degree between two models at the corresponding position. Based on the similarity matrix of critical parameter positions, the server selects several clients with the most similar model parameter distributions to each client to construct a set of similar clients. The specific steps are as follows:
[0034] S4.1. The server calculates the critical parameter similarity matrix between each client according to the critical parameter mask matrix. The similarity calculation between the i-th client and the j-th client is as follows:
[0035]
[0036] Here, ‖.‖ 1 represents the l 1 norm, and m represents the total number of parameters.
[0037] S4.2. According to the critical parameter similarity matrix O t , calculate the similarity threshold in the current round:
[0038]
[0039] where β is an adjustable hyperparameter, β ∈ (0, 1), and N represents the number of clients. This threshold can adjust its size according to the rounds, so that as the training converges, the constraint on the judgment of similar clients is reduced, thereby enhancing the transfer of knowledge between different clients.
[0040] S4.3. Divide similar clients according to Thre t to construct a set of similar clients:
[0041]
[0042] S5. The server calculates the personalized aggregation matrix based on the cosine similarity between the key parameters of the similar client models, and performs weighted averaging of the key parameters of the similar client models of each client to guide the personalized aggregation of the key parameters. The specific steps are as follows:
[0043] S5.1. The server constructs a corresponding personalized key parameter aggregation weight optimization problem according to the set of similar clients of each client:
[0044]
[0045] where sim(·,·) represents the cosine similarity function, is the weight for aggregating the key parameters of the i-th client using the j-th client, where n j is the size of the dataset of the j-th client, and α is a predefined hyperparameter used to balance the influence of the dataset size and model similarity in the calculation of the collaboration graph weight.
[0046] S5.2. Solve this quadratic programming problem to obtain the personalized aggregation weights of the key parameters for aggregating the key parameters, and the remaining parameters are aggregated using averaging. The aggregation process is expressed as:
[0047]
[0048] where J represents a matrix of all 1s.
[0049] S6. The server downloads the personalized model parameters of each aggregated client to the corresponding client, and repeats steps S1 to S5 until a predetermined number of communication rounds is reached.
[0050] The personalized federated learning system described in the present invention includes a server and a plurality of clients communicatively connected to the server, and uses the method for federated learning.
[0051] Those skilled in the art can make various corresponding changes and deformations according to the technical solutions and concepts described above, and all such changes and deformations should fall within the protection scope of the claims of the present invention.
Claims
1. A personalized federated learning method based on parameter similarity, characterized in that: The following steps are involved: S1. Each client collects data locally and builds a local data set. Each client uses its data set to train a local model independently locally, and uses a proportional function to the number of iterations to weight and accumulate the critical value of the parameters of each iteration; S2. Set the proportion of key parameters to obtain a binary key parameter mask matrix; S3. Each client uploads the trained model parameters and the binary key parameter mask matrix to the server; S4. The server calculates the overlap rate between the key mask matrices of all clients based on the l1 norm, and obtains the similarity matrix of the key parameter position, in which each element represents the similarity degree between the two models represented by the corresponding position; Based on the key parameter position similarity matrix, the server constructs a similar client set for each selected client whose model parameter distribution is most similar to that of the client; S5. The server calculates a personalized aggregation matrix based on the cosine similarity between key parameters of similar client models, and performs a weighted average of key parameters of similar client models of each client to guide personalized aggregation of key parameters; S6. The server transmits the aggregated personalized model parameters of each client to the corresponding client, and repeats steps S1 to S5 until a predetermined number of communication rounds is reached.
2. The personalized federated learning method based on parameter similarity according to claim 1, characterized in that: The step S1 is specifically as follows: The specific steps are: S1.
1. Each client uses the local data set to train the model and records the gradient of each iteration and the parameter amplitude before the iteration; S1.
2. Each client calculates the sensitivity value of each parameter based on the gradient and parameter amplitude before the iteration; the parameter sensitivity matrix is: Where ⊙ represents the multiplication of corresponding elements, t represents the round in which the server completes the aggregation operation, k represents the number of iterations in each round, i represents the client sequence number, θ represents the model parameters, and the parameter sensitivity matrix Each element in reflects the training status of the corresponding parameter before the kth iteration; S1.
3. Calculate the weighted coefficient of the sensitivity value obtained in each iteration according to the number of training iterations Where K i represents the total number of iterations in a single round of the i-th client; S1.4.Accumulate the parameter sensitivity matrix according to the weighting coefficient to obtain the parameter criticality matrix; since the deviation of each client model on the local data set is more serious in the early stage of training after each model is decentralized, and the fitting is better in the later stage, the sensitivity value in the later stage can better reflect the importance of the parameter; the parameter criticality matrix obtained by accumulating the increasing weighting coefficient is as follows: This matrix reflects the training status of each parameter of each client model in the current round.
3. The personalized federated learning method based on parameter similarity according to claim 2, characterized in that: The step S2 is specifically as follows: setting the hyperparameter τ=0.5 to represent the proportion of the key parameters, according to Take the parameter with the largest importance value before τ and mark it as 1, and mark the rest as 0 to get the binary key parameter mask matrix 4. The personalized federated learning method based on parameter similarity according to claim 3, characterized in that: The step S4 is specifically as follows: S4.
1. The server masks the matrix according to the key parameters Calculate the key parameter similarity matrix between each client, where the similarity between the i-th client and the j-th client is calculated as follows: ‖.‖1 represents the l1 norm, and m represents the total number of parameters; S4.
2. Based on the key parameter similarity matrix O t , calculate the similarity threshold for the current round: in β is an adjustable hyperparameter, β∈(0,1); S4.
3. Based on Thr t , divide similar clients and build a similar client set:
5. The personalized federated learning method based on parameter similarity according to claim 4, characterized in that: The specific steps of step S5 are: S5.
1. The server constructs the corresponding personalized key parameter aggregation weight optimization problem based on the similarity client set of each client: Among them, sim(·,·) represents the cosine similarity function, The weight of the key parameter aggregation of the i-th client using the j-th client, where n j is the dataset size of the jth client, and α is a predefined hyperparameter used to balance the impact of dataset size and model similarity in the calculation of collaboration graph weights; S5.
2. Solve the quadratic programming problem to obtain personalized aggregation weights of key parameters The aggregation is used for key parameters, and the remaining parameters are averaged. The aggregation process is expressed as: Where J represents the all-1 matrix and N represents the number of clients.
6. A personalized federated learning system based on parameter similarity, comprising a server and a plurality of clients communicating with the server, characterized in that: Federated learning is performed using the method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Federal learning local model parameter aggregation method
CN115021905A
Federated learning anti-reasoning attack privacy protection method based on double perturbation
CN115481431A