Individualized federal learning method and system based on base model decomposition
By adopting a personalized federated learning method based on base model decomposition in federated learning, using coefficient vectors for training and transmission, the cost of distribution offset and transmission is solved, and efficient personalized model training and transmission is achieved.
Patent Information
- Application Number
- CN202510240246.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-20
AI Technical Summary
In federated learning, the costly problems of distribution offsets and transmission limit their application in complex environments.
Using a personalized federated learning method based on base model decomposition, the parameters of each layer of the model are represented as a linear combination of a set of base models, and coefficient vectors are used for training, aggregation and distribution, replacing the direct aggregation and distribution of the original model parameters.
It improves the information transmission efficiency of federated learning, realizes efficient training and transmission of personalized models, and is suitable for horizontal federated learning scenarios where model isomorphic and data distribution is heterogeneous.
Smart Images

Figure CN120180301A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of federated learning, and in particular to a personalized federated learning method and device based on base model decomposition. Background Art
[0002] With the wide application of Deep Learning in multiple fields, its powerful performance has been proven to play an important role in natural language processing, object detection, image recognition, and medical health. Deep learning can automatically extract useful features from high-dimensional data through its powerful representation and learning capabilities, greatly reducing the workload of manually designing and extracting features, and thus training end-to-end and high-precision machine learning models. However, deep learning usually relies on a large amount of data during the training process, which means that data collectors need to pay huge costs for data acquisition and storage.
[0003] As data privacy protection has increasingly become the focus of global attention, it has become more and more difficult to centrally obtain a large amount of data or directly transmit data. How to utilize distributed data to integrate its valuable information and build a more powerful deep learning model while ensuring data privacy has become a new research hotspot. In such an environment, researchers have proposed Federated Learning to address these challenges.
[0004] Federated learning encrypts and transmits model parameters through encryption algorithms such as homomorphic encryption, rather than directly transmitting data, ensuring that the data remains local, thereby effectively protecting data privacy. Federated learning has become an emerging learning method and has received the attention of many scholars and researchers. For example, Patent CN118917441B proposes a federated learning method for edge heterogeneous environments. For edge heterogeneous scenarios, through screening and detection, federated aggregation is performed on clients that meet the standards; Patent CN118842654B proposes a federated model aggregation encryption method and system, focusing on the encryption blinding of model parameters to ensure the security of private data without increasing resource overhead; Patent CN118798326B proposes a transformer fault diagnosis method, terminal, and medium based on personalized federated learning, disclosing a transformer fault diagnosis method, terminal, and medium based on personalized federated learning, and applying federated learning to the technical field of transformer fault diagnosis.
[0005] Academia usually classifies federated learning into three main types according to the applicable scenarios: horizontal federation, vertical federation, and transfer federation. This application mainly focuses on the horizontal federation field, which is more commonly used in academic research and practical applications. In particular, this application mainly focuses on two main problems faced in federated learning: distribution shift and high transmission cost. These two problems limit the further application of federated learning in complex environments.
[0006] First of all, a significant challenge in federated learning is distribution shift, that is, the difference in data distribution among different clients. In traditional centralized learning, data usually comes from the same source and has a consistent distribution, and the model can be effectively trained and generalized to new data. However, in federated learning, due to the significant differences in the data sources and distributions of each participating party, this distribution shift causes the global model to perform inconsistently on different clients and cannot well adapt to the specific data characteristics of each client. This makes it necessary to customize personalized models for each client, and existing federated learning methods often cannot effectively meet this personalized need.
[0007] Secondly, the high transmission cost is also a major problem in federated learning. The core of federated learning is to transmit the model updates of each client to the central server for aggregation, and this process requires frequent data transmission. In traditional federated learning methods, a large number of model parameters or gradients need to be uploaded after each client training. Especially when the model scale is large, the transmission cost will be very high. Especially in a low-bandwidth network environment, the transmission efficiency is low, resulting in a slow and costly training process. This problem has great challenges in real-life applications, especially in the fields of healthcare and finance.
[0008] Although a variety of federated methods have been applied in reality, there are still some deficiencies in terms of generality, pertinence, and applicability:
[0009] 1) Distribution shift: The data distribution differences among different clients are large, and personalized models need to be customized to improve accuracy. However, traditional federated learning methods often cannot effectively handle data heterogeneity, resulting in a decline in model performance.
[0010] 2) Transmission cost: Transmitting a large number of model parameters between the client and the server leads to high communication costs. Especially in an environment with limited transmission bandwidth, existing methods are difficult to meet the needs of practical applications.
[0011] Therefore, to solve these problems, there is an urgent need to design a new federated learning method that can effectively address the issues of distribution shift and high transmission cost, thereby promoting the application and development of federated learning in real-world complex scenarios. Summary of the Invention
[0012] Aiming at the above problems existing in the prior art, the purpose of the present invention is to provide an efficient personalized federated learning method based on base model decomposition that can effectively solve the problems of low communication efficiency and poor personalization ability in the horizontal federated scenario.
[0013] Another object of the present invention is to provide a personalized federated learning system based on base model decomposition.
[0014] To solve the above problems, the present invention adopts the following technical solutions: A personalized federated learning method based on base model decomposition, comprising the steps of:
[0015] Step 1, federated preparation: Each client obtains the collected original data, and at the same time, the server distributes the common base model to the client, and the client determines the coefficient vector of the corresponding layer and the number of base models to be used according to the number of parameters of each layer of the model;
[0016] Step 2, local coefficient vector training: Train the coefficient vector of each local layer;
[0017] Step 3, coefficient vector aggregation and distribution: Transmit the coefficient vectors of each client to the server, perform average aggregation on the coefficient vectors at the server, and then return the aggregated vectors to each client;
[0018] Step 4, repeat Step 2 and Step 3 until convergence or the maximum iteration period is reached;
[0019] Step 5, end.
[0020] In some embodiments, the federated preparation in Step 1 specifically includes the steps of:
[0021] (11) Each client obtains the collected original data D i ;
[0022] (12) At the same time, the server distributes the common base model to the client;
[0023] (13) The client initializes the local base model, and determines the coefficient vector of the corresponding layer and the number of base models required according to the number of parameters of each layer of the model.
[0024] In some embodiments, the local coefficient vector training in Step 2 specifically comprises the steps of:
[0025] (21) Problem description: Combining the data information of all clients, training a model f for each client i ,
[0026] Suppose there are N different clients, denoted as {C1, C2,..., C N}, and the data corresponding to each federation is {D1, D2,..., D N}, the data distributions are different, and at the same time, the spaces where the original data is located are the same, and the labels of the data are in the same space, that is, X i = X j , Y i = Yj , where \(i\neq j\); among which \(X\) i and \(X\) j represent the feature vectors of two different samples, and \(Y\) i and \(Y\) j represent the labels corresponding to the said samples; the data of each federated end is divided into three parts, specifically training data validation data test data There are the total number of samples \(n\) of class \(i\) i equal to the sum of the number of samples in its training set \(n\) i tr , validation set \(n\) i val and test set \(n\) i te ; the complete data set \(D\) of class \(i\) is i the union of its training set \(D\) i tr , validation set \(D\) i val and test set \(D\) i te ; combining the data information of all clients, training a model \(f\) for each client i , and finding the objective function, where \(l\) is the objective loss function;
[0027]
[0028] (22) Local coefficient vector training
[0029] Each network layer of the client model is represented as a linear combination of a group of randomly initialized base models, and any client model is represented by adjusting the relevant coefficient vectors;
[0030] Denote the parameters of the model \(f\) as \(W\), the model has \(L\) layers, and \(W\) (l) represents the parameters of the \(l\)-th layer of the model. Assume that there are \(m\) l base models for the \(l\)-th layer, and the corresponding parameters are Then, \(W\) (l) is represented as,
[0031]
[0032] where represents the coefficient of the \(i\)-th base model corresponding to the \(l\)-th layer. By optimizing instead of directly optimizing \(W\) ι , the parameters of the network are represented as,
[0033] \(W = \{W\) (1) , \(W\) (2) , …, \(W\) (L)}.
[0034] Each W (l) is represented by ;
[0035] Assume the model is approximately orthogonal, that is
[0036]
[0037] The parameters of the k-th client are represented as follows
[0038]
[0039] The overall parameters are represented as
[0040] W k ={W (1,k) , W (2,k) , …, W (L,k)}.
[0041] Use numerical values to represent the entire network; m << d, where d represents the original parameter dimension
[0042] Assume W *k is the optimal parameter of the k-th client, and the optimization for the k-th client is expressed as
[0043]
[0044] Utilize the optimization of the coefficient vector α k to replace the optimization of the original model parameters, reducing the parameter dimension of the optimization
[0045] Assume the loss function is l. For the k-th client, there is the following objective
[0046]
[0047] where y i is the corresponding label. According to the above formula, optimize each α k .
[0048] In some embodiments, the coefficient vector aggregation and distribution described in step three specifically includes
[0049] (31) The client transfers the coefficient vector to the server. The server aggregates the coefficient vectors from all clients, calculates the average value, and determines the new global coefficient vector for each layer
[0050]
[0051] where α i(l,k) Denoted as the \(i\)-th parameter or eigenvalue of the \(k\)-th local model in the \(l\)-th layer, \(\alpha\) i (l , \(g\) lobal) Denoted as the \(i\)-th parameter or eigenvalue of the global model in the \(l\)-th layer, \(N\) represents the total number of models, \(L\) represents the total number of layers of the model, \(m\) l Denotes the number of parameters or features of the \(l\)-th layer, \(i\) represents the index of the specific parameter position within a layer, and \(l\) represents the layer number of the model.
[0052] In some embodiments, each client locally maintains the scalar weights of the final linear layer without sharing them with the server. Specifically, it includes the steps of:
[0053] (32) During the federated learning process, at the end of each training cycle, each client uploads the coefficient vector \(\alpha\) shared by its previous layers (l,k) to the server, excluding the coefficient vectors of the final layer and the BN layer; the server aggregates the coefficient vectors from all clients, calculates their average, and updates the global coefficient vector of each layer, excluding the BN layer and the last layer, that is
[0054]
[0055] In some embodiments, step three further includes:
[0056] (33) After obtaining the global coefficient vector , distribute it to each client to update the coefficient vector of the corresponding layer.
[0057] Another object of the present application is to provide a personalized federated learning system based on base model decomposition, characterized in that
[0058] It includes clients for obtaining the collected raw data. At the same time, the server distributes the common base model to the clients. The clients determine the coefficient vector of the corresponding layer and the number of base models to be used according to the number of parameters of each layer of the model;
[0059] A base model decomposition-based representation module: used to represent the parameters of each layer of the model as a linear combination of a group of base models, use the coefficient vector of the linear combination of the base models as the representation of the model, and train the local coefficient vector to train the coefficient vector of each local layer;
[0060] A base model decomposition-based training module: trains the coefficient vector of the base model to replace the training of all parameters of the entire model;
[0061] Personalized Federated Aggregation Module Based on Layer Substrate Model: It is used to maintain the partial parameters of the model of local clients and does not upload the batch normalization layer parameters and the coefficient vector of the last layer of the aggregated local model.
[0062] In some embodiments, for the client, according to the number of parameters of each layer of the model, the coefficient vector of the corresponding layer and the number of substrate models to be used are determined specifically as follows:
[0063] Each client prepares its own data D i ;
[0064] Each client divides the data into training data verification data test data
[0065] The server distributes the substrate model to each client.
[0066] Each client initializes local parameters, the coefficient vector α (l,k) .
[0067] In some embodiments, the local coefficient vector training includes;
[0068] Obtain the updated global coefficient vector from the server side in the previous round
[0069] According to Update the coefficient vector of the local part except BN and the last layer;
[0070] Using and Train the coefficient vector of the local client
[0071] By adjusting the parameter α (l,k) That is, the model weight W (l,k) , so that the model prediction result f k (x i ) and the real label y i The loss l between them is minimized.
[0072] In some embodiments, after obtaining the coefficient vector α of the client (l,k) , through the server side, the coefficient vectors of each client are transmitted, aggregated and distributed for information interaction between each client.
[0073] Compared with the prior art, the beneficial technical effects of the present invention are as follows:
[0074] 1. Through this system, the present application can, in the scenario of horizontal federated learning with model isomorphism and data distribution heterogeneity, through federated preparation, local coefficient vector training, coefficient vector aggregation and distribution, use a set of basis models to represent the client models as coefficient vectors, and through the training, aggregation and distribution of the coefficient vectors, replace the direct aggregation and distribution of the original model parameters, greatly improving the information transmission efficiency of federated learning;
[0075] 2. By maintaining the localization of some parameters, the present application ensures that each client model is developed for the local data distribution and has excellent personalization effects, mainly including three key points: the model representation method based on layer basis model decomposition, the model training method based on layer basis model decomposition, and the personalized federated aggregation method based on layer basis models.
[0076] 3. The system provided by the present application can, in the realistic horizontal federated learning scenario, based on basis model decomposition, through the training, transmission, aggregation and distribution of coefficient vectors, efficiently extract local data knowledge, transmit client information, and fuse the knowledge of each client, so as to obtain a more accurate personalized model. This system is general, robust and stable. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] Figure 1 is the working flow chart of the personalized federated learning method based on basis model decomposition according to the embodiment of the present application;
[0078] Figure 2 is the schematic diagram of data distribution according to the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0079] The technical solution of the present invention will be further clearly and detailedly described below in conjunction with the embodiments and the drawings.
[0080] Embodiment
[0081] As Figure 1 shown, the efficient personalized federated learning method based on basis model decomposition provided by the present application includes three stages: federated preparation, local coefficient vector training, and coefficient vector aggregation and distribution.
[0082] In the federated preparation stage, each client participating in the federated learning prepares its own data. At the same time, the server distributes the common basis model to each client, determines the coefficient vectors of the corresponding layers and the number of basis models used according to the number of parameters of each layer of the model, facilitating the subsequent training and parameter transmission;
[0083] In the local coefficient vector training stage, similar to the conventional model training, the coefficient vectors of each client are updated, while the basis model remains unchanged;
[0084] In the coefficient vector aggregation and distribution phase, the coefficient vectors of each client are transmitted to the server side, where the coefficient vectors are averaged and aggregated, and then the aggregated vectors are returned to each client. Note that in this phase, to enhance the personalization effect, only the coefficient vectors of some layers are transmitted; repeat the second and third phases until convergence or the maximum number of iterations is reached.
[0085] Specifically, it includes the steps: there are multiple clients, each client has an equal status, and there is a central server for information transmission, aggregation, and distribution;
[0086] Step 1: Federal preparation. Each client obtains the collected original data, and at the same time, the server distributes the common base model to the clients. Each client determines the coefficient vector of the corresponding layer and the number of base models to be used according to the number of parameters in each layer of the model;
[0087] The federal preparation stage of this application includes:
[0088] 1) Each client prepares its own data D i ;
[0089] 2) Each client divides the data into training data verification data test data
[0090] 3) The server distributes the base model to each client,
[0091] 4) Each client initializes local parameters, such as the coefficient vector α (l,k) .
[0092] Step 2: Local coefficient vector training. Train the coefficient vector of each local layer;
[0093] The specific steps of the local coefficient vector training are as follows:
[0094] (21) Problem description. Combine the data information of all clients to train a model f for each client i ,
[0095] Suppose there are N different clients, denoted as {C1, C2,..., C N}}, and the data corresponding to each federation is {D1, D2,..., D N}}. The data distributions are different, and at the same time, the spaces where the original data is located are the same, and the labels of the data are in the same space, that is, X i = X j , Y i = Y j , i ≠ j; where X i 、Xj denote the feature vectors of two different samples, Y i and Y j denote the labels corresponding to the said samples; the data of each federated end is divided into three parts, specifically training data validation data test data There are the total number of samples n of class i i equal to its training set n i tr and validation set n i val and test set n i te The sum of the number of samples; the complete data set D of class i i is its training set D i tr and validation set D i val and test set D i te Combining the data information of all clients, train a model f i for each client, and find the objective function, where l is the objective loss function;
[0096]
[0097] (22) Local coefficient vector training
[0098] Each network layer of the client model is represented as a linear combination of a group of randomly initialized base models, and any client model is represented by adjusting the relevant coefficient vectors;
[0099] Denote the parameters of the model f as W, the model has L layers, and W (l) represents the parameters of the l-th layer of the model. Assume that there are m l base models for the l-th layer, and the corresponding parameters are Then, W (l) is represented as,
[0100]
[0101] where represents the coefficient of the i-th base model corresponding to the l-th layer. By optimizing instead of directly optimizing w ι , the parameters of the network are represented as,
[0102] W = {W (1) , W (2) , …, W (L)}.
[0103] where each W(l) Represented by ;
[0104] Assume the model is approximately orthogonal, i.e.,
[0105]
[0106] The parameters of the k-th client are represented as follows
[0107]
[0108] The overall parameters are represented as
[0109] W k ={W (1,k) ,W (2,k) ,...,W (L,k)}.
[0110] Use numerical values to represent the entire network; m << d, where d represents the original parameter dimension;
[0111] Assume W *k is the optimal parameter of the k-th client, and the optimization for the k-th client is expressed as
[0112]
[0113] Use the coefficient vector α k to replace the optimization of the original model parameters, reducing the parameter dimension of the optimization;
[0114] Assume the loss function is l, and for the k-th client, there is the following objective
[0115]
[0116] where y i is the corresponding label. According to the above formula, optimize each α k
[0117] Use this stage to update and train the coefficient vector, so as to extract the information of the local client data, facilitating the construction of information transfer and communication between clients in the next step.
[0118] Step 3: Coefficient vector aggregation and distribution. Transfer the coefficient vectors of each client to the server side, perform average aggregation on the coefficient vectors at the server side, and then return the aggregated vectors to each client;
[0119] After obtaining the coefficient vector α (l,k) of each client, through the server side, transfer, aggregate and distribute the coefficient vectors of each client, so as to realize the information interaction of each client.
[0120] The specific implementation includes the following steps:
[0121] (31) The client transfers the coefficient vector to the server. The server aggregates the coefficient vectors from all clients, calculates the average value, and determines the new global coefficient vector for each layer.
[0122]
[0123] where α i (l,k) represents the i-th parameter or eigenvalue in the l-th layer of the k-th local model, and α i (l ,global l) represents the i-th parameter or eigenvalue in the l-th layer of the global model, N represents the total number of models, L represents the total number of layers of the model, m l represents the number of parameters or features in the l-th layer, i represents the index of the specific parameter position within a certain layer, and l represents the layer number of the model.
[0124] Preferably, each client locally maintains the scalar weights of the final linear layer without sharing them with the server. Specifically, it includes the steps:
[0125] (32) During the federated learning process, at the end of each training cycle, each client uploads the coefficient vector α (l,k) shared in its previous layers to the server, excluding the coefficient vectors of the final layer and the BN layer; the server aggregates the coefficient vectors from all clients, calculates their average value, and updates the global coefficient vector for each layer, but excluding the BN layer and the last layer, that is
[0126]
[0127] In some embodiments, step three further includes:
[0128] (33) After obtaining the global coefficient vector , distribute it to each client to update the coefficient vector of the corresponding layer;
[0129] Step four: Repeat step two and step three until convergence or the maximum number of iteration cycles is reached;
[0130] Step five: End.
[0131] Comparative example
[0132] To further verify the effectiveness of the proposed efficient personalized federated learning method and system based on base model decomposition and to illustrate the usage method of the present invention, the inventors also conducted experiments on real datasets. The real dataset uses OrganAMNIST in the commonly used public dataset MedMNIST [downloaded from
[0133] https: / / github.com / MedMNIST / MedMNIST / ?tab=readme-ov-file].
[0134] 1) Data acquisition
[0135] MedMNIST is a large-scale MNIST-like standardized biomedical image collection, including 12 2D datasets and 6 3D datasets. All images are 28×28 (2D) or 28×28×28 (3D). In this embodiment, 1 dataset with the largest number of classes is selected from the 12 2D datasets: OrganAMNIST. This dataset is all about abdominal CT images, contains 11 classes, and has 58,850 samples. To introduce heterogeneity among clients in FedBD, the Dirichlet distribution method is used to partition the client data into non-independent and identically distributed. As Figure 2 shown, the distribution of 20 clients is displayed. For each client, the data is divided into a training set, a validation set, and a test set in a ratio of 4:3:3.
[0136] 2) Comparison methods
[0137] To prove the effect of the present invention, a comparison is made with the classical federated learning method FedAVG [B. McMahan, E. Moore, D.
[0138] Ramage, S. Hampson, and B. A. y Arcas, “Communication-efficient learning of deep networks from decentralized data,” in Artificial intelligence and statistics. PMLR, 2017, pp.
[0139] 1273–1282].
[0140] 3) Evaluation metrics
[0141] In federated learning, evaluating model performance requires considering both accuracy and the number of communication parameters. To balance communication cost and high accuracy, a new evaluation metric called the Model Efficiency Index (MEI) is introduced. The formula for MEI is defined as follows:
[0142]
[0143] Here, MEI measures model efficiency by taking the square root of the ratio of the average accuracy to the number of communication parameters. This metric implies that for a given accuracy level, a model with fewer parameters will have a higher MEI. Therefore, MEI effectively measures the overall performance of models in federated learning.
[0144] 4) Result Analysis
[0145] Method Number of parameters Average precision MEI FedAVG 44599 88.40 4.45 The method of this patent 4895 91.78 13.69
[0146] Table 1 Overall Experimental Results
[0147]
[0148] Table 2 Precision Effects of Each Client
[0149] From Table 1 and Table 2, the following key observations can be obtained:
[0150] 1) Communication Efficiency: The proposed method only transmits 4,895 parameters (about 10% of the number transmitted by the traditional method), while achieving an accuracy comparable to that of models with more parameters.
[0151] This shows that the proposed method not only maintains high accuracy but also performs well in terms of communication efficiency, which is a key factor in a realistic environment where network resources may be limited. This efficiency is crucial in a realistic environment where network resources may be limited.
[0152] 2) Performance on Specific Clients: The proposed method achieves the highest accuracy on some clients. A possible explanation is that the model structure is simplified and has fewer parameters. The reduced complexity helps to minimize the risk of overfitting, enabling the proposed method to perform well on the data of specific clients. This observation indicates that models with too many parameters may become redundant and prone to overfitting, especially when dealing with heterogeneous data distributions.
[0153] Finally, it is necessary to point out here that: The above description is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.
Claims
1. A personalized federated learning method based on basis model decomposition, characterized in that: Includes steps: Step 1: Federation preparation: Each client obtains the collected raw data, and the server distributes the common base model to the client. The client determines the coefficient vector of the corresponding layer and the number of base models to be used according to the parameter quantity of each layer of the model. Step 2: Local coefficient vector training: training the coefficient vector of each local layer; Step 3: Coefficient vector aggregation and distribution: the coefficient vectors of each client are transmitted to the server, the coefficient vectors are averaged and aggregated on the server, and the aggregated vectors are returned to each client; Step 4: Repeat steps 2 and 3 until convergence or the maximum iteration cycle is reached; Step 5: End.
2. According to claim 1, a personalized federated learning method based on basis model decomposition is characterized in that: The federal preparation described in step 1 specifically includes the following steps: (11) Each client obtains the collected raw data D i ; (12) At the same time, the server distributes the common base model to the client; (13) The client initializes the local base model and determines the coefficient vector of the corresponding layer and the required number of base models according to the parameter quantity of each layer of the model.
3. The personalized federated learning method based on basis model decomposition according to claim 1, characterized in that: The local coefficient vector training described in step 2 includes the following specific steps: (21) Problem description: Combine the data information of all clients and train a model f for each client. i , Suppose there are N different clients, denoted as {C1, C2, ..., C N }, the data corresponding to each federation is {D1, D2, ..., D N }, the data distribution is different, but the original data is in the same space, and the data labels are in the same space, that is, X i =X j , Y i =Y j , i≠j; where X i , X j Represents the feature vectors of two different samples, Y i , Y j Indicates the label corresponding to the sample; the data of each federation end is divided into three parts, specifically training data Verify data Test Data have The total number of samples n of category i i Equal to its training set n i tr , validation set n i val and test set n i te The sum of the number of samples of category i; the complete data set D i is its training set D i tr , validation set D i val and the test set D i te Combine the data information of all clients and train a model f for each client i , find the objective function, where is the target loss function; (22) Local coefficient vector training Each network layer of the client model is represented as a linear combination of a set of randomly initialized base models, and any client model is represented by adjusting the relevant coefficient vector; The parameter of model f is W, and the model has L layers, W (l) Represents the parameters of the lth layer of the model. Assume that there are m l The corresponding parameters of the basic model are Then, W (l) It is expressed as, in Represents the coefficient of the i-th basis model corresponding to the l-th layer, through optimization Instead of directly optimizing W l , the parameters of the network are expressed as, In={In (1) ,IN (2) ,...,IN (L) }. Each W (1) Depend on express; Assume that the model is approximately orthogonal, that is The parameters of the kth client are expressed as follows, The overall parameters are expressed as, IN k ={W (1,k) ,IN (2,k) ,...,IN (L,k) }. use The value represents the entire network; m<<d, d represents the original parameter dimension; Assume W *k is the optimal parameter for the kth client. The optimization for the kth client is expressed as, Using the coefficient vector α k The optimization replaces the optimization of the original model parameters and reduces the dimension of the optimized parameters; Assume the loss function is For the kth client, there are the following goals: Among them, y i is the corresponding label. According to the above formula, for each α k Make optimizations.
4. According to the personalized federated learning method based on basis model decomposition according to claim, it is characterized in that: The coefficient vector aggregation and distribution described in step 3 specifically includes: (31) The client converts the coefficient vector Passed to the server, the server aggregates the coefficient vectors from all clients, calculates the average, and determines the new global coefficient vector for each layer. Among them, α i (l,k) Represented as the i-th parameter or eigenvalue of the k-th local model in the l-th layer, α i (l,global) It is represented as the i-th parameter or feature value of the global model in the l-th layer, N represents the total number of models, L represents the total number of layers of the model, and m l It represents the number of parameters or features of the lth layer, i represents the index of the specific parameter position within a certain layer, and l represents the number of layers of the model.
5. The personalized federated learning method based on basis model decomposition according to claim 4, characterized in that: Each client maintains the scalar weights of the final linear layer locally and does not share it with the server. Specifically, it includes the following steps: (32) In the federated learning process, at the end of each training cycle, each client shares the coefficient vector α of its previous layers. (l,k) Upload to the server, excluding the coefficient vectors of the final layer and the BN layer; the server aggregates the coefficient vectors from all clients, calculates their average, and updates the global coefficient vector of each layer, excluding the BN layer and the last layer, that is, 6. The personalized federated learning method based on basis model decomposition according to claim 4, characterized in that: The step three also includes: (33) Get the global coefficient vector After that, Distribute to each client and update the coefficient vector of the corresponding layer.
7. A system for implementing the personalized federated learning method based on basis model decomposition as claimed in claim 1, characterized in that: It includes a client for acquiring the collected raw data, and a server distributes a common base model to the client at the same time, and the client determines the coefficient vector of the corresponding layer and the number of base models to be used according to the parameter amount of each layer of the model; A representation module based on basis model decomposition: used to represent each layer parameter of the model as a linear combination of a set of basis models, use the coefficient vector of the linear combination of the basis models as the representation of the model, train the local coefficient vector, and train the coefficient vector of each local layer; Training module based on basis model decomposition: training the coefficient vector of the basis model instead of training all parameters of the entire model; Personalized federated aggregation module based on layer-based model: used to maintain some model parameters of the local client, without uploading the batch normalization layer parameters and the last layer coefficient vector of the aggregated local model.
8. The federated learning system based on basis model decomposition according to claim 7, characterized in that: The client determines the coefficient vector of the corresponding layer and the number of base models to be used according to the parameter amount of each layer of the model: Each client prepares its own data D i ; Each client divides the data into training data according to its needs Verify data Test Data The server distributes the base model to each client. Each client initializes local parameters, coefficient vector α (l,k) .
9. The federated learning system based on basis model decomposition according to claim 7, characterized in that: The local coefficient vector training includes: Get the global coefficient vector updated by the server in the previous round in accordance with Update the local coefficient vector except BN and the last layer; use as well as Train the coefficient vector on the local client By adjusting the parameter α (l,k) That is, the model weight W (l,k) , so that the model predicts the result f k (x i ) and the true label y i The loss l between them is minimal.
10. The federated learning system based on basis model decomposition according to claim 7, characterized in that: Get the client's coefficient vector α (l,k) Afterwards, the coefficient vectors of each client are transmitted, aggregated and distributed through the server for information exchange between the clients.