Personalized federal learning method based on singular value decomposition

By adopting the methods of singular value decomposition and self-knowledge distillation in personalized federated learning, selecting suitable clients and compressing parameter transmission, the problems of low training efficiency and insufficient memory in heterogeneous environments are solved, and efficient personalized federated learning is achieved.

CN120764632APending Publication Date: 2025-10-10BEIJING UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510882788.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-28
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

In a dynamic heterogeneous environment, traditional random or uniform sampling client selection strategies are difficult to adapt to changes in resource status, resulting in low training efficiency. In addition, the personalized forgetting problem increases memory usage and the risk of falling behind, affecting the effectiveness of personalized federated learning.

Method used

A personalized federated learning method based on singular value decomposition is adopted. The server monitors client resources and transmission rates, selects appropriate clients to participate in training, and combines singular value decomposition with self-knowledge distillation to transmit compressed key parameters, thereby reducing the overall delay and dropout rate.

Benefits of technology

It improves the training efficiency and model generalization ability of personalized federated learning in heterogeneous systems, reduces the device dropout rate and total latency, and improves the personalized performance and convergence speed of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120764632A_ABST
    Figure CN120764632A_ABST
Patent Text Reader

Abstract

The invention relates to a singular value decomposition-based personalized federated learning method, which is used for solving the problem of uncertainty of a system caused by equipment isomerism and the problem that a client falls behind due to extra burden caused by a storage and transmission mechanism in the prior art. Firstly, a customer selection strategy based on resource awareness is determined, and a server side establishes a quantitative scoring system by comprehensively analyzing computing resources and communication conditions of clients. And screening a plurality of clients with optimal comprehensive performance from all online clients to participate in local training. During local training of the client, historical weight parameters are reserved by adopting self-knowledge distillation (SKDSVD) combined with a singular value decomposition method, residual errors after decomposition are personalized model parameters and are stored in the local client, and universal parameters are transmitted to the server for aggregation. The SKDSVD algorithm reduces the training time delay and the fall-behind rate of the local client while optimizing the global model of the server to obtain the optimal convergence precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of Internet of Things and mobile edge computing, and specifically designs a personalized federated learning method based on singular value decomposition. Background Art

[0002] With the widespread adoption of smartphones and wearable devices, massive amounts of decentralized user data are continuously being generated. In practical applications, federated learning networks may involve a large number of participating IoT devices. Differences in device hardware capabilities, such as CPU, memory, and signal strength, can result in varying storage, computing, and transmission capabilities. Traditional AI applications typically require uploading private data to a central server for centralized training. However, this approach faces challenges with centralized data storage, which can easily lead to data leakage, and edge devices' insufficient computing power to meet training requirements. Federated learning (FL), a distributed machine learning paradigm, has emerged as an effective solution to these challenges by coordinating decentralized devices to collaboratively train a global model while maintaining data locality. However, FL trains a global shared model by aggregating model updates from various clients, but performs poorly in heterogeneous (non-independent and identically distributed) data scenarios. Therefore, personalized federated learning (PFL) for client-side training has become a research hotspot. Its core goal is to allow clients to perform local adaptations based on the global model.

[0003] Bai et al. proposed the pFedSD algorithm framework. The client uses a historical model trained locally using self-knowledge distillation (SD) as a teacher model to distill the current client model. By reusing historically updated training information, the client reduces communication costs with the server, ensuring generalization while improving personalized performance. It is worth noting that in personalized federated learning, the learning efficiency of the model is constrained by two key factors: the computing and storage resources of the client device and the characteristics of the training data. These factors together determine the upper limit of the model's personalization capabilities.

[0004] In the personalized federated learning process, in order to achieve local model compression and support self-distillation strategies, the client needs to retain some historical model information for residual modeling and inference assistance, which results in non-negligible memory usage. However, on resource-constrained devices, insufficient memory may prevent the device from completing local training; if it is directly eliminated, the accuracy of the overall model may be impaired. If the device is available, when its available memory falls below the set threshold, the operating system may swap some memory pages to disk, thereby triggering page replacement and causing high latency during training. In addition, due to resource competition caused by background applications running, the actual memory budget available for training on the device in different rounds fluctuates dynamically. On the other hand, due to differences in wireless communication capabilities between mobile devices, the heterogeneity of transmission rates will significantly affect the overall communication latency.

[0005] Through an in-depth analysis of existing research, we found that there are still two key issues that need to be addressed: (1) In dynamic heterogeneous environments, due to the uncertainty introduced to the system by the heterogeneity of device transmission rates and memory, traditional random or uniform sampling client selection strategies are difficult to adapt to the dynamically changing resource status, resulting in low client training efficiency. (2) To solve the personalized forgetting problem, the pFedSD algorithm designed a mechanism for storing and transmitting historical models. However, this type of solution will bring additional burdens to resource-constrained devices and may increase the risk of falling behind. At the same time, the overhead of transmission parameters and the actual transmission rate of the device closely affect the total latency. These problems restrict the practical application effect of personalized federated learning in resource-constrained scenarios.

[0006] In order to address the above problems, the goal of the present invention is to propose a personalized federated learning method based on singular value decomposition, taking into account the heterogeneity of device resources and the non-independent and identically distributed data. By designing a resource-aware client dynamic sampling strategy, appropriate edge devices are selected to participate in the PFL process in each training iteration, thereby improving the model's personalization capability and convergence efficiency, and reducing the dropout rate caused by insufficient resources. At the same time, the residual parameters after singular value decomposition are used as a personalized historical model, and the generalization ability of the local model is improved through self-distillation. At the same time, only the compressed key singular vectors and singular value matrices need to be transmitted to reduce the number of parameters in the communication process, thereby reducing the total system training delay and the probability of dropout. Summary of the Invention

[0007] In order to solve the problem of insufficient memory resources caused by inefficient client selection strategies and personalized forgetting in dynamic heterogeneous environments, the present invention proposes a personalized federated learning method based on singular value decomposition. The server monitors the memory of the client device and the actual transmission rate of the device to select sampling. The left and right singular value matrices and the diagonal matrix composed of singular values ​​are obtained by performing singular value decomposition (SVD) on the fully connected layer of the model. The key singular vectors and the singular value matrix are uploaded to the server to reduce the amount of parameter transmission. The client saves the calculated residual as a personalized parameter, that is, the difference between the client's model and the transmitted model. The purpose of the invention is to reduce the total delay and device dropout rate through reasonable client selection and local training, solve the incompatibility problem of the existing technology in heterogeneous system scenarios, and improve the efficiency of personalized federated learning.

[0008] The present invention provides a personalized federated learning method based on singular value decomposition, which is used to determine the client set selected for each round of training, and combines singular value decomposition with self-knowledge distillation method to build a local historical model. In order to achieve this goal, a client selection strategy based on resource perception is first determined. The server side comprehensively analyzes the computing resources and communication conditions of each client, establishes a quantitative scoring system, and selects several clients with the best overall performance to participate in local training to reduce the falling behind rate caused by resource bottlenecks. During local training, the singular value decomposition technology is combined with the self-knowledge distillation method, and the singular value vector is transmitted to the server side in the form of singular values ​​to reduce the total delay. The residual parameters generated by the singular value decomposition are used as the personalized historical model, and the model's fitting ability for local data is improved through self-distillation. Through the parameter singular value decomposition processing, the personalized characteristics of the model are retained, and the sharing and fusion of global knowledge are realized, thereby improving the training efficiency in a heterogeneous device environment.

[0009] The method of the present invention and its implementation principle are described in detail below.

[0010] Step 1. Each client node requests participation from the centralized server to initiate a PFL communication round.

[0011] The entire framework consists of a server S and N edge devices within its radius R, namely C1, C2, ... C N , and N ≥ 2. Each edge device has its own private data set where |D i | is the size of the data sample. x i,l is the lth data sample of device i, y i,l For data sample x i,l The actual label. In the initial state, the server and client use the same training model Historical model parameters are stored locally Initialization parameters: Set the number of communications to T max , the proportion of clients participating in local training of federated learning is recorded as r. max =50,r∈{0.6,0.8,1}, and the initial iteration round is set to t=1.

[0012] Step 2. After receiving the confirmation and device information of the interested client set, the server starts the client selection process, selects some devices to participate in federated learning, and broadcasts the global model And record it as the local training model

[0013] Step 2.1. When the communication round is t=1, there is no communication delay. Therefore, only all client devices C are calculated. i The calculation delay is shown in formula (1).

[0014]

[0015] CT i comp (t) represents the client C in communication round t i Estimated ideal computation latency. tnCPU i Indicates that the client executes each data sample D i The number of CPU clock cycles required, f i Indicates the computing power of the edge device, CPU frequency (in GHz). The above device parameters are set when they are initially configured. During local training, information such as model weights needs to be temporarily stored in the client's local memory. Therefore, in the tth round of training, client C i Memory usage It is expressed as shown in formula (2).

[0016]

[0017] Where |·| represents the size of the memory occupied. Client C i Local memory in round t The sources of occupation include: local training dataset | D i |, received global model parameters Self-distilled historical residual model Cache and Dynamic Fluctuation Memory | DFM i |, historical residual model at t = 1 In a highly dynamic training environment, effective scheduling based on the remaining available memory of the device and the real-time computing efficiency directly affects the accuracy of the model. i Remaining memory at round t The calculation is shown in (3).

[0018]

[0019] Among them CM i Represents client C i The initial memory size. This article introduces the remaining memory The computing power scoring mechanism driven by the client C i Remaining memory And process data sample D i The computation time CST i comp (t) as a numeracy score The core of is shown in formula (4).

[0020]

[0021] The value of ε is 1e-10 to avoid the denominator being 0. and max_CST i comp (t) represents all clients C i The maximum value of the remaining memory and computing time. The server to the client C i The calculated scores are sorted in descending order by N select =N*r select the top N customers select Customer engagement training.

[0022] Step 2.2. When the communication round is t>1, calculate the communication score and calculation score for all client devices, and sum them as the final score of the device. The calculation is shown in formula (5).

[0023]

[0024] in and They are respectively represented as client C in the tth round of training i Transmission model information delay and propagation distance delay, Represents client C i Estimate the size of the model parameters to be transferred, d i Represents client C i The distance to the server, ls represents the speed of light. Calculate the data rate r i The calculation is shown in formula (6).

[0025]

[0026] Among them B irepresents the bandwidth capacity of the edge device, N0 represents the noise power, and p i represents the transmission power density between the edge device and the server, h i Represents the channel gain. i 2 Calculate h i The value of B i ,N0,p i , respectively 50MHz, 10 -13 W, 2.2W. This paper takes the actual transmission rate of the client device as r i t And process data sample D i Communication time CST i comp (t) as the communication ability score The core of is shown in formula (7).

[0027]

[0028] The value of ε is 1e-10. Represents client C i The actual transmission rate at the current training round t, Indicates the maximum transmission rate of all clients in the current round. CST i comm (t) represents client C i Communication time for transmitting the model to the server, Indicates that all clients C when the communication round is t i The maximum communication time required to transmit the model to the server. The estimated communication time refers to the time it takes to upload the model parameters. The server's transmission rate is much higher than the client's. Therefore, the download time of all clients can be ignored without affecting fairness.

[0029] Step 2.3. Calculate client C i Comprehensive score As shown in formula (8).

[0030]

[0031] Step 2.4. Based on the comprehensive score of the client edge device Select the top N in descending order select =N*r customers.

[0032] Step 2.5. Server-side model Will be broadcast to the selected N select Client edge device, local client receives And use it as a training model

[0033] Step 3. Client local training process.

[0034] Step 3.1 Selected N select Client C i Activate and load the local private dataset D i Get personalized local training. Calculate the client C when the communication round is t i Number of iterations for local training B represents the local batch size, and the number of iterations is rounded up to ensure that all data is used.

[0035] Step 3.2. The process of training a local model using self-distillation.

[0036] 1) The data uses a hybrid enhancement strategy to generate enhanced samples. Client C i During t rounds of local training, the self-distillation method is used to optimize the local model. The training task loss is based on the student model output with mixed label y i,l Cross entropy Calculation, the calculation is shown in formula (9).

[0037]

[0038] in Represents the student model parameters After training, the sample x i,l The predicted output, y i,l is the real data label.

[0039] 2) In order to retain historical personalized knowledge, load the personalized teacher model of the previous round And extract the soft prediction of its last fully connected layer As the distillation target, it is shown in formula (10).

[0040]

[0041] Where τ is the temperature coefficient that adjusts the confidence smoothness of the soft label, τ∈{1,3}. Represents the predicted output of the last fully connected layer of the teacher model.

[0042] 3) Joint optimization and personalized retention. By continuously aligning the current model with the historical personalized model and the global model, client-level knowledge accumulation is achieved to adapt to the local data distribution. Teacher model As a carrier of local historical knowledge, through KL divergence KL i(·||·) constrains the output distribution of the student model to align with the teacher model to avoid the forgetting problem caused by the global model. Jointly optimize the cross entropy loss and KL divergence loss, the total loss function The calculation is shown in formula (11).

[0043]

[0044] Where ρ is the distillation weight, ρ∈{0.01.0.1,0.5,1}, which controls the retention strength of historical knowledge. is represented as the soft prediction of the student model, Represents the student model The output of the last fully connected layer, τ, is consistent with the teacher model.

[0045] 4) Backpropagation. Update the student model through Stochastic Gradient Descent (SGD) The process is shown in formula (12).

[0046]

[0047] Where η represents the learning rate, which is set to 0.01. Represents client C i The gradient of the loss function at training epoch t.

[0048] 5) After each round of local training is completed, client C i The model parameters The fully connected layer performs singular value decomposition and is decomposed into and The decomposition steps are as follows:

[0049] From the model Select the fully connected layer weight matrix The singular value decomposition is performed as shown in formula (13).

[0050]

[0051] in and Orthogonal matrices are left singular matrices and right singular matrices. is diag(σ1,σ2,…,σ r ),σ1≥σ2≥,…,≥σ r ≥0, Diagonal matrix, (·) T Transpose the matrix. By setting the truncation rank Decompose the singular values ​​into and The initial value is set to 2, and the amount of parameters transmitted through this value directly affects the total delay. Figure 3 The general parameters in the figure are shown in Figure 1. The model corresponding to the singular values ​​greater than the threshold is retained, that is, the global shared feature matrix learned during the training process. The calculation formula is shown in (14).

[0052]

[0053] in and For the left and right singular value matrices and arranged from large to small Singular values The local client will The decomposed matrix is ​​used as the locally trained model The local client obtains the retained personalized parameter part by calculating the residual parameter The calculation formula is shown in (15).

[0054]

[0055] in It mainly captures the fine-grained representation unique to local data and serves as a personalized teacher model for the next round of training. use.

[0056] Step 4. Record the target results of local training.

[0057] Compared with the traditional method of directly transmitting the complete model parameters, this solution adopts the following compression strategy: This paper only performs singular value decomposition on the fully connected layer weight matrix, and only uploads the truncated left / right singular matrices and the front Singular values, by reducing the amount of uploaded data to achieve the goal of reducing communication delay, while retaining the main information in the model. For each communication round of local training, it is necessary to quantitatively evaluate the delay and prediction loss value. During the training process, the client C i The total delay in round t is expressed as The calculation formula is shown in formula (16).

[0058]

[0059] where actCST i comm (t) and actCT i comp (t) are client C iThe real training communication time and training computation time at the tth round, which is obtained by experimental simulation training. and actCT i comp (t) Calculate as shown in equations (17) and (18).

[0060]

[0061] Wherein indicates the parameter size of the model of the tth round of real transmission, actTT i t indicates the training time of each round of the real device, indicates the number of iterations of the local training of the client C i Since the frequency of the device is limited by heat dissipation, power consumption, etc. to the ideal value, the real training computation time is larger, and the time completed by each communication round is limited by the slowest client. Therefore, the total training delay C total is the sum of the longest client in all rounds, as shown in equation (19).

[0062]

[0063] Wherein N select indicates the selected client participating in the training. Calculate the average model prediction loss of all selected client training as shown in equation (20).

[0064]

[0065] Wherein is the prediction loss value of the client C i in the tth round.

[0066] Step 5. The server will receive all the client models received by the server for aggregation processing. First, restore the matrix after singular value decomposition of the full connection layer in each client uploaded Then, in order to quantitatively evaluate the contribution of the user participating in the model training of a round, the server side aggregation when participating in the training of the client C i weight factor λ i is calculated as shown in equation (21).

[0067]

[0068] Wherein |D i | indicates the data set size owned by the client C i indicates the data set size owned by the client C i ​The comprehensive score is calculated by formula (8). Traverse each client model, align the parameters of each layer of the global model and the local model to achieve parameter-by-parameter aggregation, and add the local model parameters according to the weight to The process is shown in formula (22).

[0069]

[0070] Determine whether the iteration termination condition t=T is met max =50. If true, the training ends and the final global model parameters are obtained. Otherwise, the number of iteration steps is updated to t=t+1, and Broadcast to the local machine and used as the model parameters for training And continue the federated training process until the termination condition is met.

[0071] This model is used to solve the uncertainty problem of the system caused by device heterogeneity, as well as the problem that the storage and transmission mechanisms in existing technologies cause additional burdens and cause clients to fall behind. Dropout rate DR and total training delay C total Optimization target, optimization variable is N select and its corresponding As shown in formula (23).

[0072]

[0073] in Indicates the model parameters uploaded by local training in the tth communication round, N select The dropout rate (DR) is defined as the ratio of the number of mobile devices that cannot participate in federated learning training due to insufficient local memory to the total number of candidate mobile devices.

[0074] Beneficial effects

[0075] Compared to existing technologies, this paper fully accounts for the heterogeneity of edge devices by designing a personalized federated learning method based on heterogeneous device capability assessment, self-knowledge distillation training, and model parameter singular value decomposition compression. This method effectively reduces the dropout rate and overall latency by introducing a client selection mechanism. It also combines singular value decomposition technology with a self-distillation strategy to mitigate forgetting during personalized model training and reduce transmission latency, thereby improving the overall model's convergence speed and personalized performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] Figure 1 This is an example diagram of the system model;

[0077] Figure 2 Select for the client;

[0078] Figure 3 is the singular value decomposition process;

[0079] Figure 4 Comparison chart of average training loss values ​​of SKDSVD algorithm and BiasPrompt+ proposed in this paper;

[0080] Figure 5 The accuracy comparison chart of the SKDSVD algorithm and BiasPrompt+ proposed in this article. DETAILED DESCRIPTION

[0081] In order to make the purpose, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings and examples.

[0082] The present invention designs a personalized federated learning method based on singular value decomposition. Figure 1 Figure 2 is an example of a system model. In a personalized federated learning architecture in a heterogeneous environment, each client node submits metadata containing the local device hardware configuration and computing resources to the server, initiating a federated learning participation request. The server designs a comprehensive device scoring function based on multi-dimensional evaluation indicators, which is used to screen out a subset of clients with balanced capabilities to optimize the utilization of the system communication bandwidth, reduce the dropout rate, and reduce synchronization latency. After the selected client device receives the global model sent by the server, it uses self-knowledge distillation to collaboratively train the global model and the local historical model to adapt the local model. The local parameter matrix obtained from the training is subjected to singular value decomposition to decouple the device personalized parameters and global shared parameters. Personalized parameters are retained locally to maintain model adaptability, while global shared parameters are uploaded to the server for cross-client model aggregation to ensure personalized performance while achieving global knowledge sharing.

[0083] The specific steps are as follows:

[0084] Step 1. Based on the personalized federated learning architecture in a heterogeneous environment, establish a system model, set initialization parameters, and transmit the information of each client device to the server.

[0085] In this embodiment, the entire system model consists of a server S and N edge devices within a radius R of the server S. Polar coordinates are used to generate uniformly distributed client locations, as shown in formula (1).

[0086]

[0087] Where U(0,1) is a random number uniformly distributed between 0 and 1, and U(·) is a uniform distribution. i , and Customer C i The radial distance, azimuth and polar angle are used to generate a set of points uniformly distributed in a sphere with a radius of R, namely C1, C2, ... C N , and N ≥ 2. Convert spherical coordinates to rectangular coordinates Calculate the distance between each user and the server Each edge device has its own private dataset where |D i | is the size of the data sample. x i,l is the lth data sample of device i, y i,l For data sample x i,l The actual label. In the initial state, the server and client use the same training model Historical model parameters are stored locally Initialization parameters: Set the number of communications to T max , the proportion of clients participating in local training of federated learning is recorded as r. max =50, r∈{0.6,0.8,1}, and the initial value of the iteration number is set to t=1.

[0088] Step 2. The server receives the client's device information and scores the device status. Select some devices to participate in federated learning and broadcast the global model. And record it as the local training model

[0089] S1. When the communication round is t=1, only all client devices C are calculated i The calculation delay of each data D is calculated by the client. i The number of CPU cycles required tnCPU i The computing power of the edge device CPU frequency f i Ratio evaluation of edge device computational latency CT i comp (t).

[0090] During local training, model weights and other information need to be temporarily stored in the client's local memory. In a highly dynamic training environment, effective scheduling based on the device's remaining available memory and real-time computing efficiency directly affects the accuracy of the model. i Remaining memory And process data sample D i The computation time CST i comp (t) as a numeracy score core.

[0091] Server to Client C i The calculated scores are sorted in descending order by N select =N*r select the top N customers select Customer engagement training.

[0092] S2. When the communication round is t>1, the communication score and calculation score of all client devices are calculated, and the sum of the two is the final score of the device. i comm (t) is the transmission model information delay and propagation distance delay This article uses the actual transmission rate of the client device And process data sample D i Communication time CST i comp (t) as the communication ability score core.

[0093] S3. By calculating the client C i The sum of the communication capability score and computing capability score is recorded as client C i Comprehensive score

[0094] S4. Based on the comprehensive score of the client edge device Select the top N in descending order select =N*r customers.

[0095] S5. Server-side model Will be broadcast to the selected N select Client edge device, local client receives And use it as a training model

[0096] Instructions attached Figure 2 The figure shows the client selection process of personalized federated learning in a heterogeneous environment, namely steps S1-S5.

[0097] Step 3. The client receives the global model from the server and uses the local private dataset for personalized local training.

[0098] S1. Selected N select Client C i Activate and load the local private dataset D i Get personalized local training. Calculate the client C when the communication round is t i Number of iterations for local training B represents the local batch size, and the number of iterations is rounded up to ensure that all data is used.

[0099] S2. Client C i During t rounds of local training, the self-distillation method is used to optimize the local model. The training task loss is based on the student model output with mixed label y i,l Cross entropy Calculation. In order to retain historical personalized knowledge, load the personalized teacher model of the previous round And extract the soft prediction of its last fully connected layer as the distillation target.

[0100] S3. By continuously aligning the current model with the historical personalized model and the global model, client-level knowledge accumulation is achieved to adapt to the local data distribution. As a carrier of local historical knowledge, through KL divergence KL i (·||·) constrains the output distribution of the student model to align with the teacher model to avoid the forgetting problem caused by the global model. Jointly optimize the loss function including cross entropy loss and KL divergence loss

[0101] S4. Update the student model via Stochastic Gradient Descent (SGD)

[0102] S5. After each round of local training is completed, client C i The model parameters The fully connected layer performs singular value decomposition and is decomposed into and The decomposition steps are as follows:

[0103] From the model Select the fully connected layer weight matrix And perform singular value decomposition on it. By setting the singular value threshold Decompose the singular values ​​into and The initial value is set to 2. Figure 3 The process of singular value decomposition compressing local weights is explained. The model corresponding to the singular value greater than the threshold is retained, that is, the global shared feature matrix learned during the training process. The client obtains the retained personalized parameter part by calculating the residual parameter locally. It mainly captures the fine-grained representation unique to local data and uses it as a historical model for the teacher model PM in the next round of local training. i .

[0104] Step 4. Compared with the traditional method of directly transmitting the complete model parameters, this solution adopts the following compression strategy: This paper only performs singular value decomposition on the fully connected layer weight matrix, and only uploads the truncated left / right singular matrices and the front singular values, by reducing the amount of uploaded data to achieve the goal of reducing communication delay, while retaining the main information in the model. For each communication round of local training, the delay C total and the predicted loss value Conduct quantitative assessment.

[0105] Step 5. The server receives all client models Perform aggregation processing. First, the data uploaded by each client is The matrix after the singular value decomposition of the fully connected layer is restored and reconstructed. Then, in order to quantitatively evaluate the contribution of users participating in a round of model training, the server-side aggregated client C participating in the training is calculated. i Weight factor λ i Traverse each client model, align the parameters of each layer of the global model and the local model to achieve parameter-by-parameter aggregation, and add the local model parameters to the weighted

[0106] Determine whether the iteration termination condition t=T is met max =50. If true, the training ends and the final global model parameters are obtained. Otherwise, the number of iteration steps is updated to t=t+1, and Broadcast to the local machine and used as the model parameters for training And continue the federated training process until the termination condition is met.

[0107] This model is used to solve the uncertainty problem of the system caused by device heterogeneity, as well as the problem that the storage and transmission mechanisms in existing technologies cause the client to fall behind due to the additional burden. The dropout rate DR is defined as the ratio of the number of mobile devices that cannot participate in federated learning training due to insufficient local memory to the total number of candidate mobile devices. Therefore, considering the average model training loss Dropout rate DR and total training delay C total Optimization target, optimization variable is N select and its corresponding

[0108] The model used in the experiment is a simple convolutional neural network for image classification tasks, which includes two convolutional layers and two pooling layers to extract image features, then performs classification through a fully connected layer, and finally outputs the classification results. The dataset is the FMNIST image dataset of public handwritten digits, which covers 70,000 front-facing images of different products from 10 categories, with 55,000 / 5,000 / 10,000 training, validation, and test data divided into 28x28 grayscale images. Through simulation Figure 4 and Figure 5 Average model training loss and accuracy for parameters N = 20, r = 0.6, τ = 3, ρ = 0.1, η = 0.01, and B = 64. As can be seen from the figure, compared with the BiasPrompt+ algorithm, the proposed SKDSVD algorithm does not sacrifice model accuracy after singular value decomposition and can quickly reach convergence. Therefore, the proposed method can achieve better results during the iterative process.

Claims

1. A personalized federated learning method based on singular value decomposition, characterized in that: The following steps are involved: Step 1. Establish a personalized federated learning architecture in a heterogeneous environment; the entire framework consists of a server S and N edge devices within its radius R, namely C1, C2, ... C N , and N ≥ 2; each edge device has its own private data set |D i | is the size of the data sample; x i,l is the lth data sample of device i, y i,l For data sample x i,l The true label of Initially, the server and client use the same training model. Historical model parameters are stored locally Initialization parameters: Set the maximum number of communications to T max , the proportion of clients participating in local training of federated learning is recorded as r; where, r∈{0.6,0.8,1}, T max =50, the initial value of the iterative communication round is set to t=1; Step 2. Dynamically select clients based on resource awareness and the comprehensive score of client edge devices Ranking Select Top N select Clients participate in training to reduce the dropout rate caused by insufficient resources; the server-side model Will be broadcast to the selected local client and used as the local training model Step 3. Based on the self-distillation model training, the client performs singular value decomposition on the fully connected layer of the model parameters, and uploads the general model parameters corresponding to the singular value threshold to the server; The model parameters before decomposition minus the uploaded general model parameters are called residual parameters, which are used as the local historical model for the teacher model in the next round of training. Step 4. The server aggregates the received parameters by weight and then determines whether the iteration termination condition t=T is met. max And the target converges; If true, the training ends and the final global model parameters are obtained Otherwise, the number of iterations is updated to t=t+1, and the federated training process continues to execute step 2.

2. The personalized federated learning method based on singular value decomposition according to claim 1, characterized in that: Furthermore, the selection based on the client device resources includes the following steps: The server receives the device information from the client and scores the device status; selects some devices to participate in federated learning and broadcasts the global model. And record it as the local training model S1. When the communication round is t=1, there is no communication delay; therefore, only all client devices C are calculated. i The calculation delay of is shown in formula (1); where |D i | represents the size of the data sample, Indicates client C in communication round t i Estimated ideal computation latency; tnCPU i Indicates that the client executes each data sample D i The number of CPU clock cycles required, f i Indicates the computing power of the edge device CPU frequency (unit: GHz). The above device parameters are set when they are initially configured. During local training, information such as model weights needs to be temporarily stored in the client's local memory. Therefore, in the tth round of training, client C i Memory usage It is expressed as shown in formula (2); Where |·| represents the size of the memory occupied; client C i Local memory in round t The sources of occupation include: local training dataset | D i |, received global model parameters Self-distilled historical residual model Cache and Dynamic Fluctuation Memory | DFM i |, historical residual model at t = 1 In a highly dynamic training environment, effective scheduling based on the remaining available memory and real-time computing efficiency of the device directly affects the accuracy of the model; Client C i Remaining memory at round t The calculation is shown in (3); Among them, CM i Represents client C i Initial memory size; introduce remaining memory The computing power scoring mechanism driven by the client C i Remaining memory And process data sample D i Computation time As a numeracy score The core of is shown in formula (4); The value of ε is 1e-10 to avoid the denominator being 0; and Represents all clients C i The maximum value of the remaining memory and computing time; the server to the client C i The calculated scores are sorted in descending order by N select =N*r select the top N customers select Customer engagement training; S2. When the communication round is t>1, the communication score and calculation score of all client devices are calculated, and the sum of the two is the final score of the device; the transmission delay between the client edge device and the server The calculation is shown in formula (5); in and They represent the client C in the tth round of training. i Transmission model information delay and propagation distance delay, Represents client C i Estimate the size of the model parameters to be transferred, d i Represents client C i The distance to the server, ls represents the speed of light; the rate of calculating data r i The calculation is shown in formula (6); Among them B i represents the bandwidth capacity of the edge device, N0 represents the noise power, and p i represents the transmission power density between the edge device and the server, h i represents the channel gain; through 1 / d i 2 Calculate h i The value of B i ,N0,p i , respectively 50MHz, 10 -13 W, 2.2W; actual transmission rate of the client device And process data sample D i Communication time As a communication ability score The core of is shown in formula (7); The value of ε is 1e-10. Represents client C i The actual transmission rate at the current training round t, Indicates the maximum transmission rate of all clients in the current round; Represents client C i Communication time for transmitting the model to the server, Indicates that all clients C when the communication round is t i The maximum communication time required to transmit the model to the server. The estimated communication time refers to the time required to upload the model parameters. S3. Computing client C i Comprehensive score As shown in formula (8); S4. Based on the comprehensive score of the client edge device Select the top N in descending order select =N*r customers; S5. Server-side model Will be broadcast to the selected N select Client edge device, local client receives And use it as a training model 3. The personalized federated learning method based on singular value decomposition according to claim 1, characterized in that: Furthermore, the self-distillation model training combined with singular value decomposition on the client includes the following steps: S1. Selected N select Client C i Activate and load the local private dataset D i Conduct personalized local training; through Calculate the client C when the communication round is t i Number of iterations for local training B represents the local batch size, and the number of iterations is rounded up to ensure that all data is used; S2. Client C i During t rounds of local training, the self-distillation method is used to optimize the local model. The training task loss is based on the student model output with mixed label y i,l Cross entropy Calculation, the calculation is shown in formula (9); in Represents the student model parameters After training, the sample x i,l The predicted output, y i,l is the real data label; in order to retain historical personalized knowledge, load the personalized teacher model of the previous round And extract the soft prediction of its last fully connected layer As the distillation target, as shown in formula (10); Where τ is the temperature coefficient that adjusts the confidence smoothness of the soft label, τ∈{1,3}; Represents the predicted output of the last fully connected layer of the teacher model; S3. By continuously aligning the current model with the historical personalized model and the global model, client-level knowledge accumulation is achieved to adapt to the local data distribution; teacher model As a carrier of local historical knowledge, through KL divergence KL i (·||·) constrains the output distribution of the student model to align with the teacher model to avoid the forgetting problem caused by the global model; jointly optimize the cross entropy loss and KL divergence loss, the total loss function The calculation is shown in formula (11); Where ρ is the distillation weight, ρ∈{0.01.0.1,0.5,1}, which controls the retention strength of historical knowledge; is represented as the soft prediction of the student model, Represents the student model The output of the last fully connected layer, τ, is consistent with the teacher model; S4. Update the student model via Stochastic Gradient Descent (SGD) The process is shown in formula (12); Where η represents the learning rate, which is set to 0.

01. Represents client C i The gradient of the loss function at training round t; S5. After each round of local training is completed, client C i The model parameters The fully connected layer performs singular value decomposition and is decomposed into and The decomposition steps are as follows: From the model Select the fully connected layer weight matrix Perform singular value decomposition on it as shown in formula (13); in and The orthogonal matrices are left singular matrices and right singular matrices; is diag(σ1,σ2,…,σ r ),σ1≥σ2≥,…,≥σ r ≥0, Diagonal matrix, (·) T Transpose the matrix; by setting the truncation rank Decompose the singular values ​​into and The initial value is set to 2, which directly affects the total delay by affecting the amount of parameters transmitted; Figure 3 of the specification illustrates the process of singular value decomposition compressing local weights; the general parameter part in the figure The model corresponding to the singular value greater than the threshold is retained, that is, the global shared feature matrix learned during the training process; The calculation formula is shown in (14); in and For the left and right singular value matrices and arranged from large to small Singular values The local client will The decomposed matrix is ​​used as the locally trained model Transmit to the server; the client obtains the retained personalized parameter part by calculating the residual parameter locally The calculation formula is shown in (15); in It mainly captures the fine-grained representation unique to local data and serves as a personalized teacher model for the next round of training. use; The following compression strategy is used: only perform singular value decomposition on the fully connected layer weight matrix, and only upload the truncated left / right singular matrix and the front singular values, by reducing the amount of uploaded data to achieve the goal of reducing communication delay, while retaining the main information in the model; for each communication round of local training, it is necessary to quantitatively evaluate the delay and prediction loss value. During the training process, the client C i The total delay in round t is expressed as The calculation formula is shown in formula (16); in and Represents client C i The actual training communication time and training calculation time in round t are obtained from experimental simulation training; and The calculation is shown in formulas (17) and (18); in Indicates the size of the parameters of the model for the actual transmission of round t, Indicates that each real device The training time of each round, Indicates that client C is in communication round t i The number of iterations of local training; d i ,ls,r i They represent the distance between the device and the server, the speed of light, and the data transmission rate respectively. Due to the limitations of heat dissipation and power consumption, the frequency of the device cannot reach the ideal value, which makes the actual training calculation time longer. At the same time, the time to complete each communication round is limited by the slowest client. Therefore, the total training delay C total is the sum of the longest time spent by the client in all rounds, as shown in formula (19); where N select Represents the clients selected for training; calculates the average model prediction loss of all selected clients As shown in formula (20); in For client C i The predicted loss value at the tth round.

4. The personalized federated learning method based on singular value decomposition according to claim 1, characterized in that: Furthermore, the specific steps of using weight aggregation to obtain a global model include the following: The server receives all client models Perform aggregation processing; first, the data uploaded by each client The matrix after the singular value decomposition of the fully connected layer is restored and reconstructed; then, in order to quantitatively evaluate the contribution of users participating in a round of model training, the client C participating in the training is aggregated on the server side. i Weight factor λ i Calculate as shown in formula (21); where |D i | indicates client C i The size of the dataset you have; Represents customer C i The comprehensive score is calculated by formula (8); traverse each client model, align the parameters of each layer of the global model and the local model to achieve parameter-by-parameter aggregation, and add the local model parameters according to the weight to The process is shown in formula (22); Determine whether the iteration termination condition t=T is met max =50; if true, the training ends and the final global model parameters are obtained. Otherwise, the number of iteration steps is updated to t=t+1, and Broadcast to the local machine and used as the model parameters for training And continue the federated training process until the termination condition is met; This model is used to solve the uncertainty problem of the system caused by device heterogeneity, as well as the problem that the storage and transmission mechanisms in existing technologies cause additional burdens that cause clients to fall behind; taking into account the average model training loss Dropout rate DR and total training delay C total Optimization target, optimization variable is N select and its corresponding As shown in formula (23); in Indicates the model parameters uploaded by local training in the tth communication round, N select The clients selected to participate in the training; the dropout rate DR is defined as the ratio of the number of mobile devices that cannot participate in the federated learning training due to insufficient local memory to the total number of all candidate mobile devices.

Citation Information

Cited By

  • Federal learning method and system oriented to Internet of Things communication privacy protection, and storage medium

    CN121543125A