An accelerated federated learning method based on parameter selection and presynchronization

By selecting and pre-synchronizing some important parameters in the mobile edge network, and combining this with deep Q-networks to optimize the aggregation frequency, the communication cost and privacy protection issues of federated learning in heterogeneous systems are resolved, resulting in a more efficient training process.

CN116611502BActive Publication Date: 2026-05-12CHINA THREE GORGES UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA THREE GORGES UNIV
Filing Date
2023-05-15
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In mobile edge networks, the efficient deployment of federated learning is affected by system heterogeneity, bandwidth limitations, and differences in device computing power, resulting in high communication costs and privacy protection challenges.

Method used

An accelerated federated learning method based on parameter selection and pre-synchronization is adopted. By selecting some important parameters for transmission at the user end and pre-synchronizing them between base stations, the aggregation frequency is optimized by combining the alternating minimization algorithm of deep Q-networks to minimize training loss and completion time.

Benefits of technology

It effectively reduces the transmission overhead of federated learning, improves training efficiency, and reduces the average FL completion time and total training loss by 33.83%, adapts to system heterogeneity, and ensures privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116611502B_ABST
    Figure CN116611502B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of Internet of Things, in particular to an acceleration federated learning method based on parameter selection and pre-synchronization. The acceleration federated learning method accelerates the training of federated learning based on a federated learning framework of parameter selection and pre-synchronization. The method distributes the latest global model to all base stations through a center server, and each base station distributes the received global model to connected users. The users train the global model through a local training round set, and calculate the gradient and model parameters of the received global model. The users transmit selected model parameters to the connected base stations, and the base stations broadcast the model parameters between the base stations through a wired link. The model parameters are pre-synchronized until a predetermined number of pre-synchronization rounds is reached. After receiving the aggregation results from all the base stations, the center server updates the global model. The selected parameter transmission and parameter pre-synchronization can reduce system overhead while maintaining the training performance of the federated learning model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of Internet of Things (IoT) technology, and more specifically, to an accelerated federated learning method based on parameter selection and pre-synchronization. Background Technology

[0002] With the rapid development of the Internet of Things (IoT) and social networking applications, the data generated by terminal devices such as smartphones, IoT devices, and sensors is growing exponentially. The success of IoT depends on dynamic sensing and intelligent decision-making. However, due to the limitations of computing and communication resources in terminal devices, IoT devices cannot simultaneously provide real-time and high-precision results. To address these challenges, Mobile Edge Computing (MEC) has been proposed. MEC pushes computing, caching, and networking functions to the network edge to perform task processing and provide services, avoiding unnecessary transmission latency. However, MEC requires raw data from IoT devices, which raises serious privacy concerns and communication costs.

[0003] To address the challenges of privacy protection and big data, Federated Learning (FL) has emerged as an attractive distributed learning paradigm. Federated Learning is a model training technique where data owners can train models locally without sharing data, helping to ensure privacy and reduce communication costs. Model training in Federated Learning is accomplished through a collaborative process. In this process, users send their current model, which is then updated using a locally collected dataset via stochastic gradient descent (SGD). After a certain number of local training sessions, users send their updated model weights to an aggregator, which updates the global model. Federated Learning has proven its effectiveness in various statistically heterogeneous environments, such as imbalanced, non-independent, and identically distributed (non-IID) data. However, efficient deployment of Federated Learning in mobile edge networks requires consideration of system heterogeneity. This is because in mobile edge environments, system bandwidth is limited and shared by all connected mobile devices, with potential mutual interference. Furthermore, due to mobility and channel fading, selected devices may have different computing capabilities and dynamic wireless channel conditions. To address the aforementioned issues, an effective federated learning framework should be designed to dynamically adjust the aggregation frequency and strategy to adapt to the heterogeneity of the system. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide an accelerated federated learning method and a novel fine-grained federated learning framework based on parameter selection and pre-synchronization (PSPFL) to accelerate federated learning training. Specifically, this invention designs a selected parameter transmission to perform fine-grained partitioning of the federated learning model on the user. Base stations (BSs) synchronize the parameters uploaded by users to each other. Then, a deep Q-network with Alternating Minimization (DQNAM) is proposed to select the optimal aggregation frequency on the central server to minimize training loss and federated learning completion time.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] According to a first aspect of the present invention, an accelerated federated learning method based on parameter selection and pre-synchronization is provided. This accelerated federated learning method accelerates the training of federated learning based on a parameter selection and pre-synchronization federated learning framework, and includes the following steps:

[0007] The central server distributes the latest global model to all base stations, and each base station distributes the received global model to the connected users.

[0008] The user trains the global model using a local training set and calculates the gradient and model parameters of the received global model.

[0009] Users will select some model parameters to transmit to the connected base stations, and the base stations will broadcast the model parameters between the base stations via wired links;

[0010] The model parameters are pre-synchronized until a predetermined number of pre-synchronization rounds are reached. After receiving the aggregation results from all base stations, the central server updates the global model.

[0011] As a further aspect of the present invention, when a user trains and calculates the gradient and model parameters of the received global model using a local training round set, the local training round set... Among them, J l It is the number of local training rounds in the l-th federated learning round.

[0012] As a further aspect of the present invention, after the base station broadcasts the model parameters between base stations via a wired link, the base station is also used to distribute the model parameters to connected users, wherein, after the model parameters are pre-synchronized, the number of pre-synchronization rounds in each federated learning round l is defined as... Among them, I l This indicates the number of pre-synchronization rounds in the l-th federated learning round.

[0013] As a further aspect of the present invention, during the training of the global model, multiple federated learning rounds are performed between the user and the central server. In each federated learning round... Where l is the total number of FL rounds; and the computation time of user m in the l-th FL cycle is... for:

[0014]

[0015] In the formula, β represents the number of CPU cycles required for each data sample, and J l This indicates the number of local training iterations. This represents the computing resources of user m, in units of [cycles / s].

[0016] Each user A three-element tuple {c m D m ,f m}, where c m D represents the magnitude of the model parameters. m Let f represent the size of the dataset for each user m, and f m This represents the computing resources of each user m; a group of users

[0017] As a further aspect of the present invention, during the training of the global model, in the i-th pre-synchronization round of the l-th FL round, the variation range of the local model parameters of user m during the training of the local model is... for:

[0018]

[0019] In the formula, This represents the local model parameters.

[0020] As a further aspect of the present invention, the accelerated federated learning method based on parameter selection and pre-synchronization uses an adaptive significance threshold σ' to represent the attenuation significance threshold during federated learning. The adaptive significance threshold σ' is defined as:

[0021]

[0022] Among them, I l It is the parameter pre-synchronization loop number in the l-th FL loop.

[0023] As a further aspect of the present invention, in the model parameter pre-synchronization, each user only uploads important parameters to the corresponding base station, and the transmission time of user m in the i-th parameter pre-synchronization round is... The calculation is as follows:

[0024]

[0025] in, R is the size of the transmission parameters of user m in the i-th parameter pre-synchronization; m E is the transmission rate of user m; m,b Q represents the channel bandwidth between user m and BS b; m,b G represents the transmission power of user m; m,b ε is the channel gain between user m and BSb; 2 This represents the standard deviation of the Gaussian channel noise.

[0026] As a further aspect of the present invention, when a predetermined number of pre-synchronization rounds are reached, each user m uploads all local model parameters to the corresponding base station, and the parameter transmission time of user m in the i-th parameter pre-synchronization round is... The calculation is as follows:

[0027]

[0028] Obtain the completion time of the federated learning round r; where the completion time of user m's lFL round is calculated as follows:

[0029]

[0030] Under the synchronous federated learning settings, export the completion time of federated learning round l:

[0031] T l =max{T l,1 ,T l,2 ,…,T l,M}

[0032] As a further aspect of the present invention, the model parameters are pre-synchronized until a predetermined number of pre-synchronization rounds are reached. This also includes jointly minimizing the FL training loss and training time to determine the optimal local training rounds and parameter pre-synchronization rounds.

[0033] Initialize the pre-synchronization round number I and the local training loop number J in the federated learning round;

[0034] The local training cycle number J and the pre-synchronization round number I alternately perform optimization operations. The alternating minimization algorithm is used to obtain the common minimization of FL training loss and training time through pre-training.

[0035] As a further aspect of the present invention, when the central server updates the global model after receiving the aggregation results from all base stations, the improvement of the deep Q network includes experience replay and a separate target network; wherein, during experience replay, in each federated learning round, the deep Q network stores the transformation set in the replay memory and randomly extracts a local training round set to train the global model; during the separate target network, the target network is cloned by the main network in each transformation set, and the loss is minimized using the action values ​​of the target network.

[0036] In a second aspect, the present invention provides a computer device, including a memory, a processor, and a computer program running on the processor, wherein the processor executes the program to implement the steps of the above-described accelerated federated learning method based on parameter selection and pre-synchronization.

[0037] Thirdly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described accelerated federated learning method based on parameter selection and pre-synchronization.

[0038] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:

[0039] 1. This invention proposes a novel federated learning framework based on parameter selection and pre-synchronization, which jointly optimizes the federated learning completion time and training loss. To the best of our knowledge, this work is the first to jointly optimize training loss and time consumption through a fine-grained federated learning framework.

[0040] 2. This invention innovatively proposes the selection of parameter transmission and parameter pre-synchronization, which allows for the selection of some important parameters for transmission and aggregation, thereby reducing transmission overhead and accelerating federated learning training.

[0041] 3. This invention proposes an accelerated federated learning method based on parameter selection and pre-synchronization to obtain near-optimal local training rounds and parameter pre-synchronization rounds. Comprehensive simulation experiments were conducted to verify the effectiveness of the proposed method on real-world datasets.

[0042] These or other aspects of this application will become more apparent from the following description of embodiments. It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the application. Attached Figure Description

[0043] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. In the drawings:

[0044] Figure 1The flowchart illustrates an accelerated federated learning method based on parameter selection and pre-synchronization in an exemplary embodiment of the present invention.

[0045] Figure 2 The diagram illustrates the workflow of PSPFL in an accelerated federated learning method based on parameter selection and pre-synchronization in an exemplary embodiment of the present invention.

[0046] Figure 3 This illustration schematically demonstrates an accelerated federated learning method based on parameter selection and pre-synchronization on MNIST, FashionMNIST, and CIFAR10 in an exemplary embodiment of the present invention. Performance graph;

[0047] Figure 4 The illustration schematically shows a global loss function graph between PSPFL and FedAvg on MNIST with different numbers of BS in an accelerated federated learning method based on parameter selection and pre-synchronization in an exemplary embodiment of the present invention.

[0048] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0049] The present application will now be further described in conjunction with the accompanying drawings and specific embodiments. It should be noted that, without conflict, the various embodiments or technical features described below can be arbitrarily combined to form new embodiments.

[0050] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.

[0051] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0052] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.

[0053] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0054] Efficiently deploying federated learning in mobile edge networks requires consideration of system heterogeneity. This is because in mobile edge environments, system bandwidth is limited and shared by all connected mobile devices, potentially leading to mutual interference. Furthermore, due to mobility and channel fading, selected devices may have varying computational capabilities and dynamic wireless channel conditions. To address these issues, this invention provides an accelerated federated learning method based on parameter selection and pre-synchronization.

[0055] In some implementations, accelerated federated learning methods based on parameter selection and pre-synchronization can be applied to computer devices, such as PCs, laptops, mobile terminals, or other devices with display and processing capabilities, but are not limited to these.

[0056] Please refer to Figure 1 , Figure 1 This is a flowchart of the accelerated federated learning method based on parameter selection and pre-synchronization according to this application. In the embodiments of this application, the accelerated federated learning method accelerates the training of federated learning based on the federated learning framework of parameter selection and pre-synchronization. The accelerated federated learning method based on parameter selection and pre-synchronization includes the following steps:

[0057] Step S10: The central server distributes the latest global model to all base stations, and each base station distributes the received global model to the connected users.

[0058] Step S20: The user trains the global model through the local training round set and calculates the gradient and model parameters of the received global model.

[0059] Step S30: The user will select some model parameters and transmit them to the connected base stations. The base stations will broadcast the model parameters between the base stations through wired links.

[0060] Step S40: Pre-synchronize the model parameters until a predetermined number of pre-synchronization rounds are reached. After receiving the aggregation results from all base stations, the central server updates the global model.

[0061] In this application's accelerated federated learning method based on parameter selection and pre-synchronization (PSPFL), a novel fine-grained FL framework is proposed to accelerate FL training. Specifically, the invention designs a selected parameter transfer to perform fine-grained partitioning of the FL model on the user. Base stations (BSs) synchronize the parameters uploaded by the user with each other. Then, a deep Q-network employing the alternating minimization method (DQNAM) is proposed to select the optimal aggregation frequency on the central server to minimize training loss and FL completion time.

[0062] In this invention, the accelerated federated learning method accelerates federated learning training based on a parameter selection and pre-synchronization federated learning framework. This framework accelerates FL (Finite Flow), with PSPFL being a fine-grained framework that allows model parameters to be divided into fine-grained partitions. Specifically, the user selects a subset of parameters to transfer, rather than all local model parameters. The BS (Base Flow) can synchronize these subsets of parameters from the user with each other and then distribute them to the user.

[0063] In this invention, a group of users is considered. The client's dataset is and Each user m in the dataset has its own dedicated training dataset. Includes D data samples, among which and Let represent the feature vector and corresponding label of the d-th training sample at user m, respectively. The local model parameters for user m are ω. m The local loss function for user m is expressed as f(ω). m ;x d ,y d ), where (x d ,y d )yes The d-th data sample in the dataset. This invention defines the set of BS. let Represents the set of relationships between users and BS, where e m,b =1 indicates that user m is connected to BS b; otherwise, e m,b =0. Furthermore, this invention uses This represents the set of users connected to BS b.

[0064] The FL training process requires multiple FL loops between the user and the central server to achieve the desired model accuracy.

[0065] Specifically, in each FL round, Where L is the total number of FL wheels. To characterize user state information, this invention provides information for each user. A three-element tuple {c m D m ,f m}, where c m D represents the magnitude of the model parameters. m This indicates the size of its dataset, while f m This indicates its computing resources.

[0066] In this embodiment, the workflow of the accelerated federated learning method based on parameter selection and pre-synchronization is as follows: Figure 2 As shown. Each round of federated learning in the accelerated federated learning method based on parameter selection and pre-synchronization proposed in this invention is described as follows:

[0067] 1) The Central Server (CS) distributes the latest global model to all base stations (hereinafter referred to as BS) (the model parameters are randomly initialized in the first FL round). Then, each BS distributes the received global model to its connected users via radio broadcast.

[0068] 2) Each user calculates the gradients and model parameters of the model received from the connected base station through local training. Local training round set. J l It is the number of local training rounds in the l-th FL round.

[0069] 3) In order to save communication costs, in the framework of this invention, each user will select some important parameters and transmit them to the corresponding base station.

[0070] 4) Base stations broadcast important parameters to each other via wired links and aggregate them.

[0071] 5) After synchronization, the base station distributes the model parameters to its users. This invention uses steps 3)-5) as parameters for pre-synchronization, and defines the number of pre-synchronization rounds in each FL round as... Among them, I l This indicates the number of pre-synchronization wheels in the l-th FL wheel.

[0072] 6) Repeat steps 3)-5) until the preset parameter pre-synchronization loop is reached. Then, each BS aggregates the entire updated model parameters and loss gradients from the user and uploads the aggregated results to the CS.

[0073] 7) After receiving the aggregation results from all BSs, CS updates the global model using an aggregation algorithm (e.g., the FedAvg algorithm).

[0074] Unlike existing FL frameworks, the accelerated federated learning framework based on parameter selection and pre-synchronization of this invention involves two novel designs: selected parameter transfer and parameter pre-synchronization, to reduce system overhead while maintaining the training performance of FL models.

[0075] In this embodiment, the global model is trained through multiple federated learning rounds between the user and the central server. In each federated learning round, Where L is the total number of FL rounds; and the computation time for user m in the l-th FL cycle is... for:

[0076]

[0077] In the formula, β represents the number of CPU cycles required for each data sample, and J l This indicates the number of local training iterations. [cycles / s] represents the computing resources of user m;

[0078] Each user A three-element tuple {c m D m ,f m}, where c m D represents the magnitude of the model parameters. m Let f represent the size of the dataset for each user m, and f m This represents the computing resources of each user m; a group of users

[0079] During the training of the global model, in the i-th pre-synchronization round of the l-th FL round, the range of variation of the local model parameters for user m is as follows: for:

[0080]

[0081] In the formula, This represents the local model parameters.

[0082] To distinguish whether the range of parameter variation is significant, this invention defines σ as a significance threshold. Parameters with a range of variation greater than σ are considered important parameters. It is worth noting that as the FL process progresses, the range of parameter variation will decrease. Therefore, to ensure that the FL converges to the desired point, the significance threshold needs to be attenuated during the FL process. Therefore, this invention employs an adaptive significance threshold σ', defined as:

[0083]

[0084] Among them, I l It is the parameter pre-synchronization loop number in the l-th FL loop.

[0085] This invention can also describe parameter transmission time. For the channel model between a user and its connected BS, this invention considers using Orthogonal Frequency Division Multiple Access (OFDMA) as the multiple access scheme in uplink transmission. The user's transmission time can be divided into two cases. In parameter pre-synchronization, each user only uploads important parameters to the corresponding BS. Then, the transmission time of user m in the i-th parameter pre-synchronization round... It can be calculated as follows:

[0086]

[0087] in, R is the size of the transmission parameters of user m in the i-th parameter pre-synchronization. m Let E be the transmission rate of user m, which can be determined using Shannon's theorem. m,b Q represents the channel bandwidth between user m and BSb. m,b G represents the transmission power of user m. m,b ε is the channel gain between user m and BSb. 2 This represents the standard deviation of the Gaussian channel noise.

[0088] When the predetermined number of pre-synchronization rounds is reached, each user m needs to upload all local model parameters to its corresponding BS. Combining the above two cases, the parameter transmission time for user m in the i-th parameter pre-synchronization round is... The calculation is as follows:

[0089]

[0090] Then, the completion time of FL round r can be obtained. When the number of local training rounds J... l When fixed, the local computation time in each pre-synchronization round is also constant. Then, the completion time for the lFL round with respect to user m is calculated as follows:

[0091]

[0092] Under synchronous FL settings, the completion time of FL wheel l can be derived as follows:

[0093] T l =max{T l,1 ,T l,2 ,…,T l,M}

[0094] The optimization objective of this invention is to minimize the weighted sum of FL completion time and training loss, that is: the objective function P1 is to minimize the weighted sum of FL completion time and training loss, as shown in the following formula:

[0095]

[0096] The constraint requires that the completion time of any user in any FL round should not exceed the maximum allowed time T. max ,Right now:

[0097]

[0098] The constraints also include ensuring that the obtained training loss does not exceed the maximum tolerable training loss F. max ,Right now:

[0099] F(ω* )≤F max .

[0100] This invention now analyzes the objective function P1. For the total FL training time, this invention... The time after the lFL round can be rewritten as:

[0101]

[0102] However, due to the various combinations of control variable M, T in Equation 7 l It remains difficult to analyze. Regarding analyzability, this invention defines T... l An approximation of . Let This represents the average time required for all M users to complete one round of local training, and This is the average transmission time. Therefore, the present invention can approximate the FL completion time as:

[0103]

[0104] Regarding the approximation of the training loss, by utilizing... To represent the average reduction of the loss function in each local training round, this invention approximates F(ω). * )for:

[0105]

[0106] Where F0 is the initial loss function. Next, It can be approximated as

[0107] Therefore, problem P1 can be restated as:

[0108]

[0109] The constraints are:

[0110]

[0111] F(ω * )≤F max

[0112] Clearly, P2 is a nonlinear integer programming problem, a typical NP-complete problem with a time complexity of O(n log n). Therefore, the present invention cannot use traditional optimization methods to solve the problem P2. In this invention, a Deep Q-Network Alternating Minimization (DQNAM) method is designed to solve the P2 problem.

[0113] In a typical Fully Flow (FL) model, the number of local training epochs is usually a fixed hyperparameter. However, the required number of local training epochs varies across different FL epochs. Setting a fixed number of local training epochs can lead to unnecessary computational overhead without further improving the training accuracy of the FL model. Furthermore, since aggregation effects and the magnitude of important parameters vary from epoch to epoch, dynamically adjusting the number of parameter pre-synchronization epochs during the FL process is also crucial.

[0114] Therefore, this invention proposes an optimization problem P2 to find the optimal number of local training epochs and parameter pre-synchronization epochs, which together minimize the FL training loss and training time. However, it is a non-convex integer programming problem and cannot be solved by classical optimization methods. Therefore, this invention proposes an alternating minimization algorithm to obtain them through pre-training.

[0115] The basic idea of ​​the Alternating Minimization (AM) algorithm is to alternately optimize one variable while treating another variable as a constant. This optimization operation is then performed alternately on all variables. In the problem of this invention, I and J are first initialized to I... 0 and J 0 Therefore, the optimization objective of the weighted sum of training time and training loss for the FL model is expressed as:

[0116]

[0117] This invention chooses to first optimize J, while simultaneously repairing the parameter pre-synchronization round I. Formally, this invention has...

[0118]

[0119] Then, the present invention further optimizes I, while fixing the local training round J, that is:

[0120]

[0121] For n≥1, the AM algorithm can be formulated as follows:

[0122]

[0123]

[0124] As can be seen from the model of this invention, the interaction process between the BS and the user can be viewed as a Markov decision process. This invention represents the MDP as a tuple. in Representing the state space, Let u represent the action space and u represent the reward. The next step in this invention is to find a strategy that can select the optimal set of actions and maximize cumulative reward in a dynamic environment. It is important to note that this invention clarifies the concepts corresponding to the three core elements of MDP:

[0125] (1) State Space: The state space in the l-th FL loop consists of the loss function set loss={loss0,loss1,...,loss...} m loss M The number of transmitted parameters P = {P1,...,P} m ,...,P M The loss consists of}, where loss0 is the global loss, and loss m It is the loss function for user m, while P m These are the transmission parameters for user m. Then, the state space in the l-th FL loop can be represented as s. l ={loss l ,P l}, and the total state space is

[0126] (2) Action Space: The action space consists of two parts: the user's local training wheel and parameter pre-synchronization wheel Let a l ={I l J l Let} represent the action in the l-th FL cycle. Therefore, the action space can be represented as:

[0127] (3) Reward: The reward function can be used express.

[0128] This invention proposes that the Deep Q-Network (DQN) agent uses the traditional Deep Neural Network (DNN) architecture, and its input and output in the l-th FL loop are the state parameters s. l and the probability of taking the action, a l The actor network performs actions by exploring or exploiting probabilities. The exploitation process is based on the Bellman equation to find the optimal value function Q. * (s,a). To improve the stability of reinforcement learning, DQN employs two improvements: 1) experience replay and 2) a separate target network.

[0129] 1) In each round, the DQN agent will transform the set (s l ,a l ,u l ,sl+1 The data is stored in its replay memory and a mini-batch is randomly drawn to train the DNN. This method alleviates the strong correlation between samples, which also reduces the variance of the updates.

[0130] 2) The target network is cloned from the main network in each C set to minimize the loss using the action values ​​of the target network. This approach prevents overly frequent updates and reduces training divergence and oscillations.

[0131] This invention refers to a weighted neural network function approximator as a Q-network for estimating the action value function:

[0132]

[0133] The descent gradient is (z) k -Q(s k ,a k ;θ)) 2 .

[0134] Therefore, this invention proposes a novel FL framework based on parameter selection and pre-synchronization, which jointly optimizes FL completion time and training loss. To the best of our knowledge, this work is the first to jointly optimize training loss and time consumption through a fine-grained FL framework.

[0135] This invention innovatively proposes selective parameter transmission and parameter pre-synchronization, which allows for the selective transmission and aggregation of some important parameters, thereby reducing transmission overhead and accelerating FL.

[0136] This invention proposes a novel heuristic method, DQNAM, to obtain near-optimal local training rounds and parameter pre-synchronization rounds.

[0137] Comprehensive simulation experiments were conducted to verify the effectiveness of the proposed method on real datasets. The results show that the proposed method outperforms the standard FLACTION (Flexible Training) method, reducing the total FLACTION completion time and training loss by an average of 33.83%.

[0138] In this embodiment, the appendix is ​​now used in conjunction with the attached... Figure 2 The specific embodiments of the present invention are described as follows: Figure 2 A workflow diagram of PSPFL is provided. Unless otherwise specified, the number of central server, BS, and users per BS are 1, 3, and 3, respectively. The batch size for local training is 10, and the optimizer uses SGD. For the FedAvg, FedProx, and FedCS algorithms, the local training rounds are 5. To demonstrate its advantages in communication efficiency and accelerating FL training, PSPFL is compared with several benchmarks:

[0139] 1) FedAvg: The classic FL scheme, in which each terminal device should transmit the entire model parameters, while the BS only transmits the model parameters, and the central server is only responsible for global model aggregation.

[0140] 2) AM: This method uses the same framework as that proposed in this invention (i.e., FL with selective parameter transfer and pre-synchronization). However, unlike the method of this invention that determines the local training round and pre-synchronization round through DQNAM, it only uses the alternating minimization algorithm.

[0141] 3) FedCS: FedCS selects users with high model iteration efficiency to aggregate and update the global model in each FL round, which can optimize the convergence efficiency of FL.

[0142] 4) FedProx: FedProx can dynamically update the local training rounds for different users in each FL round.

[0143] FedProx adds a proximal term to the user's optimization objective function, which makes the optimization algorithm more stable and FL converges faster.

[0144] This invention first changes the transmission rate from 3MB / s to 5MB / s to compare PSPFL and the baseline algorithm. Figure 3 The results show that, across various datasets, the PSPFL proposed in this invention achieves the best results in terms of training loss and FL completion time. In this respect, it consistently outperforms the baseline. Specifically, compared to standard FedAvg, FedCS, FedProx, and AM, PSPFL's performance is superior. The average reductions were 33.83%, 27.27%, 23.18%, and 15.01%, respectively.

[0145] like Figure 3 As shown in (a), compared to FedAvg, FedCS, and FedProx, PSFL's... These represent reductions of 28.92%, 27.98%, and 22.83%, respectively. Both FedAvg and FedProx transmit all model parameters for the user and do not select users. FedProx adds a proximate term for the user, which makes FL converge faster and the loss function smaller. FedCS selects users with better computational and communication resources, thus its FL completion time per round is lower than FedAvg. As a heuristic algorithm, AM can only obtain local optima; therefore, despite considering parameter transmission and parameter pre-synchronization selection, It is still 13.56% larger than PSPFL.

[0146] Figure 3 (b) and Figure 3(c) shows the performance on PSPFL and other benchmarks on Fashion MNIST and CIFAR10. As the complexity of datasets and models increases, The value is greater than that in the MNIST dataset. Furthermore, as the transmission rate increases, the parameter transmission time decreases, and the completion time for each FL round also decreases. The overall trend is downward.

[0147] Figure 4 The global loss function between PSPFL and FedAvg on MNIST with different numbers of business units (BSs) is shown. Each BS has 3 users. It can be seen that as the number of BSs increases, the total number of users also increases, thus the global loss function gradually decreases. Compared to FedAvg, PSPFL has a lower global loss function when the number of BSs and FL rounds are the same. In summary, increasing the number of BSs does not affect the transmission parameters. On the other hand, the more BSs and users, the larger the dataset participating in FL training, and the better the model's training performance, which is consistent with the normal FL architecture.

[0148] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. An accelerated federated learning method based on parameter selection and pre-synchronization, characterized in that, This accelerated federated learning method accelerates training based on parameter selection and a pre-synchronized federated learning framework. The accelerated federated learning method includes the following steps: The central server distributes the latest global model to all base stations, and each base station distributes the received global model to the connected users. The user trains the global model using a local training set and calculates the gradient and model parameters of the received global model. Users will select some model parameters to transmit to the connected base stations, and the base stations will broadcast the model parameters between the base stations via wired links; The model parameters are pre-synchronized until a predetermined number of pre-synchronization rounds are reached. After receiving the aggregation results from all base stations, the central server updates the global model. Specifically, when a user trains and computes the gradients and model parameters of the received global model using a local training round set, the local training round set... ,in, It is the first The number of local training rounds in each federated learning round; after the base station broadcasts the model parameters between base stations via a wired link, the base station is also used to distribute the model parameters to connected users, wherein the model parameters are pre-synchronized and distributed to each federated learning round. The number of pre-synchronization rounds is defined as follows: ,in, Indicates the first The number of pre-synchronization rounds in each federated learning round; During the training of the global model, at the first... The first federal learning round Users in each pre-synchronization round The range of variation of local model parameters during local model training for: , In the formula, Indicates local model parameters; The accelerated federated learning method based on parameter selection and pre-synchronization employs an adaptive saliency threshold. The adaptive significance threshold represents the attenuation threshold during federated learning. Defined as: , in, It is the first The number of parameter pre-synchronization loops in each federated learning loop.

2. The accelerated federated learning method based on parameter selection and pre-synchronization according to claim 1, characterized in that, During the training of the global model, multiple federated learning rounds are performed between the user and the central server. In each federated learning round, ,in It is the total number of federal learning rounds; in the... Users in a federated learning cycle Calculation time for: , In the formula, This represents the number of CPU cycles required for each data sample. This indicates the number of local training iterations. [cycles / s] represents the user Computing resources; Each user A three-element tuple is defined ,in Indicates the magnitude of the model parameters. Represents each user The size of the dataset, and Represents each user Computing resources; a group of users .

3. The accelerated federated learning method based on parameter selection and pre-synchronization according to claim 2, characterized in that, In model parameter pre-synchronization, each user only uploads important parameters to the corresponding base station. Users in the pre-synchronization wheel of parameters Transmission time The calculation is as follows: , in, It is the first User parameters in pre-synchronization The size of the transmission parameters; User The transmission rate; Indicates user and BS Channel bandwidth between; Indicates user The transmission power; User and BS Channel gain between; This represents the standard deviation of the Gaussian channel noise.

4. The accelerated federated learning method based on parameter selection and pre-synchronization according to claim 3, characterized in that, When the predetermined number of pre-synchronization rounds are reached, each user Upload all local model parameters to the corresponding base station. Users in the pre-synchronization wheel of parameters parameter transmission time The calculation is as follows: , Obtain the Federal Learning Wheel The completion time; among which, the user The The completion time for the federated learning round is calculated as follows: , Export the federated learning wheels under the synchronous federated learning settings. Completion time: 。 5. The accelerated federated learning method based on parameter selection and pre-synchronization according to claim 1, characterized in that, The model parameters are pre-synchronized until a predetermined number of pre-synchronization rounds are reached. This also includes jointly minimizing the federated learning training loss and training time to determine the optimal local training rounds and parameter pre-synchronization rounds. Pre-synchronization rounds in federated learning rounds and the number of local training loops Perform initialization; Local training loop count With pre-synchronization wheel rounds The optimization operation is performed alternately, and the alternating minimization algorithm is used to obtain the common minimization of federated learning training loss and training time through pre-training.

6. The accelerated federated learning method based on parameter selection and pre-synchronization according to claim 5, characterized in that, After receiving the aggregated results from all base stations, when the central server updates the global model, the improvements of the Deep Q-Network include experience replay and a separate target network. During experience replay, in each federated learning round, the Deep Q-Network stores the transition set in the replay memory and randomly extracts a set of local training rounds to train the global model. During the separate target network, the target network is cloned by the main network in each transition set, and the loss is minimized using the action values ​​of the target network.