Hybrid aggregation mode federated learning optimization method based on queuing network

By adopting a hybrid aggregation mode based on queuing network in federated learning and dynamically selecting synchronous or asynchronous aggregation mode, the problems of training efficiency and accuracy in traditional federated learning methods are solved, and more efficient and accurate global model training is achieved.

CN119990264APending Publication Date: 2025-05-13NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510232736.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the traditional federated learning method, the synchronous aggregation mode reduces the overall training efficiency due to the large delay of individual clients, while the asynchronous aggregation mode has the problems of large communication overhead and low global model training accuracy.

Method used

A hybrid aggregation mode federated learning optimization method based on queuing network is proposed. By calculating the average time required by the client to complete a local training in synchronous and asynchronous aggregation mode, dynamically selecting the optimal mode to reduce the time spent on a single round of training and improve the convergence speed and training accuracy of the global model.

Benefits of technology

Under reasonable communication overhead, dynamically selecting the optimal aggregation mode, reducing the time-consuming of single-round training, improving the convergence speed and training accuracy of the global model, and solving the problems of training efficiency and accuracy in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990264A_ABST
    Figure CN119990264A_ABST
Patent Text Reader

Abstract

The invention discloses a mixed aggregation mode federated learning optimization method based on a queuing network. The method comprises the following steps: firstly, constructing a federated learning model; k clients are selected to participate in the current round of training, and the average time needed by the selected clients to complete local training in the synchronous aggregation mode is calculated; calculating the average time required for completing local training by the selected client in the asynchronous aggregation mode by using an approximation algorithm of a closed queuing network theory; and comparing the average time in the synchronous and asynchronous aggregation modes, if the former is not greater than the latter and the number of completed training rounds is greater than 80% of the preset total number of training rounds, selecting the synchronous aggregation mode to carry out the training round, otherwise, selecting the asynchronous aggregation mode. And repeating the steps until a preset total training round number is reached. According to the method, under the reasonable communication overhead, the convergence speed and training precision of the global model are effectively improved, meanwhile, the calculation overhead is small, the result is accurate, and the method has wide universality and applicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of federated learning, and in particular to a hybrid aggregation mode federated learning optimization method based on a queuing network. Background Art

[0002] With the popularity of edge devices such as mobile phones and wearable devices in modern society, private data from distributed sources is also growing rapidly. In this digital age, companies are constantly using big data and artificial intelligence to optimize all aspects of life. Although abundant data provides huge opportunities for artificial intelligence applications, most of this data is highly sensitive and exists in the form of isolated islands. It is difficult to centrally manage, and privacy and security issues are becoming increasingly prominent. These many problems restrict the development of artificial intelligence technology. To this end, Brendan McMahan and others from Google proposed federated learning technology. Federated learning is a distributed machine learning method. In federated learning, multiple clients (such as mobile devices, edge devices, etc.) use their own local data to train public models under the guidance of a central server, instead of centralizing data on a central server, thereby protecting user privacy. Traditional federated learning usually adopts a synchronous aggregation mode: after the server sends the global model to all selected clients, the client starts local training. After the local training is completed, the client sends the updated local model parameters to the server. The server waits for all clients to send their local models to the server and then performs weighted average aggregation. However, the synchronous aggregation mode may reduce the overall training efficiency due to the large delay of individual clients. Therefore, Michael R Sprague et al. proposed federated learning in an asynchronous aggregation mode. The main features of the asynchronous aggregation mode are: the server sends the global model to all selected clients, and the client sends the updated local model parameters to the server after local training. The server immediately performs weighted aggregation for each local model parameter sent by a client, and then sends the new global model to the client. This method can improve the update frequency and training efficiency of the model. However, compared with the synchronous aggregation mode, it also has problems such as high communication overhead and low final training accuracy of the global model. Summary of the invention

[0003] In view of the deficiencies in the prior art, the present invention discloses a hybrid aggregation mode federated learning optimization method based on a queuing network to solve the problems raised in the above background technology.

[0004] To achieve the above object, the present invention provides the following technical solution: a hybrid aggregation mode federated learning optimization method based on a queuing network, comprising the following steps:

[0005] S1. Build a federated learning model consisting of a server and M clients, define the neural network model parameter size as Z; select client i, i∈[1,M], and the time it takes to complete a local training is t i , let the sending bandwidth of the i-th client be B i , the communication time to upload the model to the server is The sending bandwidth of the server is B S , the time for distributing the model is The total number of training rounds of the federated learning model is set to R;

[0006] S2. In each round of training, the server randomly selects K clients to form a set G = {c1,…,c K}, together with the server, form N = K + 1 nodes, and the server is denoted as c N =S;

[0007] S3. Calculate the average time T required for the selected client to complete a local training in the synchronous aggregation mode syn ;

[0008] S4. Secondly, use the approximate algorithm of the closed queuing network theory to calculate the average time T required for the selected client to complete a local training in the asynchronous aggregation mode. asyn ;

[0009] S5. Compare the average time required to complete a local training in the synchronous aggregation mode and the asynchronous aggregation mode. If the former is not greater than the latter and the number of completed training rounds is greater than 80% of the preset total number of training rounds, then select the synchronous aggregation mode for this round of training; otherwise, select the asynchronous aggregation mode for this round of training;

[0010] S6. Repeat the above steps until the preset total number of training rounds R is reached.

[0011] Preferably, in step S3, the average time T for the selected client to complete local training and upload the model to the server is syn , the specific steps include:

[0012] For the selected K clients, the total time for each client to complete local training and upload the model to the server is t c1 +d c1 , t c2 +d c2 ,...,t cK +d cK According to the average method, the average time for K clients to complete local training and upload the model is the sum of the total time of each of these clients divided by the number of clients K, that is:

[0013]

[0014] in:

[0015] K represents the number of clients randomly selected by the server to participate in this round of training;

[0016] c j represents the jth client selected, j∈[1,K];

[0017] t cj It represents the time required for the selected j-th client to complete a local training;

[0018] d cj represents the communication time for the selected j-th client to upload the neural network model to the server, and Where Z is the size of the neural network model parameters, B cj is the sending bandwidth of the jth client;

[0019] d S Indicates the time required for the server to distribute a model.

[0020] Preferably, in step S4, an approximate algorithm of closed queuing network theory is used to calculate the average time T required for the selected client to complete a local training in the asynchronous aggregation mode. asyn , specifically including the following steps:

[0021] S4-1. Initialize the average number of visits to each node in the federated learning model

[0022]

[0023] S4-2. Calculate the initial average number of jobs per node

[0024]

[0025] S4-3. Calculate the average response time of each node

[0026]

[0027] S4-4. Calculate the throughput λ:

[0028]

[0029] S4-5. Calculate the current average number of jobs per node

[0030]

[0031] S4-6, initialize the allowable error range ε, set ε = 0.06, if Then jump to S4-7, otherwise jump to S4-8;

[0032] S4-7, Order Then jump to S4-3;

[0033] S4-8. Get the average time required for all currently selected clients to complete a local training in asynchronous aggregation mode

[0034] Preferably, in step S5, if T syn ≤Y asyn Or if the number of completed training rounds is greater than 80% of the preset total number of training rounds R, a round of synchronous federated training is performed, which includes the following steps:

[0035] The server will use the global model w G Sent to all selected clients; client c j ∈G starts local training and updates the local model after local training. Sent to the server; after receiving the local models sent by all clients, the server aggregates them. The aggregation formula is: Get a new global model.

[0036] Preferably, in step S5, if T syn ≤T asyn Or if the number of completed training rounds is greater than 80% of the preset total number of training rounds R, a round of asynchronous federated training is performed, which includes the following steps:

[0037] The server will use the global model w G Sent to all selected clients; client c j ∈G starts local training and updates the local model after local training. Send to the server; the server immediately aggregates each local model sent by the client. , and its aggregation formula is:

[0038]

[0039] Among them, α is the weight of model aggregation. When the server aggregates for the first time, α=1, and in other cases α=0.5. Then the new global model is sent to the client; until time After that, this round of training ends, and the time

[0040] Compared with the prior art, the present invention has the following beneficial effects:

[0041] 1. The present invention proposes a federated learning optimization method of a hybrid aggregation mode, which aims to combine the advantages of synchronous and asynchronous aggregation modes, dynamically select the optimal mode under reasonable communication overhead, reduce the time consumption of a single round of training, and improve the convergence speed and training accuracy of the global model.

[0042] 2. The present invention adopts the approximate algorithm of closed queuing network theory to calculate the average time required for all currently selected clients to complete a local training in asynchronous aggregation mode, which can obtain more accurate calculation results with less computing overhead and has wide versatility and applicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.

[0044] In the attached picture:

[0045] Figure 1 It is a flow chart of the hybrid aggregation mode federated learning optimization method based on queuing network of the present invention. DETAILED DESCRIPTION

[0046] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0047] Example: Figure 1 As shown, a hybrid aggregation mode federated learning optimization method based on a queuing network specifically includes the following steps:

[0048] Step 1: First, build a federated learning model consisting of a server and M clients. Let Z be the size of the neural network model parameters. For client i, i∈[1,M], t i represents the time required for the i-th client to complete a local training. Let the sending bandwidth of the i-th client be B i , so the communication time for uploading the neural network model to the server is Assume the sending bandwidth of the server is B S , then the time required for the server to distribute a model is Considering that the processing capacity of the server is much greater than that of the client, the present invention ignores the time of the server aggregation model. In addition, the total number of training rounds of the above federated learning model is R.

[0049] Step 2: In each round of training, the server randomly selects K clients to participate in this round of training. Let the selected client set be G, G = {c1,…cj …,c K},c j ∈[1,M]; these K clients plus the server have a total of K+1 nodes, let N=K+1, and denote the server as c N =S;

[0050] Step 3: Calculate the average time T required for all currently selected clients to complete a local training in the synchronous aggregation mode syn For the selected K clients, the total time for each client to complete local training and upload the model to the server is t c1 +d c1 , t c2 +d c2 ,...,t cK +d cK According to the average method, the average time for K clients to complete local training and upload the model is the sum of the total time of each of these clients divided by the number of clients K, that is:

[0051]

[0052] in:

[0053] K represents the number of clients randomly selected by the server to participate in this round of training;

[0054] c j represents the jth client selected, j∈[1,K];

[0055] t cj It represents the time required for the selected j-th client to complete a local training;

[0056] d cj represents the communication time for the selected j-th client to upload the neural network model to the server, and Where Z is the size of the neural network model parameters, B cj is the sending bandwidth of the jth client;

[0057] d S Indicates the time required for the server to distribute a model.

[0058] Step 4: Use the approximate algorithm of the closed queuing network theory to calculate the average time T required for all currently selected clients to complete a local training in the asynchronous aggregation mode. asyn The specific process is shown in steps 4-1 to 4-8:

[0059] Step 4-1: Initialize the average number of visits to each node in the federated learning model

[0060]

[0061] Step 4-2: Calculate the initial average number of jobs per node

[0062]

[0063] Step 4-3: Calculate the average response time of each node

[0064]

[0065] Step 4-4, calculate the throughput λ:

[0066]

[0067] Step 4-5: Calculate the current average number of jobs per node

[0068]

[0069] Step 4-6: Initialize the allowable error range ε, set ε = 0.06, if If yes, jump to step 4-7, otherwise jump to step 4-8;

[0070] Step 4-7, command Then jump to step 4-3;

[0071] Step 4-8: Get the average time T required for all currently selected clients to complete a local training in asynchronous aggregation mode asyn :

[0072]

[0073] Step 5: If T syn ≤T asyn Or if the number of completed training rounds is greater than 80% of the preset total number of training rounds R, jump to step 6; otherwise, jump to step 7;

[0074] Step 6: Perform a round of synchronous federated training, that is, the server will use the global model w G Sent to all selected clients; client c j ∈G starts local training and updates the local model after local training. Sent to the server; after receiving the local models sent by all clients, the server aggregates them. The aggregation formula is: Get the new global model, and jump to step 8 after completion;

[0075] Step 7: Perform a round of asynchronous federated training, that is, the server will use the global model w G Sent to all selected clients; client c j ∈G starts local training and updates the local model after local training. Sent to the server; the server aggregates each local model sent by a client immediately, such as , and its aggregation formula is:

[0076]

[0077] Among them, α is the weight of model aggregation. When the server aggregates for the first time, α=1, and in other cases α=0.5. Then the new global model is sent to the client; until time After that, this round of training ends and jumps to step 8, where the time The calculation method is

[0078] Step 8: If the preset total number of training rounds R has been reached, the training ends; otherwise, jump to step 2 to enter the next round of training.

[0079] In order to verify that the federated learning optimization method based on the hybrid aggregation mode of the queuing network of the present invention can effectively speed up the training speed of the federated learning model, a verification example is listed for comparison and explanation:

[0080] In this verification example, the federated learning system consists of a server and M = 100 clients, the neural network model parameter size Z = 1.46 (MB), and the total number of training rounds is preset to R = 500;

[0081] In this verification example, the number of clients selected for training is K = 5, the set of selected clients is G, G = {c1, c2, c3, c4, c5}, and the local training time of the selected clients is They are: 1.0, 4.0, 3.0, 1.0, 3.0 (seconds), the sending bandwidth of each client They are: 3.0, 6.0, 5.0, 4.0, 7.0 (Mbps), then the time required for each client to send the local model to the server They are 3.89, 1.95, 2.34, 2.92, and 1.67 (seconds) respectively. The current server's sending bandwidth is B S =5.0Mbps, the time required for the server to send a model is At this time, there are K+1 nodes in the federated learning model, let N=K+1, and record the server as c N =S;

[0082] Calculate the average time required for all currently selected clients to complete a local training in synchronous aggregation mode:

[0083]

[0084] The average time T required for all currently selected clients to complete a local training in asynchronous aggregation mode is calculated using the approximate algorithm of closed queuing network theory. asyn :

[0085] Initialize the average number of visits to each node of the federated learning model They are: 1.0, 1.0, 1.0, 1.0, 1.0, 5.0;

[0086] Calculate the initial average number of jobs per node

[0087]

[0088] Calculate the average response time for each node

[0089]

[0090] Similarly:

[0091] They are 9.92, 8.90, 6.53, 7.78 (seconds) respectively;

[0092]

[0093] Calculate the throughput λ:

[0094]

[0095] Calculate the current average number of jobs per node

[0096]

[0097] Similarly:

[0098] They are 0.82, 0.73, 0.54, 0.64, and 1.61 respectively;

[0099] Assume the allowable error range ε ​​= 0.06, make Repeat the above calculation steps until the result is within the error range. At this time, the average response time of each node is They are; 6.87, 9.17, 7.79, 5.09, 6.44, 6.83 (seconds).

[0100] The average time required for the client to complete a local training in asynchronous aggregation mode is

[0101] At this time T asyn <T syn The number of completed training rounds is less than 80% of the preset total number of training rounds. Therefore, the asynchronous aggregation mode is selected until the time After that, this round of training ends and the client is reselected for the next round of training. The calculation method is as follows:

[0102]

[0103] Reselect K = 5 clients to start training, and use the same calculation method to get T syn =11.92 (seconds), T asyn =23.48 (seconds).

[0104] At this time T asyn >T syn ,Therefore, the synchronous aggregation mode is selected. After the current round of training is completed, the client is reselected for the next round of training;

[0105] Repeat the above steps until the number of training rounds reaches the preset total number of training rounds.

[0106] In the simulation experiment of this method, the experimental environment used is as follows: the experiment is carried out under the CIFAR-10 data set, and a federated learning model is composed of a server and 100 clients. The parameter size of the neural network model is 1.46 (MB), the bandwidth of the server is 5.0 (Mbps), the training time of the client is in the range of 1 to 30 (seconds), and the bandwidth of the client is in the range of 0.5 to 10.0 (Mbps). Each round of training is completed by the server randomly selecting K = 5 clients to participate, and the total number of training rounds is preset to R = 500. Table 1 shows the performance comparison of this method with the synchronous and asynchronous aggregation modes:

[0107] Table 1 Comparison of the accuracy of global models under different aggregation modes

[0108]

[0109] It can be seen from Table 1 that, compared with the synchronous and asynchronous aggregation modes, the hybrid aggregation mode proposed in the present invention can obtain the highest global model accuracy during the training process, thereby verifying the effectiveness of the present method.

[0110] In summary, the present invention proposes an optimization method for hybrid aggregation mode in federated learning. By calculating the average time required to complete a local training under different aggregation modes and combining the ratio of the number of completed training rounds to the total number of preset training rounds, the current aggregation mode is determined. Under reasonable communication overhead, the convergence speed and training accuracy of the global model are effectively improved, and more accurate calculation results can be obtained with less computing overhead. It has wide versatility and applicability.

[0111] Finally, it should be noted that the above description is only a preferred example of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A hybrid aggregation mode federated learning optimization method based on queuing network, characterized in that: The following steps are involved: S1. Build a federated learning model consisting of a server and M clients, define the neural network model parameter size as Z; select client i, i∈[1,M], and the time it takes to complete a local training is t i , let the sending bandwidth of the i-th client be B i , the communication time to upload the model to the server is The sending bandwidth of the server is B S , the time for distributing the model is The total number of training rounds of the federated learning model is set to R; S2. In each round of training, the server randomly selects K clients to form a set G = {c1,…,c K }, together with the server, form N = K + 1 nodes, and the server is denoted as c N =S; S3. Calculate the average time T required for the selected client to complete a local training in the synchronous aggregation mode syn ; S4. Secondly, use the approximate algorithm of the closed queuing network theory to calculate the average time T required for the selected client to complete a local training in the asynchronous aggregation mode. asyn ; S5. Compare the average time required to complete a local training in the synchronous aggregation mode and the asynchronous aggregation mode. If the former is not greater than the latter and the number of completed training rounds is greater than 80% of the preset total number of training rounds, then select the synchronous aggregation mode for this round of training; otherwise, select the asynchronous aggregation mode for this round of training; S6. Repeat the above steps until the preset total number of training rounds R is reached.

2. The hybrid aggregation mode federated learning optimization method based on queuing network according to claim 1 is characterized by: In step S3, the average time T for the selected client to complete local training and upload the model to the server syn , the specific steps include: For the selected K clients, the total time for each client to complete local training and upload the model to the server is t c1 +d c1 , t c2 +d c2 ,...,t cK +d cK According to the average method, the average time for K clients to complete local training and upload the model is the sum of the total time of each of these clients divided by the number of clients K, that is: in: K represents the number of clients randomly selected by the server to participate in this round of training; c j represents the jth client selected, j∈[1,K]; t cj It represents the time required for the selected j-th client to complete a local training; d cj represents the communication time for the selected j-th client to upload the neural network model to the server, and Where Z is the size of the neural network model parameters, B cj is the sending bandwidth of the jth client; d S Indicates the time required for the server to distribute a model.

3. The hybrid aggregation mode federated learning optimization method based on queuing network according to claim 1 is characterized by: In step S4, the average time T required for the selected client to complete a local training in the asynchronous aggregation mode is calculated using the approximate algorithm of the closed queuing network theory. asyn , specifically including the following steps: S4-1. Initialize the average number of visits to each node in the federated learning model S4-2. Calculate the initial average number of jobs per node S4-3. Calculate the average response time of each node S4-4. Calculate the throughput λ: S4-5. Calculate the current average number of jobs per node S4-6, initialize the allowable error range ε, set ε = 0.06, if Then jump to S4-7, otherwise jump to S4-8; S4-7, Order Then jump to S4-3; S4-8. Get the average time required for all currently selected clients to complete a local training in asynchronous aggregation mode 4. The hybrid aggregation mode federated learning optimization method based on a queuing network according to claim 1 is characterized in that: In step S5, if T syn ≤T asyn Or if the number of completed training rounds is greater than 80% of the preset total number of training rounds R, a round of synchronous federated training is performed, which includes the following steps: The server will use the global model w G Sent to all selected clients; client c j ∈G starts local training and updates the local model after local training. Sent to the server; after receiving the local models sent by all clients, the server aggregates them. The aggregation formula is: Get a new global model.

5. The hybrid aggregation mode federated learning optimization method based on queuing network according to claim 4 is characterized by: In step S5, if T syn ≤T asyn Or if the number of completed training rounds is greater than 80% of the preset total number of training rounds R, a round of asynchronous federated training is performed, which includes the following steps: The server will use the global model w G Sent to all selected clients; client c j ∈G starts local training and updates the local model after local training. Send to the server; the server immediately aggregates each local model sent by the client. Its aggregation formula is: Among them, α is the weight of model aggregation. When the server aggregates for the first time, α=1, and in other cases α=0.

5. Then the new global model is sent to the client; until time After that, this round of training ends, and the time