Asynchronous federated learning optimization method and system

Through collaborative training between the central server and client, the number of training layers and layers is dynamically adjusted, which solves the problem of model training efficiency decline caused by client heterogeneity, and improves the training efficiency and effect of heterogeneous devices in federated learning.

CN120297435APending Publication Date: 2025-07-11BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510184068.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The problem of model training efficiency decline caused by client heterogeneity exists in federated learning, especially when the resources and computing power differences between clients are large, which affects training speed and efficiency.

Method used

The central server initializes the global model and layer index table. The client trains and sends model parameters and training time according to the layer index table. The central server determines the importance of the layer based on the model parameters, convergence index and obsoleteness index, calculates the slicing ratio, updates the layer index table and sends it to the client, and the client trains based on the updated layer index table.

Benefits of technology

Improve the training efficiency and effect of heterogeneous clients in federated learning, adapt to the processing capabilities of different devices, and optimize the training process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297435A_ABST
    Figure CN120297435A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an asynchronous federated learning optimization method and system, and the method comprises the steps: a central server initializes a global model and a layer index table, and transmits the global model and the layer index table to each client; the client trains all the layers according to the layer index table to obtain updated model parameters of all the layers, and the model parameters and local training time are sent to the central server; the central server adaptively selects training layers and the number of the layers of each client according to the importance degree of each layer in each round of training and the segmentation ratio of the client, updates a layer index table, updates a global model according to model parameters, and sends the updated global model and the layer index table to the client; and based on the updated global model, the client trains the corresponding layer in the table according to the updated layer index table to obtain the updated model parameters of the corresponding layer and the local training time of the round. The method can adapt to the processing capability of the heterogeneous client, improve the training efficiency and effect of the heterogeneous device in federal learning, and optimize the training process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of machine learning, and in particular, to an asynchronous federated learning optimization method and system. Background Art

[0002] Federated learning is a distributed machine learning method that allows multiple clients to train models based on local datasets. Each client only needs to upload the parameters of the local model, and the central server generates a global model by aggregating the local models. During the federated learning process, there are differences among the clients participating in model training in terms of hardware, computing resources, etc. The heterogeneity of client devices will affect the efficiency and performance of federated learning. Especially when there are large differences in resources and computing power among clients, it will lead to a decline in the speed and efficiency of training the model. Summary of the Invention

[0003] In view of this, the purpose of the embodiments of the present application is to propose an asynchronous federated learning optimization method and system to solve the problem of the decline in model training efficiency caused by client heterogeneity.

[0004] Based on the above purpose, the embodiments of the present application provide an asynchronous federated learning optimization method, including:

[0005] The central server initializes the global model, generates an initial layer index table, and sends the global model and the initial layer index table to each client; the initial layer index table includes the layer information of all layers of the global model;

[0006] The client receives the global model and the layer index table, trains all layers according to the layer index table, obtains the updated model parameters of each layer, and sends the updated model parameters of each layer and the local training time to the central server;

[0007] The central server receives the model parameters and local training time of each client, updates the global model according to the model parameters, determines the importance degree of each layer according to the number of parameters, convergence index, and obsolescence index of each layer, calculates the splitting ratio according to the local training time, and updates the layer index table of each client according to the importance degree and splitting ratio of each layer, and sends the updated global model and the updated layer index table to the corresponding client; wherein, the convergence index of each layer is calculated according to the gradient of each layer, and the obsolescence index is determined according to the time difference between two adjacent updates of each layer;

[0008] The client receives the updated global model and the updated layer index table, and based on the updated global model, trains the corresponding layer in the table according to the updated layer index table to obtain the updated model parameters of the corresponding layer and the local training time of this round.

[0009] Optionally, the method further includes:

[0010] The client sends the model parameters of the corresponding layer after the update and the local training time of this round to the central server. The central server updates the global model according to the model parameters of the corresponding layer, determines the importance level of each layer according to the number of parameters, convergence index, and obsolescence index of each layer, calculates the splitting ratio according to the local training time of this round, updates the layer index table of each client according to the importance level of each layer and this splitting ratio, and sends the updated global model and the updated layer index table to the corresponding client; repeat the above process until the global model meets the preset training end condition.

[0011] Optionally, the client trains the corresponding layer in the table according to the layer index table to obtain the model parameters of the updated corresponding layer. The method is as follows:

[0012]

[0013] where is the updated model parameter of the l-th layer of the local model in the (k + 1)-th round of training for client l, γ is the learning rate during local model training, and L is the loss function. is the model parameter of the l-th layer of the local model for client i in the k-th round of training.

[0014] Optionally, the central server updates the global model according to the model parameters. The method is as follows:

[0015]

[0016] where m is the number of clients, Ⅱ(l ∈ layerIndex i ) is the indicator function, which takes the value of 1 when l ∈ layerIndex i and 0 otherwise; is the model parameter of the l-th layer in the k-th round of training for the central server, is the model parameter of the l-th layer in the (k + 1)-th round of training for central server s, ω i is the updated weight of client i in this aggregation.

[0017] Optionally, the method for calculating the convergence index is as follows:

[0018]

[0019] where is the gradient of the l-th layer of the local model, τ is the window size for measuring the gradient convergence situation, and j is the time from t - τ to t.

[0020] Optionally, determine the importance level of each layer according to the number of parameters, convergence index, and obsolescence index of each layer. The method is as follows:

[0021]

[0022] Among them, α, β, and δ are weights. is the number of parameters of the l-th layer. is the obsolescence index of the l-th layer.

[0023] Optionally, calculate the split ratio according to the local training time. The method is as follows:

[0024]

[0025] Among them, T i is the time required for client i to train the model in this round. T min is the minimum value of the training time of all clients in this round of model training. T max is the maximum value of the training time of all clients in this round of model training.

[0026] Optionally, update the layer index table of each client according to the importance level of each layer and the split ratio, including:

[0027] Sort each layer in descending order of importance to obtain the sorted layers.

[0028] Determine the target number of layers for each client to be selected from the sorted layers according to the split ratio of the client and the total number of layers of the global model.

[0029] Based on the sorted layers, select the layers with the target number of layers from front to back, and update the layer index table of each client according to the selected layers.

[0030] Optionally, determine the target number of layers to be selected from the sorted layers according to the split ratio and the total number of layers of the global model. The method is as follows:

[0031]

[0032] Among them, L is the total number of layers of the global model, and A is the target number of layers to be selected.

[0033] The embodiment of the present application also provides an asynchronous federated learning optimization system, including:

[0034] A central server, which is used to initialize a global model, generate an initial layer index table, and send the global model and the initial layer index table to each client; the initial layer index table includes layer information of all layers of the global model; and receive model parameters and local training times of each client, update the global model according to the model parameters, determine the importance level of each layer according to the number of parameters, convergence index, and obsolescence index of each layer, calculate a segmentation ratio according to the local training time, update the layer index table of each client according to the importance level and segmentation ratio of each layer, and send the updated global model and the updated layer index table to the corresponding client; wherein, the convergence index of each layer is calculated according to the gradient of each layer, and the obsolescence index is determined according to the time difference between two adjacent updates of each layer;

[0035] A client is used to receive the global model and the initial layer index table, train all layers according to the layer index table to obtain updated model parameters of each layer, and send the updated model parameters of each layer and the local training time to the central server; and receive the updated global model and the updated layer index table, and based on the updated global model, train the corresponding layers in the table according to the updated layer index table to obtain updated model parameters of the corresponding layers and the local training time of this round.

[0036] As can be seen from the above, in the asynchronous federated learning optimization method and system provided by the embodiments of the present application, the central server initializes the global model and the layer index table and sends them to each client; the client trains all layers according to the layer index table to obtain updated model parameters of each layer, and sends the model parameters and the local training time to the central server; the central server adaptively selects the training layers and the number of layers of each client according to the importance level of each layer in each round of training and the segmentation ratio of the client, updates the layer index table, updates the global model according to the model parameters, and sends the updated global model and the layer index table to the client; the client trains the corresponding layers in the table according to the updated layer index table based on the updated global model to obtain updated model parameters of the corresponding layers and the local training time of this round. The present application can adapt to the processing capabilities of heterogeneous clients, improve the training efficiency and effect of heterogeneous devices in federated learning, and optimize the training process. Description of the Drawings

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0038] Figure 1 It is a schematic flowchart of the method of the embodiment of the present application;

[0039] Figure 2 It is the system structure block diagram of the embodiment of the present application;

[0040] Figure 3 It is the structure block diagram of the electronic device of the embodiment of the present application. Specific embodiments

[0041] To make the purpose, technical solutions and advantages of the present disclosure clearer, the following further describes the present disclosure in detail with reference to specific embodiments and the accompanying drawings.

[0042] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure belongs. The "first", "second" and similar terms used in the embodiments of the present application do not indicate any order, quantity or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left" and "right" are only used to represent relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0043] Federated learning is a distributed machine learning method in which multiple clients cooperate to complete the training of a global model under the coordination of a central server. Each client has its own local dataset. After training the initial global model using the local dataset, the client obtains a local model and only needs to upload the model parameters (such as weights, biases, gradients, etc.) of the local model to the central server. The central server aggregates and distributes the model parameters of the local models of each client. Since the local data of the clients is not shared, the security of the user's private data can be guaranteed. The general process of the federated learning algorithm is as follows: The central server initializes the global model according to a specific goal and sends the global model to each client. The client trains the initial global model based on the local dataset to obtain a local model and uploads the model parameters of the local model to the central server; the central server aggregates the model parameters of the local models of each client to obtain an aggregated model.

[0044] Hereinafter, the technical solutions of the present application will be further described in detail through specific embodiments.

[0045] As Figure 1 、 2 shown, the embodiment of the present application provides an asynchronous federated learning optimization method, including:

[0046] S101: The central server initializes the global model, generates an initial layer index table, and sends the global model and the initial layer index table to each client; the initial layer index table includes layer information of all layers of the global model.

[0047] In this embodiment, the central server initializes the global model according to a specific goal, uses a model slicing method to slice the global model into multiple layers, generates an initial layer index table according to all layers, and this layer index table includes layer information such as the layer name, number, the name of the model parameters included, and the corresponding model parameter values of all layers. Then, the initialized global model and the layer index table are sent to all clients.

[0048] In some ways, according to the application scenario of the global model, the global model may include a pooling layer, a fully connected layer, an activation layer, etc. The global model can be divided into multiple layers according to the composition levels, and then a layer index table is generated according to all layers. For example: the global model includes two fully connected layers fc1 and fc2, and the information items included in the layer index table include [fc1.weight, fc1.bias, fc2.weight, fc2.bias]. Among them, fc1.weight represents the name of the weight parameter of the fully connected layer fc1, and the weight parameter of the fully connected layer fc1 can be obtained according to this parameter. fc1.bias represents the name of the bias parameter of the fully connected layer fc1, and the bias parameter of the fully connected layer fc1 can be obtained according to this parameter.

[0049] S102: The client receives the global model and the initial layer index table, trains all layers according to this layer index table, obtains the updated model parameters of each layer, and sends the updated model parameters of each layer and the local training time to the central server.

[0050] In this embodiment, the client receives the global model and the initial layer index table. Since this layer index table includes all layers of the global model, the client needs to use the local dataset to perform the first round of training on all layers of the global model. After training, the first-round local model and the updated model parameters of all layers are obtained. Then, the updated model parameters of all layers and the local training time used for the first-round training are sent to the central server.

[0051] In some embodiments, before starting training, the client traverses each layer of the global model. For the currently traversed layer, if the current layer is recorded in the layer index table, mark this layer as a training layer; if not recorded in the layer index table, mark it as a non-training layer. After traversing, the client performs model training on all training layers. After training, the updated model parameters of each training layer are obtained, and there is no need to train the non-training layers, and no gradient calculation and parameter update will be performed on the non-training layers.

[0052] In some ways, the client trains the corresponding layer in the layer index table according to the layer index table, that is, trains each training layer to obtain updated model parameters of the corresponding layer. The method is as follows:

[0053]

[0054] Among them, is the updated model parameter of the l-th layer of the local model by client i in the (k + 1)-th round of training. γ is the learning rate during local model training, and L is the loss function. is the model parameter of the l-th layer of the local model by client i in the k-th round of training. The l-th layer is the training layer of client i.

[0055] S103: The central server receives the model parameters and local training times of each client, updates the global model according to the model parameters, determines the importance of each layer according to the number of parameters, convergence index, and obsolescence index of each layer, calculates the splitting ratio according to the local training time, and updates the layer index table of each client according to the importance of each layer and the splitting ratio. Then, the updated global model and the updated layer index table are sent to the corresponding clients. Among them, the convergence index of each layer is calculated according to the gradient of each layer, and the obsolescence index is determined according to the time difference between two adjacent updates of each layer.

[0056] In this embodiment, after each client trains to obtain the local model and the model parameters of each layer according to the layer index table, the model parameters and the local training time used by the client in this round of model training are sent to the central server, so that the central server updates the global model according to the model parameters of each client.

[0057] The method for the central server to update the global model when receiving the model parameters uploaded by m clients at time t is as follows:

[0058]

[0059] Among them, Ⅱ(l ∈ layerIndex i ) is an indicator function, which takes the value of 1 when l ∈ layerIndex i , and 0 otherwise. That is, when the l-th layer is the training layer of client i, the value of the indicator function is 1, and when the l-th layer is the non-training layer of client i, the value of the indicator function is 0. is the model parameter of the l-th layer by the central server in the k-th round of training. is the model parameter of the l-th layer by the central server s in the (k + 1)-th round of training, that is, the updated model parameter.

[0060] ω iis the updated weight of client i in this aggregation, which to a certain extent represents the contribution of client i's update to the global model in this round. Its calculation method is as follows:

[0061]

[0062] where n i is the data volume of the local dataset of client i, and N is the total data volume of the local datasets of all clients. is the freshness of client i on the l-th layer of the local model. σ is a hyperparameter that controls the influence degree of the local training time of the client. T is the current time of this update, and T i l is the time when client i last updated the l-th layer.

[0063] In some ways, during the asynchronous federated learning process, when the central server receives the model parameters sent by some clients, even if there are still some clients that have not completed the training of the training layer, the central server will update the global model according to the received partial model parameters. That is, when a client completes local training, the central server immediately updates the global model without waiting for all clients to complete training.

[0064] In some embodiments, the central server calculates the convergence index of each layer according to the gradient of each layer, calculates the obsolescence index according to the time difference between two adjacent updates of each layer of the model, and determines the importance degree of each layer according to the number of parameters, convergence index, and obsolescence index of each layer.

[0065] In some ways, the central server quantifies the convergence of the layer gradient according to the direction of the layer gradient update. The method for calculating the convergence index of each layer is as follows:

[0066]

[0067] where is the gradient of the l-th layer of the local model (which is one of the model parameters), j is the time from t - τ to t; τ is the window size used to measure the gradient convergence, which is a hyperparameter representing the gradients of the previous τ model updates. For example, when τ takes the value of 1, the convergence index of this layer is calculated according to the previous update and the current update of this layer.

[0068] The method for calculating the obsolescence index of each layer is as follows:

[0069]

[0070] where T is the current time of this update, is the time when client i last updated the l-th layer, and ΔT l is the time difference between two adjacent updates of the l-th layer.

[0071] The maximum - minimum normalization method is used to normalize ΔT l to obtain the obsolescence index The method is as follows:

[0072]

[0073] where L is the total number of layers of the global model, is the minimum value of ΔT in the L - th layer l and is the maximum value of ΔT in the L - th layer is the maximum value of ΔT in the L - th layer l and is the maximum value of ΔT in the L - th layer

[0074] The method for calculating the number of parameters of each layer is as follows:

[0075]

[0076] where Q l is the number of parameters of the l - th layer (for example, when the parameters of a certain layer include a weight matrix of size m×n and a bias vector of size q, the number of parameters is (m×n + q)), Q min is the minimum value of the number of parameters in all layers of the model, Q max is the maximum value of the number of parameters in all layers of the model, is the normalized number of parameters of the l - th layer

[0077] The central server calculates the importance degree of each layer according to the calculated number of parameters, convergence index, and obsolescence index of each layer. The method is as follows:

[0078]

[0079] where α, β, and δ are the weights of the number of parameters, obsolescence index, and convergence index of the l - th layer respectively

[0080] The central server calculates the splitting ratio of the client according to the local training time used by the client for local model training. The method is as follows:

[0081]

[0082] where μ i is the splitting ratio of client i, T i is the local training time used by client i in this round of model training, T min is the minimum value of the local training time of all clients in this round of model training, T max is the maximum value of the local training time of all clients in this round of model training

[0083] In some embodiments, the central server updates the layer index tables of each client according to the importance levels of each layer and the segmentation ratios of each client, including:

[0084] Sort each layer in descending order of importance level to obtain the sorted layers.

[0085] Based on the segmentation ratio of the client and the total number of layers of the global model, determine the target number of layers for each client to be selected from the sorted layers.

[0086] Based on the sorted layers, select the layers of the target number of layers from front to back, and update the layer index tables of each client according to the selected layers.

[0087] In this embodiment, after determining the importance levels of all layers of the model, sort all layers in descending order of importance level to obtain the sorted all layers. For each client, based on the segmentation ratio of the client and the total number of layers of the global model, determine the target number of layers that the client needs to process. Based on the sorted all layers, select the layers of the target number of layers from front to back, and generate the layer index table of the client according to the selected layers. That is, after each round of model training, the central server dynamically updates the layer index table of the client according to the importance levels of each layer of the model and the segmentation ratio of the client.

[0088] In some ways, based on the segmentation ratio of the client and the total number of layers of the global model, determine the target number of layers to be selected from the sorted layers. The method is:

[0089]

[0090] Among them, L is the total number of layers of the global model, and A is the target number of layers. For example, for the client with the strongest training ability, its local training time T min , then its segmentation ratio μ i is 1, and the target number of layers A that can be allocated to this client is A = L, that is, use this client to train all layers.

[0091] When the central server receives the local model and model parameters of a certain client, it immediately updates the global model according to the received model parameters, and updates the layer index table of this client according to the calculated segmentation ratio of this client and the importance levels of each layer, and sends the updated global model and layer index table to this client, without waiting for all other clients to upload parameters before updating.

[0092] S104: The client receives the updated global model and the updated layer index table, and based on the updated global model, trains the corresponding layers in the table according to the updated layer index table to obtain the updated model parameters of the corresponding layers and the local training time of this round.

[0093] In this embodiment, after the client receives the updated global model and layer index table, it traverses each layer of the updated global model, determines whether the currently traversed layer is recorded in the updated layer index table. If it is recorded, the layer is marked as a training layer; if not, it is marked as a non-training layer. After traversal, the client trains the model for all training layers, and after training, obtains the updated model parameters of each training layer and the local training time used in this round of model training.

[0094] In some embodiments, the asynchronous federated learning optimization method further includes:

[0095] The client sends the updated model parameters of the corresponding layer and the local training time of this round to the central server; the central server updates the global model according to the model parameters of the corresponding layer, determines the importance of each layer according to the number of parameters, convergence index, and obsolescence index of each layer, calculates the splitting ratio according to the local training time of this round, and updates the layer index table of each client according to the importance of each layer and this splitting ratio, and sends the updated global model and the updated layer index table to the corresponding client; repeat the above process until the global model meets the preset training end condition.

[0096] In this embodiment, the central server and each client complete the training process of the global model through multiple rounds of collaborative training. Specifically, after the central server initializes the global model and the initial layer index table, it distributes the global model and the layer index table to all clients. Each client uses the local dataset and trains all layers of the global model according to the initial layer index table, obtains the model parameters of the first round of model training, and records the local training time of the first round of model training. Each client sends the model parameters and local training time of the first round of model training to the central server.

[0097] The central server updates the global model according to the model parameters of some clients received, calculates the importance of each layer according to the number of parameters, obsolescence index, and convergence index of each layer, calculates the splitting ratio according to the local training time used by some clients, determines the training layers of some clients in the next round of training according to the importance of each layer and the splitting ratio of the client, that is, performs dynamic model splitting by combining the training results of each client in each round, dynamically divides the layers required for the next round of training of each client, generates the layer index table for the next round based on the re-divided layers, sends the updated global model to the corresponding client, and distributes the updated layer index table to the corresponding client.

[0098] The client receives the updated global model and layer index table, performs model training for the corresponding layer of the updated global model according to the updated layer index table, obtains the model parameters of this round, and records the local training time of this round, and sends the model parameters and local training time of this round to the central server.

[0099] The central server updates the global model according to the model parameters of some clients in this round, calculates the importance degree of each layer according to the number of parameters, obsolescence index and convergence index of each layer in this round, calculates the splitting ratio according to the local training time used by some clients in this round, determines the training layers of some clients in the next round of training according to the importance degree of each layer and the splitting ratio of the clients, generates the layer index table for the next round, and distributes the updated global model and the updated layer index table to the corresponding clients. By repeating the above training process, the central server and each client complete the training of the global model through collaborative training in multiple rounds.

[0100] In the asynchronous federated learning optimization method provided in the embodiments of the present application, during the collaborative training process between the central server and the clients, the central server adaptively selects the training layers of each client according to the importance degree of each layer in each round of training and the splitting ratio of the clients, and dynamically adjusts the layers and the number of layers to be trained in combination with the computing power of the clients. For example, clients with relatively weak computing power are used to train a small number of key layers with relatively high importance degree to improve the optimization quality of the key layers and enhance the contribution to the global model. The computing power of the clients is measured by the splitting ratio, and clients with strong computing power are used to train most of the key layers and general layers, which can adapt to the processing capabilities of heterogeneous clients, improve the training efficiency and effect of heterogeneous devices in federated learning, and optimize the training process.

[0101] It should be noted that the method in the embodiments of the present application can be executed by a single device, such as a computer or a server. The method in this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In such a distributed scenario, one of the multiple devices can only execute one or more steps in the method in the embodiments of the present application, and these multiple devices will interact with each other to complete the described method.

[0102] It should be noted that the above specifically describes certain embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0103] As Figure 2 shown, the embodiments of the present application provide an asynchronous federated learning optimization system, including:

[0104] A central server, which is used to initialize a global model, generate an initial layer index table, and send the global model and the initial layer index table to each client; the initial layer index table includes layer information of all layers of the global model; and receive model parameters and local training time of each client, update the global model according to the model parameters, determine the importance degree of each layer according to the number of parameters, convergence index, and obsolescence index of each layer, calculate a segmentation ratio according to the local training time, and update the layer index table of each client according to the importance degree and segmentation ratio of each layer, and send the updated global model and the updated layer index table to the corresponding client; wherein, the convergence index of each layer is calculated according to the gradient of each layer, and the obsolescence index is determined according to the time difference between two adjacent rounds of model training;

[0105] A client is used to receive the global model and the initial layer index table, train all layers according to the layer index table to obtain updated model parameters of each layer, and send the updated model parameters of each layer and the local training time to the central server; and receive the updated global model and the updated layer index table, and based on the updated global model, train the corresponding layers in the table according to the updated layer index table to obtain updated model parameters of the corresponding layers and the local training time of this round.

[0106] For the convenience of description, when describing the above system, various modules are described separately according to their functions. Of course, when implementing the embodiments of the present application, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0107] The system of the above embodiment is used to implement the corresponding method in the foregoing embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.

[0108] Figure 3 The figure shows a more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.

[0109] The processor 1010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0110] The memory 1020 can be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and called and executed by the processor 1010.

[0111] The input / output interface 1030 is used to connect to an input / output module to implement information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.

[0112] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to implement communication and interaction between this device and other devices. Among them, the communication module can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.).

[0113] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).

[0114] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, this device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solutions of the embodiments of this specification, and do not necessarily include all the components shown in the figure.

[0115] The electronic device in the above embodiment is used to implement the corresponding method in the foregoing embodiment and has the beneficial effects of the corresponding method embodiment, which will not be elaborated here.

[0116] The computer-readable medium of this embodiment includes both permanent and non-permanent, removable and non-removable media that can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device.

[0117] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; within the context of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present application as described above, and for the sake of brevity, they are not provided in detail.

[0118] In addition, for the sake of simplicity of explanation and discussion, and in order not to make the embodiments of the present application difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the devices may be shown in block diagram form in order to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that the details regarding the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present application are to be implemented (i.e., these details should be entirely within the understanding of those skilled in the art). In cases where specific details (such as circuits) are set forth to describe the exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present application can be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0119] Although the present disclosure has been described in connection with specific embodiments of the present disclosure, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) can be used with the embodiments discussed.

[0120] Embodiments of the present application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of the present application shall be included within the protection scope of this disclosure.

Claims

1. An asynchronous federated learning optimization method, characterized in that, It includes: The central server initializes the global model, generates an initial layer index table, and sends the global model and the initial layer index table to each client; The initial layer index table includes the layer information of all layers of the global model; The client receives the global model and the layer index table, trains all layers according to the layer index table, obtains the updated model parameters of each layer, and sends the updated model parameters of each layer and the local training time to the central server; The central server receives the model parameters and local training time of each client, updates the global model according to the model parameters, determines the importance degree of each layer according to the number of parameters, convergence index, and obsolescence index of each layer, calculates the splitting ratio according to the local training time, updates the layer index table of each client according to the importance degree and splitting ratio of each layer, and sends the updated global model and the updated layer index table to the corresponding client; wherein, the convergence index of each layer is calculated according to the gradient of each layer, and the obsolescence index is determined according to the time difference between two adjacent updates of each layer; The client receives the updated global model and the updated layer index table, and based on the updated global model, trains the corresponding layers in the table according to the updated layer index table, to obtain the updated model parameters of the corresponding layers and the local training time of this round.

2. The method according to claim 1, characterized in that, It also includes: The client sends the updated model parameters of the corresponding layers and the local training time of this round to the central server. The central server updates the global model according to the model parameters of the corresponding layers, determines the importance degree of each layer according to the number of parameters, convergence index, and obsolescence index of each layer, calculates the splitting ratio according to the local training time of this round, updates the layer index table of each client according to the importance degree and the splitting ratio of each layer, and sends the updated global model and the updated layer index table to the corresponding client; repeat the above process until the global model meets the preset training end condition.

3. The method according to claim 1 or 2, characterized in that, The method for the client to train the corresponding layers in the table according to the layer index table to obtain the updated model parameters of the corresponding layers is: Among them, is the updated model parameter of the l-th layer of the local model for client l in the (k + 1)-th round of training, γ is the learning rate during local model training, and L is the loss function. is the model parameter of the l-th layer of the local model for client i in the k-th round of training.

4. The method according to claim 3, characterized in that, The method for the central server to update the global model according to the model parameters is: where m is the number of clients, Ⅱ(l∈layerIndex i ) is an indicator function that takes the value 1 when l∈layerIndex i and 0 otherwise; is the model parameter of the l-th layer of the central server in the k-th round of training, is the model parameter of the l-th layer of the central server s in the (k + 1)-th round of training, ω i is the updated weight of client i in this aggregation.

5. The method according to claim 1, wherein The method for calculating the convergence index is: Among them, is the gradient of the l-th layer of the local model, τ is the window size used to measure the gradient convergence, and j is the time from t - τ to t.

6. The method according to claim 5, wherein The method for determining the importance degree of each layer according to the number of parameters, convergence index, and obsolescence index of each layer is: where α, β, δ are weights, is the number of parameters of the l-th layer, is the obsolescence index of the l-th layer.

7. The method according to claim 6, wherein The method for calculating the splitting ratio according to the local training time is: Among them, T i is the time required for client i in this round of model training, and T min is the minimum value of the time for all clients in this round of model training, and T max is the maximum value of the time for all clients in this round of model training.

8. The method according to claim 7, wherein Updating the layer index table of each client according to the importance degree and splitting ratio of each layer includes: Sort the layers in descending order of importance degree to obtain the sorted layers; Determine the target number of layers for each client that needs to be selected from the sorted layers according to the splitting ratio of the client and the total number of layers of the global model; Based on the sorted layers, select the layers of the target number of layers from front to back, and update the layer index table of each client according to the selected layers.

9. The method according to claim 8, characterized in that, The method for determining the target number of layers that needs to be selected from the sorted layers according to the splitting ratio and the total number of layers of the global model is: Wherein, L is the total number of layers of the global model, and A is the target number of layers to be selected.

10. An asynchronous federated learning optimization system, characterized in that, It includes: A central server, which is used to initialize a global model, generate an initial layer index table, and send the global model and the initial layer index table to each client; The initial layer index table includes layer information of all layers of the global model; and receive the model parameters and local training time of each client, update the global model according to the model parameters, determine the importance level of each layer according to the number of parameters, convergence index, and obsolescence index of each layer, calculate the splitting ratio according to the local training time, and update the layer index table of each client according to the importance level and splitting ratio of each layer, and send the updated global model and the updated layer index table to the corresponding client; wherein, the convergence index of each layer is calculated according to the gradient of each layer, and the obsolescence index is determined according to the time difference between two adjacent updates of each layer; A client, which is used to receive the global model and the initial layer index table, train all layers according to the layer index table to obtain updated model parameters of each layer, and send the updated model parameters of each layer and the local training time to the central server; and receive the updated global model and the updated layer index table, and based on the updated global model, train the corresponding layers in the table according to the updated layer index table to obtain updated model parameters of the corresponding layers and the local training time of this round.