Federal learning method and device
By grouping and model aggregating the client data distributions of edge nodes in federated learning, the problem of poor global model performance caused by non-independent and homogeneous data is solved, and a better personalized model training effect is achieved.
Patent Information
- Application Number
- CN202311627068.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-05-30
AI Technical Summary
In federated learning, when user data is not independently distributed simultaneously, the gradient directions of each user edge node are different after iteration, resulting in the global model generated by the central server performing poorly on each edge node and even failing to converge.
Through data interoperability between edge nodes, edge nodes can group clients with similar data distributions, thereby improving the performance of the model in better personalized model training. The specific method includes: the first edge node obtains information about the client model, groupes and aggregates, generates the edge node model, and uploads it to the server for update of the global model.
By grouping and model aggregating clients with similar data distributions, better personalized model training is achieved, and the performance and stability of the model on each side node is improved.
Smart Images

Figure CN120069118A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computing devices, and in particular, to a federated learning method and apparatus. Background Art
[0002] Federated learning is a method that can train a neural network model in recent years without the data leaving the user devices. The architecture of a federated learning system can include a server and multiple clients. The clients train a local model using the data saved locally, and upload the information of the trained local model to the server. The server aggregates the model information obtained by training multiple clients to obtain a global model.
[0003] When the user data is non-independent and identically distributed (Non-IID), due to the large differences in the gradient directions after iteration of each user edge node, the global model generated by the central server performs poorly on each edge node and may even fail to converge. Therefore, how to train a personalized model on the edge node to facilitate the prediction of local data is an urgent problem to be solved. Summary of the Invention
[0004] This application provides a federated learning method and apparatus for realizing the grouping of clients with similar data distributions among edge nodes through data intercommunication between edge nodes, and achieving better personalized model training.
[0005] In view of this, in a first aspect, the present application provides a federated learning method applied to a federated learning system. The federated learning system includes a server, a plurality of edge nodes, and a plurality of clients. The plurality of edge nodes are connected to the server, and the plurality of clients are connected to the plurality of edge nodes. The edge nodes can also be referred to as edge servers, side servers, or intermediate nodes, etc. Taking the steps executed by any one of the nodes as an example, for the sake of distinction, it is called the first edge node. The method includes: Generally, federated learning is a process of iterative learning. Taking one iteration process of the present application as an example: First, the first edge node obtains the information of the client models uploaded by at least one client. The information of the client models includes the models obtained by the at least one client through model training using the locally saved data, such as the parameters or gradients of the trained client models. The information of the models mentioned below, such as the information of the edge node models or the global model information, etc., may also include the parameters or gradients of the models. The at least one client includes the clients connected to the plurality of edge nodes, such as the clients connected to the first edge node or other edge nodes; Subsequently, the first edge node groups the at least one client according to the information of the client models uploaded by the at least one client, and obtains the edge-side grouping corresponding to the first edge node. The number of edge-side groupings can be one group or multiple groups; Subsequently, the first edge node aggregates the information of the client models uploaded by the clients in the edge-side grouping, and obtains the information of the edge node model; Subsequently, the first edge node sends the information of the edge node model to the server, so that the server performs re-aggregation after receiving the information of one or more edge node models to obtain a new global model; Subsequently, the server can send the aggregated global model to each edge node, that is, the first edge node receives the aggregated global model sent by the server; Subsequently, the first edge node sends the aggregated global model to each client connected to the first edge node, so that each client uses the aggregated global model to update the client model saved locally by each client, and uses the locally saved data to train the updated client model to obtain the trained client model. The information of the trained client model is used to be uploaded to the first edge node again for the next iteration.
[0006] In the embodiments of the present application, the clients can be grouped based on the data distribution of the information of the client models uploaded by the clients. For example, the clients with similar data distributions can be grouped into one group. When aggregating subsequently, the information of the client models uploaded by the clients within the group can be aggregated. Therefore, models with similar data distributions can be aggregated to achieve better personalized training. Moreover, the edge nodes can interact with each other. For example, they can transmit to each other the information of the client models uploaded by the accessed clients or the parameters of the local models of each edge node, etc. Therefore, when grouping the clients, cross-edge-node client grouping can be achieved, so that more accurate grouping of each client within the federated learning system can be realized, which is more conducive to personalized model training and obtaining a personalized model with better performance.
[0007] In a possible implementation manner, the foregoing first edge node groups at least one client according to the information of the client models uploaded by at least one client to obtain at least one group corresponding to the first edge node, which may include: First, the first edge node obtains the information of the client models uploaded by at least one client; Subsequently, the first edge node groups at least one client according to the information of the client models uploaded by at least one client to obtain at least one client-side group; The first edge node determines the edge-side group of the first edge node according to at least one client-side group.
[0008] In the embodiments of the present application, the clients can be grouped according to the information of the client models uploaded by each client. For example, the clients with similar data distributions can be grouped into one group, so that when aggregating the models subsequently, the models with similar distributions can be aggregated, and thus a personalized model can be aggregated.
[0009] In a possible implementation manner, the foregoing first edge node groups at least one client according to the information of the client models uploaded by at least one client to obtain at least one client-side group corresponding to the first edge node, which may include: The first edge node clusters the information of the client models uploaded by at least one client and groups at least one client according to the clustering result to obtain at least one client-side group.
[0010] In the embodiments of the present application, when grouping the clients, grouping can be performed by clustering. For example, the clients corresponding to the data in the same category can be used as a group of clients, or the clients with a distance less than a preset distance value from the clustering center point in each category can be used as a group of client groups, etc., so as to group the clients with closer data distributions into one group, realize personalized model training within the group, and retain as many personalized features as possible.
[0011] In a possible implementation manner, the foregoing first edge node clusters according to the information of the client models uploaded by at least one client, groups at least one client according to the clustering result, and obtains at least one edge-side grouping, which may include: The first edge node clusters the information of the client models uploaded by at least one client to obtain a clustering result; The first edge node groups at least one client according to the clustering result and the number of groups to obtain at least one edge-side grouping; The first edge node evaluates at least one edge-side grouping to obtain an evaluation value of at least one edge-side grouping; If the evaluation value does not meet the preset evaluation condition, the first edge node adjusts the number of groups and re-groups at least one client according to the adjusted number of groups.
[0012] In the implementation manner of this application, when determining the number of clients in each group, the optimal number of groups with the best evaluation result can be selected by evaluating the groups, so as to obtain a better grouping scheme.
[0013] In a possible implementation manner, the foregoing first edge node determines the edge-side grouping of the first edge node according to at least one edge-side grouping, which may include: The first edge node obtains the similarity between the model information saved locally and the information of the client models uploaded by each of at least one client, and the model information saved locally by the first node is the information of the edge node model obtained in the previous iteration process; The first edge node determines the edge-side grouping from at least one edge-side grouping based on the similarity between the model information saved locally and the information of the client models uploaded by each client. For example, the clients with the highest similarity of a preset number or the clients with a similarity higher than the preset similarity can be selected as the clients in the edge-side grouping.
[0014] In the implementation manner of this application, when the edge node selects the clients to access the group, the similarity between the clients in the group and the data saved by the edge node, such as the local model of the edge node, can be calculated, and then the clients to access can be selected based on this similarity. For example, the clients with a similarity higher than the preset similarity to the local model of the edge node can be selected, so as to reduce the interference to model aggregation caused by the difference between the model of the edge node and the models of each client, and obtain a personalized model that is more suitable for the clients in each group.
[0015] In a possible implementation, for the foregoing first edge node to determine the edge-side grouping of the first edge node based on at least one end-side grouping, it may further include: the first edge node further obtains communication resources corresponding to at least one client, where the communication resources include communication resources used when at least one client communicates with the first edge node, such as computing power of each client, available bandwidth of the first edge node, computing power of the first edge node, or communication latency between the first edge node and the client, etc.; subsequently, the first edge node combines the communication resources corresponding to the clients in at least one grouping, and determines the edge-side grouping according to at least one end-side grouping. For example, it can select clients with lower latency, higher computing power, or less bandwidth occupancy to access the edge-side grouping.
[0016] In the embodiments of the present application, clients accessing the edge-side grouping can be selected based on the communication resources when the edge node communicates with the clients. Therefore, when the edge node selects the clients to access, it can select clients adapted to the communication resources for access, so as to balance the communication load of the edge node and improve the learning efficiency of federated learning.
[0017] In a possible implementation, the foregoing method may further include: if the first edge node determines that at least one end-side grouping further includes an end-side grouping corresponding to a second edge node, which can also be referred to as a first end-side grouping. For example, the information of the client model uploaded by the clients in one group of end-side groupings has a high similarity with the edge node model of the second edge node, or the information of the client model uploaded by the clients in this end-side grouping belongs to the same category (i.e., the category after clustering) as the information of the client model uploaded by the clients in the edge-side grouping of the second edge node, then this end-side grouping can be divided into the edge-side grouping corresponding to the second edge node. In this scenario, the first edge node sends the information of the client model uploaded by the clients in the end-side grouping corresponding to the second edge node to the second edge node; when there are multiple clients divided into the edge-side grouping of the second edge node, the information of the client models uploaded by these multiple clients can be directly sent to the second edge node, or the first edge node aggregates the information of the client models uploaded by these multiple clients and then sends the aggregated model information to the second edge node to reduce the communication overhead.
[0018] In the embodiments of the present application, grouping can be performed across edge nodes. For example, edge nodes can transmit the local models stored by themselves. When each edge node groups the clients it accesses (or the clients it has established connections with), it can also calculate the similarity between the information of the client models uploaded by the clients and the local models of other edge nodes, and divide the clients with a high similarity (such as a similarity higher than a preset similarity value) with other edge nodes into the edge-side groupings of other edge nodes, so as to achieve cross-edge node client grouping. After grouping, the information of the client models uploaded by the clients divided into other edge nodes is sent to other edge nodes.
[0019] In a possible implementation, the first edge node sending the information of the client model uploaded by the client in the end-side packet corresponding to the second edge node may include: the first edge node sending the information of the client model uploaded by the client in the end-side packet corresponding to the second edge node through a direct connection established with the second edge node, and the direct connection may be a short-distance communication or a wired connection, etc.; or, the first edge node sending the information of the client model uploaded by the client in the end-side packet corresponding to the second edge node through the backbone network in the communication network, and the communication network is the network where the first edge node and the second edge node are located.
[0020] In the embodiments of the present application, edge nodes can communicate directly with each other or through the backbone network. Therefore, edge nodes can directly transmit the information of the client model uploaded by the client or transmit the information of the client model uploaded by the client through the backbone network, so as to realize data interaction between edge nodes.
[0021] In a possible implementation, the first edge node sending the information of the edge node model to the server may include: the first edge node sending the information of at least one network layer in the edge node model, such as the parameters or gradients of a certain layer or multiple layers of the network layer. That is, when the edge node uploads the model, it can be uploaded layer by layer, such as only uploading one or more of them, so as to reduce the communication overhead.
[0022] In the embodiments of the present application, when the first edge node uploads the edge node model to the server, it can upload all or part of the parameters of the network layer of the edge node model. For example, it can transmit the network layer with a sensitivity higher than a preset sensitivity value or a change value exceeding a preset change value during the training process, so as to reduce the communication overhead.
[0023] In a possible implementation, the first edge node grouping at least one client according to the information of the models uploaded by at least one client to obtain the edge-side packet corresponding to the first edge node may further include: the first edge node obtaining the similarity between the model information saved locally and the information of the models uploaded by at least one client, and the locally saved model information may be the information of the edge node model obtained in the previous iteration process; screening at least one client according to the similarity between the locally saved model information and the information of the models uploaded by at least one client to obtain the edge-side packet. For example, selecting the N (positive integer) clients with the highest similarity as the clients in the edge-side packet. Therefore, in the embodiments of the present application, the edge node can directly screen out the clients in the edge-side packet from one or more clients, so as to more efficiently determine the model that is more similar to the local model of the edge node.
[0024] In a possible implementation, the client model locally saved by each client is used to perform at least one of a recommendation task, a classification task, a segmentation task, or an object detection task. The globally aggregated model can be deployed on each client. The globally aggregated model obtained by the method provided in this application is more adapted to the distribution of the local data of each client, and can implement personalized model training for each client, enabling the client to perform local personalized inference tasks in combination with the globally aggregated model, such as recommendation tasks, classification tasks, segmentation tasks, or object detection tasks, etc.
[0025] In a second aspect, the present application provides a federated learning method, which is applied to a federated learning system. The federated learning system includes a server, multiple edge nodes, and multiple clients. The multiple clients are connected to the multiple edge nodes. The method includes: in one round of iteration, the server receives the information of the edge node models sent by the multiple edge nodes respectively, or the edge node models for short. The information of the edge node models includes the aggregation of the client model information uploaded by the clients in at least one edge-side group by each edge node among the multiple edge nodes. At least one edge-side group is obtained by each edge node grouping at least one terminal according to the data uploaded by at least one terminal connected to the multiple edge nodes; subsequently, the server aggregates the received information of the edge node models to obtain a global model; and sends the new globally aggregated model to each edge node below, so that each edge node distributes the new global model to each client, enabling each client to update the client model locally saved based on the new global model, and training the updated client model using the locally saved training data to obtain a trained model and uploading it to the edge node again for the next iteration.
[0026] In the implementation of the present application, the server only needs to receive the models aggregated by the edge nodes uploaded by the edge nodes, without receiving the information uploaded by all clients, reducing the communication overhead. And the edge nodes can communicate with each other, enabling cross-edge-node client grouping. Therefore, personalized grouping for clients can be achieved, enabling better personalized model training and obtaining a personalized model with better performance.
[0027] In a possible implementation, in one round of iteration, the above method may further include: the server sends the information of the global model to each edge node. Therefore, the server can distribute the new global model to the edge nodes, so as to update the global models of each edge node.
[0028] In a third aspect, the present application provides a federated learning device, which is applied to a federated learning system. The federated learning system includes a server, multiple edge nodes, and multiple clients. The multiple clients are connected to the multiple edge nodes. The federated learning device is any one of the multiple edge nodes. The device includes:
[0029] A transceiver module, configured to obtain data uploaded by at least one client, where the at least one client includes a client accessing multiple edge nodes, and the first edge node is any one of the multiple edge nodes;
[0030] A grouping module, configured to group at least one client according to the information of the client models uploaded by the at least one client, so as to obtain an edge-side grouping corresponding to the federated learning device;
[0031] An aggregation module, configured to aggregate the information of the client models uploaded by the clients in the edge-side grouping to obtain the information of the edge node model;
[0032] The transceiver module is further configured to send the information of the edge node model to a server, and the information of the edge node model is used for the server to aggregate and obtain a global model;
[0033] The transceiver module is further configured to receive the aggregated global model sent by the server;
[0034] The transceiver module is further configured to send the aggregated global model to each client accessing the first edge node, so that each client uses the aggregated global model to update the client model locally saved by each client, and trains the updated client model using the locally saved data to obtain a trained client model, and the information of the trained client model is used to be uploaded to the first edge node again.
[0035] In a possible implementation manner, the transceiver module is further configured to obtain the information of the client models uploaded by at least one client;
[0036] A grouping module, configured to group at least one client according to the information of the client models uploaded by the at least one client, so as to obtain at least one end-side grouping;
[0037] The grouping module is configured to determine the edge-side grouping of the federated learning device according to at least one end-side grouping.
[0038] In a possible implementation manner, the grouping module is specifically configured to perform clustering according to the information of the client models uploaded by at least one client, and group at least one client according to the clustering result, so as to obtain at least one end-side grouping.
[0039] In a possible implementation, the grouping module is specifically configured to: cluster based on the information of the client models uploaded by at least one client to obtain a clustering result; group at least one client according to the clustering result and the number of groups to obtain at least one end-side group; evaluate at least one end-side group to obtain an evaluation value of at least one end-side group; if the evaluation value does not meet the preset evaluation condition, adjust the number of groups, and re-group at least one client according to the adjusted number of groups.
[0040] In a possible implementation, the grouping module is specifically configured to: obtain the similarity between the model information saved locally and the information of the client models uploaded by each of at least one client, where the model information saved locally is the information of the edge node model obtained in the previous iteration process; determine the edge-side group according to at least one end-side group based on the similarity between the model information saved locally and the information of the client models uploaded by each client.
[0041] In a possible implementation, the grouping module is specifically configured to: obtain the communication resources corresponding to at least one client, where the communication resources include the communication resources used when at least one client communicates with the first edge node; combine the communication resources corresponding to the clients in at least one group, and determine the edge-side group corresponding to the first edge node from at least one end-side group.
[0042] In a possible implementation, the transceiver module is further configured to: if the first edge node determines that at least one end-side group further includes the end-side group corresponding to the second edge node, which can also be referred to as the first end-side group, for example, the similarity between the information of the client model uploaded by the clients in one of the end-side groups and the edge node model of the second edge node is relatively high, or the information of the client model uploaded by the clients in this end-side group and the information of the client models uploaded by the clients in the edge-side group of the second edge node belong to the same category (i.e., the category after clustering), then send the information of the client models uploaded by the clients in the end-side group corresponding to the second edge node to the second edge node.
[0043] In a possible implementation, the transceiver module is further configured to: send the information of the client models uploaded by the clients in the end-side group corresponding to the second edge node to the second edge node through the connection established with the second edge node; or send the information of the client models uploaded by the clients in the end-side group corresponding to the second edge node to the second edge node through the backbone network.
[0044] In a possible implementation, the transceiver module is further configured to send the information of at least one network layer in the edge node model to the server.
[0045] In a possible implementation, the grouping module is further configured to: The first edge node obtains the similarity between the locally saved model information and the information of the models uploaded by at least one client. The locally saved model information may be the information of the edge node model obtained in the previous iteration process; According to the similarity between the locally saved model information and the information of the models uploaded by at least one client, screen at least one client to obtain an edge-side grouping. For example, select the top N (positive integer) clients with the highest similarity as the clients in the edge-side grouping.
[0046] In a possible implementation, the client model locally saved by each client is used to perform at least one of a recommendation task, a classification task, a segmentation task, or an object detection task. Equivalently, the globally aggregated model can be deployed in the client. The globally aggregated model obtained by the method provided in this application is more adapted to the distribution of the local data of each client, and can realize personalized model training for each client, so that the client can combine the globally aggregated model to perform local personalized inference tasks, such as recommendation tasks, classification tasks, segmentation tasks, or object detection tasks, etc.
[0047] In a fourth aspect, the present application provides a federated learning device, which is applied to a federated learning system. The federated learning system includes a server, multiple edge nodes, and multiple clients. The multiple clients are connected to the multiple edge nodes. The federated learning device includes a server, and the device includes:
[0048] A transceiver module, configured to receive the information of the edge node models respectively sent by the multiple edge nodes. The information of the edge node models is obtained by aggregating the information of the client models uploaded by the clients in at least one edge-side grouping by each edge node in the multiple edge nodes. At least one edge-side grouping is obtained by each edge node grouping at least one terminal according to the data uploaded by at least one terminal connected to the multiple edge nodes;
[0049] An aggregation module, configured to aggregate the information of the edge node models to obtain a global model;
[0050] The transceiver module is further configured to send the new globally aggregated model to each edge node below, so that each edge node distributes the new globally aggregated model to each client, so that each client updates the client model locally saved based on the new globally aggregated model, and trains the updated client model using the locally saved training data to obtain a trained model and upload it to the edge node again for the next iteration.
[0051] In a possible implementation, the transceiver module is further configured to send the information of the global model to each edge node.
[0052] Fifth aspect, an embodiment of the present application provides a federated learning device, including: a processor and a memory. The processor and the memory are interconnected by a line, and the processor calls the program code in the memory to execute the functions related to processing in the federated learning method shown in any one of the above first aspects. Optionally, the federated learning device may be a chip.
[0053] Sixth aspect, an embodiment of the present application provides a federated learning device, including: a processor and a memory. The processor and the memory are interconnected by a line, and the processor calls the program code in the memory to execute the functions related to processing in the first method shown in any one of the above second aspects. Optionally, the federated learning device may be a chip.
[0054] Seventh aspect, an embodiment of the present application provides a federated learning system, including a plurality of servers, a plurality of edge nodes, and at least one client. The at least one client accesses the plurality of edge nodes; any one of the plurality of servers is used to implement the steps in the federated learning method shown in any one of the above second aspects; the plurality of edge nodes are used to implement the steps in the federated learning method shown in any one of the above first aspects; the at least one client is used to perform model training using the locally saved data.
[0055] Eighth aspect, an embodiment of the present application provides a digital processing chip or a chip. The chip includes a processing unit and a communication interface. The processing unit obtains program instructions through the communication interface, and the program instructions are executed by the processing unit. The processing unit is used to execute the functions related to processing in any optional implementation manner in the above first aspect or second aspect.
[0056] Ninth aspect, an embodiment of the present application provides a computer-readable storage medium, including instructions. When it runs on a computer, it causes the computer to execute the method in any optional implementation manner in the above first aspect or second aspect.
[0057] Tenth aspect, an embodiment of the present application provides a computer program product including computer programs / instructions. When it is executed by a processor, it causes the processor to execute the method in any optional implementation manner in the above first aspect or second aspect. Description of the Drawings
[0058] Figure 1 It is a schematic diagram of the architecture of a federated learning system provided by the present application;
[0059] Figure 2 It is a schematic diagram of the architecture of another federated learning system provided by the present application;
[0060] Figure 3 It is a schematic diagram of the architecture of another federated learning system provided by the present application;
[0061] Figure 4 Flow schematic diagram of a federated learning method provided by this application;
[0062] Figure 5 Flow schematic diagram of another federated learning method provided by this application;
[0063] Figure 6 Flow schematic diagram of another federated learning method provided by this application;
[0064] Figure 7 Flow schematic diagram of another federated learning method provided by this application;
[0065] Figure 8 Flow schematic diagram of another federated learning method provided by this application;
[0066] Figure 9 Flow schematic diagram of another federated learning method provided by this application;
[0067] Figure 10 Flow schematic diagram of another federated learning method provided by this application;
[0068] Figure 11 Flow schematic diagram of another federated learning method provided by this application;
[0069] Figure 12 Flow schematic diagram of another federated learning method provided by this application;
[0070] Figure 13 Structure schematic diagram of a federated learning device provided by this application;
[0071] Figure 14 Structure schematic diagram of another federated learning device provided by this application;
[0072] Figure 15 Structure schematic diagram of another federated learning device provided by this application;
[0073] Figure 16 Structure schematic diagram of another federated learning device provided by this application. Detailed implementation manners
[0074] Next, the technical solutions in the embodiments of this application will be described with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of this application.
[0075] The embodiments of the present application are mainly applied to training machine learning models used in various application scenarios of federated learning. The trained machine learning models can be applied to the above various application fields to achieve classification, regression, or other functions. The processing objects of the trained machine learning models can be image samples, discrete data samples, text samples, speech samples, etc., which are not enumerated here. Among them, the machine learning model can specifically be a neural network, a linear model, or other types of machine learning models, etc. Correspondingly, the multiple modules that make up the machine learning model can specifically be neural network modules, current model modules, or modules that make up other types of machine learning models, etc., which are not enumerated here. In the following embodiments, only the case where the machine learning model is a neural network is used as an example for illustration. When the machine learning model is of other types other than the neural network, it can be understood by analogy and will not be elaborated in the embodiments of the present application.
[0076] The embodiments of the present application can be applied to the field of federated learning and can train a neural network through the cooperation of a client and a server. Therefore, a large number of applications related to neural networks are involved, such as the neural network trained by the client during federated learning. To better understand the solutions of the embodiments of the present application, the following first introduces the relevant terms and concepts of the neural network that the embodiments of the present application may involve.
[0077] (1) Neural Network
[0078] A neural network can be composed of neural units. A neural unit can refer to an operation unit that takes xs and an intercept of 1 as inputs. The output of this operation unit can be as shown in formula (1-1):
[0079]
[0080] where s = 1, 2,... n, n is a natural number greater than 1, W s is the weight of x s , b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce non-linear characteristics into the neural network to convert the input signal in the neural unit into an output signal. The output signal of this activation function can be used as the input of the next convolutional layer. The activation function can be a sigmoid function. A neural network is a network formed by connecting multiple such single neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be a region composed of several neural units.
[0081] (2) Loss Function
[0082] During the process of training a neural network, since we hope that the output of the neural network is as close as possible to the value we really want to predict, we can compare the predicted value of the current network with the real target value, and then update the weight vector of each layer of the neural network according to the difference between the two. (Of course, there is usually an initialization process before the first update, that is, configuring parameters for each layer in the deep neural network in advance). For example, if the predicted value of the network is too high, we adjust the weight vector to make it predict lower, and keep adjusting until the deep neural network can predict the real target value or a value very close to the real target value. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or objective function. They are important equations used to measure the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. Then the training of the deep neural network becomes a process of minimizing this loss as much as possible. The loss function usually can include mean squared error, cross-entropy, logarithm, exponential and other loss functions. For example, mean squared error can be used as the loss function, defined as Specifically, the specific loss function can be selected according to the actual application scenario.
[0083] (3) Backpropagation algorithm
[0084] The neural network can use the backpropagation (BP) algorithm to correct the magnitude of the parameters in the initial neural network model during the training process, so that the reconstruction error loss of the neural network model becomes smaller and smaller. Specifically, the forward propagation of the input signal until the output will generate an error loss, and the initial neural network model parameters are updated by backpropagating the error loss information, so that the error loss converges. The backpropagation algorithm is a backpropagation movement dominated by the error loss, aiming to obtain the optimal parameters of the neural network model, such as the weight matrix.
[0085] In this application, when the client is performing model training, it can train the global model through the loss function or through the BP algorithm to obtain the trained global model.
[0086] (4) Gradient
[0087] That is, the derivative vector of the loss function with respect to the parameters. It can be used to measure the change value of the model's parameters.
[0088] (5) Non-independent-identically-distributed (Non-IID)
[0089] In the context of federated learning, it refers to the different data distributions on each edge user side. "Independent and identically distributed" means that the data distributions of each terminal are independent of each other and have the same distribution.
[0090] (6) Federated learning (FL)
[0091] A distributed machine learning algorithm that collaboratively completes model training and algorithm updates through multiple clients, such as mobile devices or edge servers, and a server without the data leaving the domain, in order to obtain a trained global model. It can be understood that during the process of machine learning, each participating party can jointly build a model by leveraging the data of other parties. Each party does not need to share data resources, that is, data joint training is carried out without the data leaving the local area to establish a shared machine learning model.
[0092] Federated learning is a way to train a neural network model without the data leaving the user device. Usually, at the beginning of training, the central server sends an initial model to each edge user, and the edge user then sends the model to each end user. Subsequently, after different end users use local data to iterate the model for several epochs, they feedback the model change amount (such as gradient or model weight parameters) back to the edge server. The edge server performs weighted averaging on the feedback gradients, updates the initial model with the obtained average gradient, and distributes the updated model to each end user to restart the next small round of iteration. After several rounds of end-edge synchronization on the edge server, the edge server feedbacks the change amount (such as gradient or model weight parameters) back to the central server. The central server performs weighted averaging on the feedback gradients, updates the initial model with the obtained average gradient, and distributes the updated model to each edge user, and then the edge user distributes it to the end user to restart the next large round of iteration.
[0093] When the user data is non-independent and identically distributed (Non-IID), due to the large differences in the gradient directions after iteration of each user edge node, the global model generated by the central server performs poorly on each edge node, or even fails to converge. In real life, the geographical distances of the end nodes physically connected to the edge nodes are relatively small, and usually the locally collected data distributions are similar.
[0094] In some existing federated learning solutions, two models can be trained simultaneously: a global model and a local model. After the end-user performs local iterations for several epochs, the end-user uploads the gradient update obtained from the global model to the server for averaging to obtain a new global model. Subsequently, the end-user downloads the global model from the server, sums the global model and the local model weighted to obtain a new local model, and prepares for the next round of local iteration. Finally, the trained local model is used as the personalized model. However, since all clients participate in the training of the global model, under non-IID or malicious attacks, the personalized model is vulnerable to the adverse effects of the global model, resulting in poor performance of the personalized model.
[0095] In some existing solutions, the server clusters nodes with data of similar distributions and trains a personalized model in each group. However, the model parameter values owned by all nodes within a group are exactly the same, resulting in the inability to optimize the personalized model for each node. In addition, a high synchronization frequency causes the models on the nodes to be averaged by the models of other nodes within the group after certain differences occur. Therefore, once grouped, it is difficult for a node to be grouped into another group.
[0096] Therefore, the present application provides a federated learning method that can group clients through edge nodes, enabling better personalized model learning.
[0097] First, the federated learning system provided by the present application is introduced.
[0098] Refer to Figure 1 , a schematic diagram of the architecture of a federated learning system provided by the present application. This system (or can also be simply referred to as a cluster) can include multiple servers, and connections can be established between these multiple servers, that is, communication can also be carried out between each server. Each server can be connected to one or more clients through an intermediate device (which can also be called an edge node or an edge server, etc.), or in other words, each edge node can access one or more clients. This intermediate device can be a routing device between the server and the client, such as a base station, a router, a switch, etc. The client accesses the corresponding edge node, and the client can be deployed in various devices, such as being deployed in a mobile terminal or a server, etc., as Figure 1 shown by client 1, client 2,..., client N-1, and client N, etc.
[0099] Specifically, servers can interact with each other, or servers, edge nodes and clients can interact through a communication network using any communication mechanism / communication standard. The communication network can be a wide area network, a local area network, a point-to-point connection, or any combination thereof. Specifically, the communication network can include a wireless network, a wired network, or a combination of a wireless network and a wired network. The wireless network includes, but is not limited to: a fifth-generation mobile communication technology (5th-Generation, 5G) system, a long term evolution (LTE) system, a global system for mobile communication (GSM), or a code division multiple access (CDMA) network, a wideband code division multiple access (WCDMA) network, wireless fidelity (WiFi), Bluetooth, Zigbee, radio frequency identification (RFID), Long Range (Lora) wireless communication, near field communication (NFC), or any combination of one or more of them. The wired network can include an optical fiber communication network or a network composed of coaxial cables, etc.
[0100] Accordingly, the edge node may include a device accessing multiple clients, such as a next generation node basestation (gNB) in a 5G system, an evolved node B (eNB) in an LTE system, a radio network controller (RNC), a node B (NB), a base station controller (BSC), a base transceiver station (BTS), a home evolved node B (or home node B, HNB), a base band unit (BBU), a transmitting and receiving point (TRP), a transmitting point (TP), a pico, a mobile switching center, or a network device in a future network, etc. It can be understood that the specific type of the access network device is not limited in this application. In systems adopting different radio access technologies, the names of the devices with the functions of access network devices may be different.
[0101] Generally, the clients can be deployed in various servers or terminals. The clients mentioned below can also refer to the servers or terminals on which the client software programs are deployed. The terminal can include a mobile terminal or a fixedly installed terminal, etc. For example, the terminal can specifically include a mobile phone, a tablet, a personal computer (PC), a smart bracelet, a speaker, a TV, a smart watch or other terminals, etc.
[0102] Specifically, the federated learning system to which the federated learning method provided in this application can be applied can include various topological relationships. For example, the federated learning system can include an architecture with two or more layers. Some possible architectures are introduced exemplarily below.
[0103] I. Three-layer architecture
[0104] As Figure 2 shown, it is a schematic structural diagram of a federated learning system provided in this application.
[0105] Among them, the federated learning system includes one or more cloud servers, one or more edge servers, and one or more clients, forming a three-layer architecture of cloud server-edge server-client.
[0106] In this system, one or more edge servers are connected to the cloud server, and one or more clients are connected to the edge server.
[0107] In the process of federated learning, the cloud server distributes the globally saved model to the edge server, and then the edge server distributes the globally saved model to the clients connected to it. The clients use the locally stored training samples to train the received globally saved model and feedback the trained globally saved model to the edge server. The edge server updates the locally saved globally saved model according to the received trained globally saved model and then feeds back the updated globally saved model by the edge server to the cloud server to complete the federated learning.
[0108] II. Architecture with three or more layers
[0109] As Figure 3 shown, it is a schematic structural diagram of another federated learning system provided by the present application.
[0110] Among them, the federated learning system includes an architecture with three or more layers. One layer includes one or more cloud servers, and multiple edge servers form an architecture with two or more layers. For example, one or more upstream edge servers form one layer of the architecture, and each upstream edge server is connected to one or more downstream edge servers. Each edge server in the last layer of the architecture formed by the edge servers is connected to one or more clients, so that the clients form one layer of the architecture.
[0111] In the process of federated learning, the upstream cloud server distributes the latest globally saved model stored locally to the edge server in the next layer. Subsequently, the edge server distributes the globally saved model layer by layer to the next layer until it is distributed to the clients. After receiving the globally saved model distributed by the edge server, the clients use the locally stored training samples to train the received globally saved model and feedback the trained globally saved model to the edge server in the upper layer. Then, after the edge server in the upper layer updates the locally saved globally saved model based on the received trained globally saved model, it can upload the updated globally saved model to the edge server in the upper layer. By analogy, until the edge server in the second layer uploads the updated globally saved model to the cloud server, and the cloud server updates the globally saved model locally based on the received globally saved model to obtain the final globally saved model, thus completing the federated learning.
[0112] The foregoing introduced the federated learning system provided by the present application. Next, in combination with this federated learning system, the flow of the method provided by the present application will be introduced.
[0113] Refer to Figure 4 , it is a schematic flow diagram of a federated learning method provided by the present application, as described below.
[0114] First, the architecture of the federated learning system provided by the present application can refer to the foregoing Figure 2 or Figure 3As introduced above, for example, a three - layer architecture is taken for exemplary introduction here. And federated learning is usually iterative in multiple rounds. In the embodiments of the present application, for example, one iteration process is taken for exemplary introduction.
[0115] For the convenience of distinction, the model aggregated by the server is called the global model, the model aggregated by the edge nodes is called the edge - node model, and the model maintained and trained by the client is called the client model. This will not be elaborated hereinafter.
[0116] 401. The server sends the global model to the edge nodes.
[0117] First, the server can send the global model to the edge nodes. The global model is the global model aggregated in the previous round of iteration or the initial model. If this round of iteration is the first iteration, the global model is an initialized model to be trained. If this round of iteration is not the first iteration, the global model can be updated according to the models uploaded by the edge nodes in the previous round of iteration.
[0118] In the federated learning system provided by the present application, the number of edge nodes can be one or more. If there are multiple edge nodes, the cloud server can send the global model to each of the multiple edge nodes respectively. Moreover, the cloud service can actively send the global model to the edge nodes or send the global model to each edge node upon the request of the edge nodes, which can be adjusted according to the actual application scenario. The present application does not make any limitation thereto.
[0119] Specifically, the server can send the structural parameters of the global model (such as the width, depth of the neural network or the size of the convolution kernel, etc.) or the initial weight parameters, etc. to the downstream devices.
[0120] Optionally, the server can also send the training configuration parameters, such as the learning rate, the number of epochs or the parameters in the category of the security algorithm, etc. to the downstream devices, so that the client for final training can use the training configuration parameters to train the global model.
[0121] In the embodiments of the present application, the mentioned models, such as the global model, edge node model, client model, or local model, etc., may specifically include neural networks such as convolutional neural networks (CNN), deep convolutional neural networks (DCNN), recurrent neural network (RNN), etc. The model to be learned can be specifically determined according to the actual application scenario, and the present application does not limit this. The foregoing models can be specifically used to perform at least one of the recommendation task, classification task, segmentation task, or object detection task. The specific application scenario of the model can be adjusted according to the actual application scenario, and the present application does not limit this.
[0122] 402. The edge node sends the model to the client.
[0123] The edge node updates the local edge node model of the edge node based on the received global model. Specifically, it can directly replace the locally stored edge node model with the global model to obtain a new local edge node model; or, it can also aggregate the received global model and the edge node model obtained in the previous iteration process saved locally, such as taking the average value or weighted fusion, etc., so as to obtain a new local model.
[0124] After obtaining the global model, the edge node can directly send the received global model to the client, or send the updated local edge node model of the edge node to the client.
[0125] 403. The client uses local data for training and uploads the information of the trained client model to the edge node.
[0126] After receiving the model sent by the edge node, the client can update the local client model of the client based on the received model. For example, fuse the locally saved client model with the global model, or replace the locally saved client model with the global model, etc. And use the training samples stored locally to train the new local client model to obtain the trained client model.
[0127] And upload the information of the trained client model to the edge node. Specifically, it can upload the weight parameters, gradient parameters, etc. of the client model.
[0128] Accordingly, the training data on each client side can be related to the inference tasks that the client model needs to perform. For example, if the client model is used to perform a recommendation task, the training data can include data related to the recommendation scenario collected on the client side, such as the data that has been recommended to the user and the feedback of the client on the recommended data; if the client model is used to perform a classification task, the training data can include the classified data collected by the client; if the client model is used to perform a segmentation task, the training data can include the segmented data collected by the client; if the client model is used to perform an object detection task, the training data can include the data with marked objects collected by the client, etc. Specifically, it can be selected according to the actual application scenario.
[0129] 404. The edge node groups the accessed clients based on the information of the client model uploaded by the clients.
[0130] After receiving the information of the trained model uploaded by one or more accessed clients, the edge node can group the one or more clients based on the information of the trained model uploaded by the one or more clients, and divide the one or more clients into at least one edge-side group.
[0131] In addition, the federated learning system can include multiple edge nodes, and data can be transmitted between the edge nodes. Therefore, when grouping the clients, each edge node can also group the clients accessing other edge nodes. Thus, the grouping of each client in the federated learning system can be realized, and the clients with a data distribution closer to each other in the system can be divided into one group across edge nodes, which can achieve better personalized model training and thus improve the performance of the personalized model.
[0132] Optionally, the edge node can first group the clients based on the information of the client model uploaded by the clients to obtain at least one client-side group, and then further determine the edge-side group corresponding to the edge node based on the at least one client-side group.
[0133] Optionally, when grouping the clients, specifically, clustering can be performed based on the client model information uploaded by the clients, and then the clients can be grouped according to the clustering result to obtain at least one client-side group. For example, clustering can be performed on the information of the client models uploaded by at least one client, and the at least one client can be grouped according to the clustering result. For example, the clients corresponding to the model information in the same category of data are divided into one group to obtain at least one client-side group. Therefore, in the implementation manner of this application, when grouping the clients, the terminals can be grouped by combining the data structure characteristics of the client model information uploaded by the clients, so that the clients with similar data structure characteristics are divided into one group.
[0134] Optionally, the number of each edge-side grouping may be a preset number or a number determined according to the evaluation result. For example, after obtaining the clustering result, at least one client may be grouped based on the clustering result and the number of groupings to obtain at least one edge-side grouping. The at least one edge-side grouping is evaluated to obtain an evaluation value for each edge-side grouping. If the evaluation value does not meet the preset evaluation condition, such as the evaluation value is less than a preset value, the number of groupings may be adjusted, and at least one client is regrouped according to the adjusted number of groupings until an edge-side grouping with an evaluation value meeting the preset evaluation condition is obtained. In the implementation manner of this application, the grouping result may be evaluated to obtain a qualified grouping result.
[0135] In a possible scenario, the edge node may also obtain the similarity between the locally saved model information and the information of the client model uploaded by each of at least one client. The locally saved model information is the information of the edge node model obtained in the previous iteration process; the first edge node determines the edge-side grouping according to at least one edge-side grouping based on the similarity between the locally saved information of the edge node model and the information of the client model uploaded by each client. For example, clients with a similarity less than the preset similarity between the model and the edge node in the edge-side grouping may be screened out to avoid the influence of data with large differences on the edge node model data of the edge node. In the implementation manner of this application, the edge node may select the clients in the edge-side grouping based on the similarity between the edge node model saved by itself and the client model uploaded by the client, so that clients with a higher model similarity may be selected as the edge-side grouping corresponding to the edge node.
[0136] Optionally, the method for the edge node to group the clients to obtain the edge-side grouping may further include: the edge node obtains the similarity between the locally saved model information and the information of the model uploaded by at least one client. The locally saved model information may be the information of the edge node model obtained in the previous iteration process; according to the similarity between the locally saved model information and the information of the model uploaded by at least one client, at least one client is screened to obtain the edge-side grouping. For example, the top N (positive integer) clients with the highest similarity are selected as the clients in the edge-side grouping. Therefore, in the implementation manner of this application, the edge node may directly screen out the clients in the edge-side grouping from one or more clients, so that a model more similar to the local model of the edge node may be determined more efficiently.
[0137] In addition, in a possible scenario, the edge node can also obtain the communication resources corresponding to each client, that is, the resources available when each client communicates with the edge node, such as network latency or computing power. The edge node can determine the clients in the edge-side group in combination with the communication resources corresponding to each client, so as to obtain the clients with higher communication efficiency with the edge node as the edge-side group, obtain an edge-side group adapted to the communication resources of the edge node, and improve the learning efficiency of federated learning.
[0138] In addition, in a possible scenario, one of the edge nodes (for example, called the first edge node for easy distinction) can identify the similarity between the information of the client model of the clients in each end-side group and the edge node model of other edge nodes (for example, called the second edge node for easy distinction) or the client model of the clients accessing the second edge node. If the similarity is higher than the preset value, it can send the information of the client model uploaded by the client with the similarity higher than the preset value to the second edge node. Therefore, the edge nodes can communicate with each other, and can send the information of the client model uploaded by the clients with similar data structure features to other edge nodes with similar data structure features, so as to aggregate the model information between the edge nodes and clients with similar data structure features.
[0139] In addition, in some possible scenarios, when model information needs to be transmitted between edge nodes, the model information to be transmitted can be directly transmitted through the connection established between the edge nodes, or the model information that needs to be transmitted between the edge nodes can be transmitted through the backbone network in the communication network. Taking the first edge node and the second edge node as an example, the first edge node sends the information of the client model uploaded by the clients in the end-side group corresponding to the second edge node to the second edge node through the direct connection established with the second edge node (such as short-distance communication or wired connection, etc.); or, the first edge node sends the information of the client model uploaded by the clients in the end-side group corresponding to the second edge node to the second edge node through the backbone network of the communication network, and this communication network is the communication network accessed by the first edge node and the second edge node.
[0140] 405. The edge node aggregates the information of the client model uploaded by the clients in the edge-side group to obtain the information of the edge node model.
[0141] After the edge node groups the clients to obtain the edge-side group, it can aggregate the information of the client model uploaded by the clients in the edge-side group to obtain the information of the edge node model.
[0142] Among them, the number of edge node models can be one or more. For example, when the edge-side grouping includes multiple types of client groupings, that is, the end-side grouping, aggregation can be performed separately based on each type of end-side grouping, and multiple edge-side models can be obtained. When the edge-side grouping only includes one type of client grouping, that is, the end-side grouping, an edge node model can be obtained through aggregation.
[0143] 406. The edge node uploads the information of the edge node model to the server.
[0144] After the edge node obtains the information of the edge node model, it can upload the information of the edge node model to the server. Specifically, it can be the weight parameters or gradient parameters of the transmission model, etc.
[0145] 407. The server aggregates the information of the received edge node models to obtain an aggregated global model.
[0146] After the server receives the information of one or more edge node models, it can aggregate the received edge node models again to obtain an aggregated global model.
[0147] 408. The server sends the aggregated global model to the edge node.
[0148] 409. The edge node sends the aggregated complete global model to the client.
[0149] For the specific implementation of steps 408 to 409, reference can be made to the descriptions of steps 401 to 402 above, and details will not be elaborated here.
[0150] After the server obtains the aggregated global model, it can continue the next round of iteration, that is, repeat step 401. Details of similar steps will not be elaborated here.
[0151] In the embodiment of the present application, the edge node can group one or more clients accessing it, and the grouping can be based on the data structure characteristics of each client. Therefore, when performing model aggregation, the client model information uploaded by the clients within the same group can be aggregated, so as to realize personalized model training within each group and make the performance of the finally obtained model better. Equivalently, the global model obtained by the method provided in the present application realizes personalized training, retains the personalized characteristics of each client to a greater extent, and the aggregated global model can be deployed as a new client model in each client and can be used to execute the local personalized inference tasks of each client, such as recommendation tasks, classification tasks, segmentation tasks or target detection tasks adapted to each client, etc., so as to realize personalized client inference tasks.
[0152] The above introduced the process of the federated learning method provided by this application. Next, in combination with more specific application scenarios, the method process provided by this application will be introduced.
[0153] Combined with the above Figure 4 As shown in the process, the federated learning method provided by this application can be divided into multiple steps. Next, each step will be introduced separately.
[0154] 501. The server sends down the global model.
[0155] As Figure 5 shown, if the current iteration is the initial stage, the server can send down the initial global model and the local model. The local model is the initial model of each local side. Specifically, the server can send down the structure parameters or weight parameters of the global model, etc.
[0156] If the current iteration is not the first iteration, the server can only send down the global model, or can also send down a new local model, which can be adjusted according to the actual application scenario. The local model is the model that needs to be updated locally by the client, or the structure in the global model that the client needs to update.
[0157] Generally, in a federated learning system, the client can train two models simultaneously: the global model and the local model. After the end - side user performs local iteration for several epochs, the gradient update amount obtained from the global model is uploaded to the server for averaging to obtain a new global model. In the subsequent iteration process, the end - side user receives the new global model, sums the new global model and the local model with weights to obtain a new local model, and prepares for the next round of local iteration. Finally, the trained local model is used as the local personalized model. In the exemplary embodiment of this application, one iteration is taken as an example for exemplary introduction, and details will not be repeated hereinafter.
[0158] In addition, when data is transmitted between various nodes in the federated learning system, such as when transmitting model information between the server and edge nodes, between various edge nodes, or between edge nodes and clients, in order to reduce communication overhead, the data to be transmitted can be quantized before transmission, that is, the number of bits occupied by the data to be transmitted is reduced, thereby reducing communication overhead. After receiving the quantized data, the receiving end can perform de - quantization on the received data to restore the full - precision model information.
[0159] 502. The edge node sends down the global model, and the client uploads the trained client model.
[0160] As Figure 6As shown in the figure, after receiving the global model sent by the server, the edge node can directly forward the global model to each client; or it can fuse the global model with the edge node model of the edge node and then send the fused model to each client.
[0161] After receiving the global model or the edge node model sent by the edge node, the client can perform weighted fusion of the received global model or edge node model with the local client model to obtain a new client model, train the new client model to obtain a trained client model, and upload the newly trained client model to the edge node. It can also directly use the received global model or edge node model as the new client model, train the new client model to obtain a trained client model, and upload the newly trained client model to the edge node.
[0162] Among them, when each client performs model update, it can update the parameters of all or part of the structure of the client model, so as to obtain a personalized client model adapted to each client locally and upload it to the edge node. Among them, when each client updates part of the structure of the client model, the updated part of the structure can be determined by the server or by the edge node.
[0163] 503. The edge node groups based on the data structure characteristics of the clients.
[0164] Such as Figure 7 As shown in the figure, after receiving the trained client models uploaded by the accessed clients, the edge node can group each client based on the information of the client models uploaded by each client. Among them, when the edge node performs grouping, the edge nodes can communicate with each other, and can transmit the information of the client models uploaded by the accessed clients, or the characteristics of the information of the client models uploaded by each client, etc. Therefore, cross-edge node client grouping can be realized, so that more clients with similar data distributions can be divided into one group, thereby improving the training effect of the personalized model and avoiding the adverse effects of client models with too large data differences on the personalized model.
[0165] In the embodiment of the present application, exemplarily, taking the steps executed by one of the edge nodes as an example for exemplary introduction, the edge node therein can be any node in the federated learning system, which will not be elaborated below.
[0166] Optionally, clients with similar data structure characteristics can be divided into one group. For example, clustering can be performed on the information of the client models uploaded by each client, and the clients corresponding to the model information in the same category can be divided into one group.
[0167] In addition, when determining the edge-side grouping, the similarity between the client model information uploaded by the clients in each grouping and the local edge-node model of the edge node can also be combined to determine the clients in the edge-side grouping. For example, the cosine similarity between the client model information uploaded by each client and the edge-node model of the edge node can be calculated, and the clients with a similarity higher than the preset similarity are classified into the edge-side grouping, that is, the client is used as the client in the edge-side grouping of the edge node. Therefore, when the edge node performs model aggregation, the local personalized model features can be retained as much as possible to achieve better performance in personalized model training.
[0168] For example, when the edge node selects an end node to establish a logical connection, the cosine similarity is used to judge the difference between the data distribution of the end node and the model parameters locally saved by the edge node, and the K (positive integer) clients with the smallest difference are selected as the edge-side grouping to fully access the end node for personalized model training within the limited bandwidth resources.
[0169] Optionally, when the edge node performs grouping, the communication resources or computing resources within the group can also be considered for grouping. For example, the edge node can evaluate the comprehensive index of an end node based on the communication-related information such as the delay and bandwidth during the client's upload and the computing power information, so as to consider whether to access the end node for synchronization within the limited bandwidth and round time. For example, clients with high transmission delay or low computing power can be avoided from accessing to improve the efficiency of federated learning.
[0170] When the edge node performs grouping, it can be grouped according to a preset quantity, or the specific grouping quantity can be selected according to the actual application scenario.
[0171] For example, the silhouette score can be used as an index to evaluate the clustering effect, which measures the similarity between data points within the same cluster and the dissimilarity between data points in different clusters. The steps to calculate the grouping number k using the silhouette score can include: selecting a range of k values, where k is the grouping number; performing kmeans clustering for each k value; calculating the silhouette score for each k; and selecting the k value corresponding to the largest silhouette score as the grouping number. Therefore, it is possible to assist in obtaining the best grouping of clients in the case where the distributions of each edge node are unknown or it is difficult to ensure the security of all edge nodes. Thus, it can automatically help the user find the most suitable grouping number when the user is unsure of the grouping number, and a relatively optimal grouping can be found to improve the performance of the personalized model.
[0172] Of course, when grouping clients, in addition to using the silhouette score to evaluate groupings of various sizes, other evaluation metrics can also be adopted, such as the average intra-cluster distance or the average inter-cluster distance. Specifically, the appropriate evaluation metric can be selected according to the actual application scenario, and no further elaboration will be provided here.
[0173] In addition, in some scenarios, edge nodes can transmit data to each other, and edge nodes can transmit the client model information uploaded by the clients accessing each edge node to each other. Therefore, when grouping edge nodes, cross-edge node grouping can also be performed. For example, the clients accessing the first edge node can be divided into the edge-side grouping of the second edge node, so that clients with closer data distributions can be grouped across edge nodes, which is more conducive to implementing personalized model training.
[0174] 504. The edge node aggregates the client models uploaded by the clients within each grouping.
[0175] As Figure 8 shown, after the edge node receives the client model information uploaded by the accessing clients and performs grouping, it can aggregate the client model information uploaded by the clients within each grouping respectively to obtain one or more edge node models.
[0176] Generally, when performing model aggregation, the weighted formula: λ.h 1 +(1 - λ).h c can be referred to, where λ is the weight of the edge node model parameter h 1 of the edge node locally, and h c is the client model parameter uploaded by the client.
[0177] Among them, when multiple edge node models are obtained, such as when the clients are divided into multiple groups and one edge node model is aggregated for each group, the multiple obtained edge node models can be aggregated again to obtain the final edge node model. Of course, multiple edge node models can also be directly uploaded to the server.
[0178] When the clients under the edge node are only divided into one group, only the client model parameters uploaded by the clients in this one group can be aggregated.
[0179] Therefore, through the grouping method, scenarios with various data distribution structures can be adapted to achieve a greater degree of personalized model training.
[0180] Generally, after each grouping, an edge node can upload only one or more layers of parameters in the model, which reduces the communication overhead and avoids information leakage caused by directly transmitting data distribution. For example, an edge node can upload only the parameters of one or more network structures in the globally updated model that are more sensitive than a preset sensitivity or whose change value exceeds a preset change value.
[0181] In addition, the data structure characteristics between clients accessing different edge nodes may also be similar. Therefore, clients accessing different edge nodes can also be divided into a group. In this scenario, edge nodes can also transmit the client model information uploaded by clients across edge node groups.
[0182] Optionally, when an edge node needs to transmit the client model information uploaded by multiple clients, the sending edge node can aggregate the client model information uploaded by the multiple clients and then transmit it to the receiving edge node, or directly transmit the client model information uploaded by the multiple clients. Specifically, an appropriate transmission method can be selected according to the actual application scenario, and this application does not make any limitations in this regard.
[0183] For example, the information of the model transmitted between edge nodes in different scenarios can be as Figure 9 or Figure 10 shown.
[0184] That is, optionally, step 5041 can be executed: transmit the client model uploaded by clients across node groups between edge nodes.
[0185] For example, Figure 9 shown, if a direct communication connection is established between edge nodes, the edge nodes can directly transmit the model information that needs to be transmitted. Such a direct communication connection can be established, for example, between edge nodes through short-distance communication (such as Bluetooth or Wi-Fi, etc.) or directly through a wired connection.
[0186] For example, Figure 10 shown, data can also be transmitted between edge nodes through a backbone network. For example, when edge nodes cannot communicate directly, the model information that needs to be transmitted between edge nodes can be transmitted through the backbone network. Such a backbone network can be the backbone network in the communication network accessed by the edge nodes.
[0187] In this scenario, in the next iteration process, the second edge node can send a new global model to the clients connected across nodes, so that the clients can use the new global model adapted to the local data distribution for downstream tasks.
[0188] 505. The edge node uploads the edge node model.
[0189] For example, Figure 11As shown, after aggregating the client models uploaded by the clients in each group at the edge node to obtain one or more edge node models, the edge node can upload the information of the one or more edge node models to the server, such as the weight parameters or gradients of the edge node models.
[0190] In the embodiments of the present application, the clients for training the personalized model are grouped according to the differences, effectively avoiding the interference of nodes with large differences or potential malicious attacks on personalized learning. Considering the impact of the differences in personalized models on model training, by clustering and grouping the edge nodes according to their model parameters, the interference caused by edge nodes with large differences to model training is eliminated. The greater the data distribution difference between clients, the more obvious the advantages of the grouped models; at the same time, both bandwidth resources and distribution differences are considered in the selection of client groups, so that the edge nodes can select high-quality end nodes as much as possible, improving the performance of the personalized model.
[0191] 506. The server aggregates to obtain a new global model.
[0192] Such as Figure 12 As shown, after the server receives one or more edge node models, it can update the local global model based on the received one or more edge node models to obtain an updated global model.
[0193] After obtaining the updated global model, the server can send the new global model to each edge node to enter the next round of iteration, that is, the server can continue to send the new global model to the edge node, that is, repeat step 501 and subsequent steps, so that each client can keep the global model updated synchronously and use the latest model for downstream tasks.
[0194] In the embodiments of the present application, the server can send the global model to the edge node, and the edge node groups based on the data structure characteristics of the client model information uploaded by each client, so as to facilitate the subsequent aggregation of the client models uploaded by clients with closer data distributions, and retain the personalized characteristics of clients with closer distributions as much as possible, realizing a personalized model with better performance. It is equivalent to restricting the logical connection between clients and edge nodes without establishing a physical connection, ensuring that the distribution differences between the end nodes selected by the edge nodes are as small as possible under limited bandwidth resources, and improving the performance of the personalized model. And the edge node can aggregate the client models uploaded by the connected clients and only needs to upload the aggregated model to the server, without uploading all the client models uploaded by the clients to the server. Therefore, the amount of data that needs to be transmitted to the server is greatly reduced, and the communication overhead and the load of the server are greatly reduced. And after grouping the clients, the global model within the group only communicates and synchronizes within the group, further reducing the communication pressure on the cloud server.
[0195] The method flow provided by the present application has been introduced above. Next, the device for executing the foregoing method will be introduced.
[0196] Refer to Figure 13 , Figure 13 which is a schematic structural diagram of a federated learning device provided by the present application. This device can be used to execute the steps performed by the edge node in the foregoing Figures 4 to 12 . The device includes the following:
[0197] A transceiver module 1301, configured to obtain data uploaded by at least one client. The at least one client includes clients accessing multiple edge nodes in the federated learning system, and the federated learning device is any one of the multiple edge nodes;
[0198] A grouping module 1302, configured to group at least one client according to the information of the client models uploaded by the at least one client, so as to obtain an edge-side grouping corresponding to the federated learning device;
[0199] An aggregation module 1303, configured to aggregate the information of the client models uploaded by the clients in the edge-side grouping to obtain the information of the edge node model;
[0200] The transceiver module 1301 is further configured to send the information of the edge node model to the server. The server is configured to aggregate the information of the edge node models uploaded by multiple received edge nodes to obtain a global model;
[0201] The transceiver module 1301 is further configured to receive the aggregated global model sent by the server;
[0202] The transceiver module 1301 is further configured to send the aggregated global model to each client accessing the first edge node, so that each client uses the aggregated global model to update the client model locally saved by each client, and trains the updated client model using the locally saved data to obtain a trained client model. The information of the trained client model is used to be uploaded to the first edge node again.
[0203] In a possible implementation manner, the transceiver module 1301 is further configured to obtain the information of the client models uploaded by at least one client;
[0204] The grouping module 1302 is configured to group at least one client according to the information of the client models uploaded by the at least one client to obtain at least one end-side grouping;
[0205] The grouping module 1302 is configured to determine the edge-side grouping of the federated learning device according to at least one end-side grouping.
[0206] In a possible implementation, the grouping module 1302 is specifically configured to cluster the information of the client models uploaded by at least one client, group the at least one client according to the clustering result, and obtain at least one edge-side grouping.
[0207] In a possible implementation, the grouping module 1302 is specifically configured to: cluster the information of the client models uploaded by at least one client to obtain a clustering result; group the at least one client according to the clustering result and the number of groupings to obtain at least one edge-side grouping; evaluate the at least one edge-side grouping to obtain an evaluation value of the at least one edge-side grouping; if the evaluation value does not meet the preset evaluation condition, adjust the number of groupings, and re-group the at least one client according to the adjusted number of groupings.
[0208] In a possible implementation, the grouping module 1302 is specifically configured to: obtain the similarity between the model information saved locally and the information of the client models uploaded by each client among the at least one client, where the model information saved locally includes the information of the edge node model obtained in the previous iteration process; based on the similarity between the model information saved locally and the information of the client models uploaded by each client, determine the edge-side grouping according to the at least one edge-side grouping.
[0209] In a possible implementation, the grouping module 1302 is specifically configured to: obtain the communication resources corresponding to at least one client, where the communication resources include the communication resources used when at least one client communicates with the federated learning device; combine the communication resources corresponding to the clients in at least one grouping, and determine the edge-side grouping according to the at least one edge-side grouping.
[0210] In a possible implementation, the transceiver module 1301 is further configured to, if the federated learning device determines that the at least one edge-side grouping further includes the edge-side grouping corresponding to the second edge node, which can also be referred to as the first edge-side grouping for easy distinction, for example, if the similarity between the information of the client model uploaded by the clients in one of the edge-side groupings and the edge node model of the second edge node is relatively high, or the information of the client model uploaded by the clients in this edge-side grouping and the information of the client models uploaded by the clients in the edge-side grouping of the second edge node belong to the same category (i.e., the category after clustering), then send the information of the client models uploaded by the clients in the edge-side grouping corresponding to the second edge node to the second edge node.
[0211] In a possible implementation, the transceiver module 1301 is further configured to: send the information of the client models uploaded by the clients in the edge-side grouping corresponding to the second edge node to the second edge node through the connection established with the second edge node; or send the information of the client models uploaded by the clients in the edge-side grouping corresponding to the second edge node to the second edge node through the backbone network.
[0212] In a possible implementation, the transceiver module 1301 is further configured to send information of at least one network layer in the edge node model to the server.
[0213] In a possible implementation, the grouping module 1302 is further configured to: obtain the similarity between the model information saved locally by the first edge node and the information of the models uploaded by at least one client, where the model information saved locally may be the information of the edge node model obtained in the previous iteration process; screen at least one client according to the similarity between the model information saved locally and the information of the models uploaded by at least one client to obtain an edge-side grouping. For example, select the N (positive integer) clients with the highest similarity as the clients in the edge-side grouping.
[0214] See Figure 14 , a schematic structural diagram of another federated learning device provided by an embodiment of the present application. This device can be used to execute the steps performed by the server in the foregoing Figures 4 to 12 . This device includes:
[0215] A transceiver module 1401, which receives the information of the edge node models respectively sent by multiple edge nodes. The information of the edge node models is obtained by aggregating, for each edge node among the multiple edge nodes, the information of the client models uploaded by the clients in at least one edge-side grouping. At least one edge-side grouping is obtained by each edge node grouping at least one terminal according to the data uploaded by at least one terminal accessing the multiple edge nodes;
[0216] An aggregation module 1402, configured to aggregate the information of the edge node models to obtain a global model;
[0217] The transceiver module 1401 is further configured to send the new aggregated global model to each edge node below, so that each edge node distributes the new global model to each client, enabling each client to update the client model saved locally based on the new global model, and training the updated client model using the training data saved locally to obtain a trained model and upload it to the edge node again for the next iteration.
[0218] In a possible implementation, the transceiver module 1401 is further configured to send the information of the global model to each edge node.
[0219] As Figure 15 shown, it is a schematic hardware structure diagram of a federated learning device 150 provided by an embodiment of the present application. This federated learning device can be used to execute the steps performed by the edge node in the foregoing Figures 5 to 12 .
[0220] Figure 15The federated learning device 150 shown may include: a processor 1501, a memory 1502, a communication interface 1503, and a bus 1504. The processor 1501, the memory 1502, and the communication interface 1503 may be connected through the bus 1504.
[0221] The processor 1501 is the control center of the federated learning device 150, and may be a general-purpose central processing unit (CPU), or other general-purpose processors, etc. Among them, the general-purpose processor may be a microprocessor or any conventional processor, etc.
[0222] As an example, the processor 1501 may include one or more CPUs, such as Figure 15 CPU 0 and CPU 1 shown in
[0223] The memory 1502 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a magnetic disk storage medium, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0224] In a possible implementation, the memory 1502 may exist independently of the processor 1501. The memory 1502 may be connected to the processor 1501 through the bus 1504 for storing data, instructions, or program code. When the processor 1501 calls and executes the instructions or program code stored in the memory 1502, the method provided in the embodiments of the present application can be implemented.
[0225] In another possible implementation, the memory 1502 may also be integrated with the processor 1501.
[0226] The communication interface 1503 is used for the federated learning device 150 to be connected to other devices through a communication network, and the communication network may be an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc. The communication interface 1503 may include a receiving unit for receiving data and a sending unit for sending data.
[0227] The bus 1504 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. This bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 15 it is only represented by a thick line in the figure, but it does not mean that there is only one bus or one type of bus.
[0228] It should be noted that Figure 15 the structure shown in the figure does not constitute a limitation on the federated learning device 150. Except Figure 15 for the components shown, the federated learning device 150 may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0229] Optionally, the aforementioned Figure 15 federated learning device shown in the figure is a chip.
[0230] As Figure 16 shown, it is a schematic diagram of the hardware structure of a federated learning device 160 provided by an embodiment of the present application. This federated learning device can be used to execute the steps performed by the server in the aforementioned Figures 5 to 12 figure.
[0231] Figure 16 The federated learning device 160 shown in the figure may include: a processor 1601, a memory 1602, a communication interface 1603, and a bus 1604. The processor 1601, the memory 1602, and the communication interface 1603 can be connected through the bus 1604.
[0232] The processor 1601 is the control center of the federated learning device 160. It can be a general-purpose central processing unit (CPU), or other general-purpose processors, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0233] As an example, the processor 1601 may include one or more CPUs, such as Figure 16 the CPU 0 and CPU 1 shown in the figure.
[0234] The memory 1602 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or can also be an electrically erasable programmable read-only memory (EEPROM), a magnetic disk storage medium, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0235] In a possible implementation, the memory 1602 can exist independently of the processor 1601. The memory 1602 can be connected to the processor 1601 through the bus 1604 and is used to store data, instructions, or program code. When the processor 1601 invokes and executes the instructions or program code stored in the memory 1602, the method provided by the embodiments of the present application can be implemented.
[0236] In another possible implementation, the memory 1602 can also be integrated with the processor 1601.
[0237] The communication interface 1603 is used to connect the federated learning device 160 to other devices through a communication network. The communication network can be an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc. The communication interface 1603 can include a receiving unit for receiving data and a transmitting unit for transmitting data.
[0238] The bus 1604 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 16 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0239] It should be noted that Figure 16 the structure shown in the figure does not constitute a limitation on the federated learning device 160, exceptFigure 16 In addition to the components shown, the federated learning device 160 may include more or fewer components than shown, or combine certain components, or have a different component arrangement.
[0240] Optionally, the aforementioned Figure 16 The federated learning device shown in is a chip.
[0241] The embodiment of the present application also provides a digital processing chip. Circuits for implementing the aforementioned processor 1501 or 1601, or the functions of processor 1501 or 1601, and one or more interfaces are integrated in the digital processing chip. When a memory is integrated in the digital processing chip, the digital processing chip can complete the method steps of any one or more of the foregoing embodiments. When a memory is not integrated in the digital processing chip, it can be connected to an external memory through a communication interface. The digital processing chip implements the actions performed by the federated learning device in the foregoing embodiments according to the program code stored in the external memory.
[0242] The embodiment of the present application also provides a computer program product, which, when running on a computer, causes the computer to execute the steps in the method described in the foregoing Figures 5 - 12 shown embodiments.
[0243] The federated learning device provided in the embodiment of the present application may be a chip, and the chip includes: a processing unit and a communication unit. The processing unit may be a processor, for example, and the communication unit may be an input / output interface, a pin, a circuit, etc. The processing unit can execute the computer execution instructions stored in the storage unit to enable the chip to execute the method described in the Figures 5 - 12 shown embodiments. Optionally, the storage unit is a storage unit inside the chip, such as a register, a cache, etc. The storage unit may also be a storage unit outside the chip and inside the radio access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0244] Specifically, the foregoing processing unit or processor may be a central processing unit (CPU), a neural-network processing unit (NPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processing unit or processor may be used to execute the foregoing Figures 5 - 12 steps of the corresponding method.
[0245] In addition, it should be noted that the device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.
[0246] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions accomplished by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits. However, for this application, software programs are a better implementation method in more cases. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, read only memory (ROM), random access memory (RAM), magnetic disk, or optical disc, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments of this application.
[0247] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0248] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0249] The terms "first", "second", "third", "fourth", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order other than that shown or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.
[0250] Finally, it should be noted that the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, and all should be covered within the protection scope of this application.
Claims
1. A federated learning method, characterized in that, applied to a federated learning system, the federated learning system includes a server, a plurality of edge nodes, and a plurality of clients. The plurality of edge nodes are connected to the server, and the plurality of clients are connected to the plurality of edge nodes. The method includes a plurality of iterative processes, where one iterative process is as follows: The first edge node obtains information of the client models uploaded by at least one client. The at least one client includes the clients connected to the plurality of edge nodes. The first edge node is any one of the plurality of edge nodes; The first edge node groups the at least one client according to the information of the client models uploaded by the at least one client to obtain an edge-side grouping corresponding to the first edge node; The first edge node aggregates the information of the client models uploaded by the clients in the edge-side grouping to obtain information of an edge node model; The first edge node sends the information of the edge node model to the server, and the server is used to aggregate the information of the edge node models uploaded by the plurality of received edge nodes to obtain information of a global model; The first edge node receives the aggregated global model sent by the server; The first edge node sends the aggregated global model to each client connected to the first edge node, so that each client uses the aggregated global model to update the client model locally saved by each client, and trains the updated client model using the locally saved data to obtain a trained client model. The information of the trained client model is used to be uploaded to the first edge node again.
2. The method according to claim 1, characterized in that, The first edge node groups the at least one client according to the information of the models uploaded by the at least one client to obtain an edge-side grouping corresponding to the first edge node, including: The first edge node groups the at least one client according to the information of the client models uploaded by the at least one client to obtain at least one end-side grouping; The first edge node determines the edge-side grouping corresponding to the first edge node according to the at least one end-side grouping.
3. The method according to claim 2, characterized in that, The first edge node groups the at least one client according to the information of the client models uploaded by the at least one client to obtain at least one end-side grouping corresponding to the first edge node, including: The first edge node clusters the information of the client models uploaded by the at least one client, and groups the at least one client according to the clustering result to obtain at least one end-side grouping.
4. The method according to claim 3, characterized in that, The first edge node clusters the information of the client models uploaded by the at least one client, and groups the at least one client according to the clustering result to obtain at least one end-side grouping, including: The first edge node clusters the information of the client models uploaded by the at least one client to obtain a clustering result; The first edge node obtains the number of groups, where the number of groups includes the number of clients included in each edge-side group; The first edge node groups the at least one client according to the clustering result and the number of groups, to obtain at least one edge-side group; The first edge node evaluates the at least one edge-side group to obtain an evaluation value of the at least one edge-side group; If the evaluation value does not meet the preset evaluation condition, the first edge node adjusts the number of groups, and re-groups the at least one client according to the adjusted number of groups.
5. The method according to any one of claims 2-4, wherein, The first edge node determines the edge-side group corresponding to the first edge node according to the at least one edge-side group, including: The first edge node obtains the similarity between the model information saved locally and the information of the client model uploaded by each client among the at least one client, and the model information saved locally is the information of the edge node model obtained in the previous iteration process; The first edge node determines the edge-side group corresponding to the first edge node from the at least one edge-side group based on the similarity between the model information saved locally and the information of the client model uploaded by each client.
6. The method according to any one of claims 2-5, wherein, The first edge node determining the edge-side group of the first edge node according to the at least one edge-side group further includes: The first edge node also obtains the communication resources corresponding to the at least one client, where the communication resources include the communication resources used when the at least one client communicates with the first edge node; The first edge node combines the communication resources corresponding to the clients in the at least one group, and determines the edge-side group according to the at least one edge-side group.
7. The method according to any one of claims 2-6, wherein, The method further includes: If the first edge node determines that the at least one edge-side group further includes a first edge-side group corresponding to a second edge node, the first edge node sends the information of the client model uploaded by the clients in the edge-side group corresponding to the second edge node to the second edge node, where the second edge node is any one of the multiple edge nodes other than the first edge node, and the information of the client model of the first edge-side group and the information of the client model of the edge-side group of the second edge node belong to the same category, and the same category is the same clustering category after clustering the information of the client model.
8. The method according to claim 7, wherein, The first edge node sending the information of the client model uploaded by the clients in the edge-side group corresponding to the second edge node to the second edge node includes: The first edge node sends the information of the client model uploaded by the clients in the edge-side group corresponding to the second edge node to the second edge node through a direct connection established with the second edge node; Alternatively, the first edge node sends information of the client model uploaded by the client in the end - side packet corresponding to the second edge node to the second edge node through the backbone network in the communication network, where the communication network is the network where the first edge node and the edge node are located.
9. The method according to any one of claims 1 - 8, wherein, the first edge node sending the information of the edge node model to the server includes: the first edge node sending information of at least one network layer in the edge node model to the server.
10. The method according to any one of claims 1 - 9, wherein, the first edge node grouping the at least one client according to the information of the models uploaded by the at least one client to obtain an edge - side packet corresponding to the first edge node, further includes: the first edge node obtaining the similarity between the model information locally saved and the information of the models uploaded by the at least one client, where the locally saved model information is the information of the edge node model obtained in the previous iteration process; screening the at least one client according to the similarity between the locally saved model information and the information of the models uploaded by the at least one client to obtain the edge - side packet.
11. The method according to any one of claims 1 - 10, wherein, the client model locally saved by each client is used to perform at least one of a recommendation task, a classification task, a segmentation task, or an object detection task.
12. A federated learning method, wherein, applied to a federated learning system, the federated learning system includes a server, a plurality of edge nodes, and a plurality of clients, each edge node is connected to the server, and the plurality of clients are connected to the plurality of edge nodes. The method includes: the server receives information of edge node models respectively sent by the plurality of edge nodes, where the information of the edge node models is obtained by aggregating information of client models uploaded by clients in at least one edge - side packet by each edge node in the plurality of edge nodes, and the at least one edge - side packet is obtained by each edge node grouping at least one terminal connected to the plurality of edge nodes according to the information of client models uploaded by the at least one terminal; the server aggregates the information of the edge node models to obtain information of a global model.
13. The method according to claim 12, wherein, the method further includes: the server sending the information of the global model to each edge node.
14. A federated learning device, wherein, applied to a federated learning system, the federated learning system includes a server, a plurality of edge nodes, and a plurality of clients, each edge node is connected to the server, and the plurality of clients are connected to the plurality of edge nodes. The federated learning device is any one of the plurality of edge nodes, and the federated learning includes a plurality of iteration processes. In one of the iteration processes, the device includes: A transceiver module, configured to obtain information of models uploaded by at least one client, where the at least one client includes clients accessing the multiple edge nodes, and the federated learning device is any one of the multiple edge nodes; A grouping module, configured to group the at least one client according to the information of the client models uploaded by the at least one client, so as to obtain an edge-side grouping corresponding to the federated learning device; An aggregation module, configured to aggregate the information of the client models uploaded by the clients in the edge-side grouping, so as to obtain information of an edge node model; The transceiver module is further configured to send the information of the edge node model to the server, and the server is configured to aggregate the information of the edge node models uploaded by the multiple received edge nodes to obtain a global model; The transceiver module is further configured to receive the aggregated global model sent by the server; The transceiver module is further configured to send the aggregated global model to each client accessing the first edge node, so that each client uses the aggregated global model to update the client model locally saved by each client, and trains the updated client model by using the data locally saved, so as to obtain a trained client model, and the information of the trained client model is used to be uploaded to the first edge node again.
15. The apparatus according to claim 14, wherein, the transceiver module is further configured to obtain the information of the client models uploaded by the at least one client; the grouping module is configured to group the at least one client according to the information of the client models uploaded by the at least one client, so as to obtain at least one end-side grouping; the grouping module is configured to determine the edge-side grouping of the federated learning device according to the at least one end-side grouping.
16. The apparatus according to claim 15, wherein, the grouping module is specifically configured to cluster the information of the client models uploaded by the at least one client, and group the at least one client according to the clustering result, so as to obtain at least one end-side grouping.
17. The apparatus according to claim 16, wherein, the grouping module is specifically configured to: cluster according to the information of the client models uploaded by the at least one client to obtain a clustering result; obtain a grouping number, where the grouping number includes the number of clients included in each end-side grouping; group the at least one client according to the clustering result and the grouping number, so as to obtain at least one end-side grouping; evaluate the at least one end-side grouping to obtain an evaluation value of the at least one end-side grouping; if the evaluation value does not meet a preset evaluation condition, adjust the grouping number, and re-group the at least one client according to the adjusted grouping number.
18. The apparatus according to any one of claims 15-17, wherein, the grouping module is specifically configured to: Obtain the similarity between the locally saved model information and the information of the client models uploaded by each of the at least one client, where the locally saved model information is the information of the edge node model obtained in the previous iteration process; Based on the similarity between the locally saved model information and the information of the client models uploaded by each of the at least one client, determine the edge-side grouping corresponding to the federated learning device from the at least one edge-side grouping.
19. The device according to any one of claims 15-18, characterized in that, The grouping module is specifically configured to: Obtain the communication resources corresponding to the at least one client, where the communication resources include the communication resources used when the at least one client communicates with the federated learning device; Combine the communication resources corresponding to the clients in the at least one grouping, and determine the edge-side grouping according to the at least one edge-side grouping.
20. The device according to any one of claims 15-19, characterized in that, The transceiver module is further configured to, if the federated learning device determines that the at least one edge-side grouping further includes a first edge-side grouping corresponding to a second edge node, send the information of the client model uploaded by the clients in the edge-side grouping corresponding to the second edge node to the second edge node, where the second edge node is any one of the multiple edge nodes other than the federated learning device, and the information of the client model in the first edge-side grouping and the information of the client model in the edge-side grouping of the second edge node belong to the same category, and the same category is the same clustering category after clustering the information of the client model.
21. The device according to claim 20, characterized in that, The transceiver module is further configured to: Send the information of the client model uploaded by the clients in the edge-side grouping corresponding to the second edge node to the second edge node through a direct connection established with the second edge node; Or, send the information of the client model uploaded by the clients in the edge-side grouping corresponding to the second edge node to the second edge node through the backbone network in the communication network, where the communication network is the network where the federated learning device and the second edge node are located.
22. The device according to any one of claims 14-21, characterized in that, The transceiver module is further configured to send the information of at least one network layer in the edge node model to the server.
23. The device according to any one of claims 14-22, characterized in that, The grouping module is further configured to: Obtain the similarity between the locally saved model information and the information of the models uploaded by the at least one client, where the locally saved model information is the information of the edge node model obtained in the previous iteration process; Screen the at least one client according to the similarity between the locally saved model information and the information of the models uploaded by the at least one client, and obtain the edge-side grouping.
24. The device according to any one of claims 14-22, characterized in that, The client model locally saved by each client is used to perform at least one of a recommendation task, a classification task, a segmentation task, or an object detection task.
25. A federated learning device, characterized in that it is applied to a federated learning system, which includes a server, multiple edge nodes, and multiple clients. Each edge node is connected to the server, and the multiple clients are connected to the multiple edge nodes. The federated learning device is the server, and the device includes: a transceiver module that receives information on the edge node models respectively sent by the multiple edge nodes. The information on the edge node models is obtained by aggregating, by each of the multiple edge nodes, the information on the client models uploaded by the clients in at least one edge-side group. The at least one edge-side group is obtained by each edge node grouping at least one terminal based on the information on the client models uploaded by the at least one terminal connected to the multiple edge nodes; an aggregation module for aggregating the information on the edge node models to obtain a global model.
26. The device according to claim 25, characterized in that the device further includes: the transceiver module is further configured to send the information on the global model to each edge node.
27. A federated learning system, characterized in that it includes multiple servers, multiple edge nodes, and at least one client, and the at least one client is connected to the multiple edge nodes; any one of the multiple servers is configured to implement the steps of the method according to any one of claims 12 to 13; the multiple edge nodes are configured to implement the steps of the method according to any one of claims 1-11; the at least one client is configured to perform model training using locally saved data.
28. A federated learning device, characterized in that it includes one or more processors, the one or more processors are coupled to a memory, and the memory stores a program. When the program instructions stored in the memory are executed by the one or more processors, the steps of the method according to any one of claims 1 to 11 are implemented.
29. A federated learning device, characterized in that it includes one or more processors, the one or more processors are coupled to a memory, and the memory stores a program. When the program instructions stored in the memory are executed by the one or more processors, the steps of the method according to claim 12 or 13 are implemented.
30. A computer-readable storage medium, characterized in that it includes a program, and when it is executed by a processing unit, it executes the method according to any one of claims 1 to 13.
31. A computer program product including computer programs / instructions, characterized in that when the computer programs / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.