Federated learning method and apparatus
By realizing data interoperability and client grouping between edge nodes in the federated learning system, the model training problem under non-independent and homogeneous data is solved, and the personalization and overall performance of the model are improved.
Patent Information
- Application Number
- PCT/CN2024/125122
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-30
- Filing Date
- 2024-10-16
- Publication Date
- 2025-06-05
AI Technical Summary
In federated learning, when user data is not independently distributed simultaneously, the gradient directions of each user edge node are different after iteration, resulting in the global model generated by the central server performing poorly on each edge node and even failing to converge.
Through data interoperability between edge nodes, edge nodes can group clients with similar data distributions, thereby performing model aggregation in better personalized model training. The specific steps include: edge nodes obtain information about the client model, group and aggregate, generate edge node models, and upload them to the server for global model updates.
Through grouping and aggregation, edge nodes can more accurately group client groups, realize better personalized model training, and improve the performance and convergence of the model on each edge node.
Smart Images

Figure CN2024125122_05062025_PF_FP_ABST
Abstract
Description
A federated learning method and device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on November 30, 2023, with application number 202311627068.8 and application name “A Federated Learning Method and Device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of computing devices, and in particular to a federated learning method and apparatus. Background Art
[0003] Federated learning is a method that has emerged in recent years, enabling neural network model training without requiring data to leave the user's device. The architecture of a federated learning system can include a server and multiple clients. Clients train local models using locally stored data and upload the trained local model information to the server. The server aggregates the model information from multiple clients to create a global model.
[0004] When user data is non-independent and identically distributed (Non-IID), the gradient directions of each user's edge nodes vary significantly after iteration, causing the global model generated by the central server to perform poorly on each edge node, or even fail to converge. Therefore, how to train personalized models on edge nodes to facilitate local data prediction is an urgent problem that needs to be solved.
[0005] Summary of the Invention
[0006] The present application provides a federated learning method and device for enabling edge nodes to group clients with similar data distributions through data intercommunication between edge nodes, thereby achieving better personalized model training.
[0007] In view of this, on the first aspect, the present application provides a federated learning method, which is applied to a federated learning system, wherein the federated learning system includes a server, multiple edge nodes and multiple clients, the multiple edge nodes are connected to the server, and the multiple clients are connected to multiple edge nodes. The edge nodes can also be called edge servers, edge servers or intermediate nodes, etc. Taking the steps executed by any one of the nodes as an example, for the convenience of distinction, it is called the first edge node. The method includes: Generally, federated learning is a process of multiple iterative learning. The present application takes one of the iterative processes as an example: First, the first edge node obtains information of a client model uploaded by at least one client, and the information of the client model includes the information obtained by the at least one client using locally stored data to perform model training, such as parameters or gradients of the trained client model, etc. The model information mentioned below, such as information of the edge node model or information of the global model, etc., may also include parameters or gradients of the model, etc. The at least one client includes clients connected to multiple edge nodes, such as clients connected to the first edge node or other edge nodes; then, the first The edge node groups at least one client according to the information of the client model uploaded by at least one client to obtain the edge group corresponding to the first edge node, and the number of edge groups can be one group or more groups; then the first edge node aggregates the information of the client model uploaded by the clients in the edge group to obtain the information of the edge node model; then the first edge node sends the information of the edge node model to the server, so that the server re-aggregates after receiving the information of one or more edge node models to obtain a new global model; then the server can send the aggregated global model to each edge node, that is, the first edge node receives the aggregated global model sent by the server; then, the first edge node sends the aggregated global model to each client connected to the first edge node, so that each client uses the aggregated global model to update the client model locally saved by each client, and uses the locally saved data to train the updated client model to obtain the trained client model, and the information of the trained client model is used to upload to the first edge node again for the next iteration.
[0008] In the implementation of the present application, clients can be grouped based on the data distribution of the client model information uploaded by the clients. For example, clients with similar data distributions can be grouped together. When subsequently aggregated, the client model information uploaded by the clients in the group can be aggregated. Therefore, models with similar data distributions can be aggregated to achieve better personalized training. In addition, edge nodes can interact with each other, such as transmitting client model information uploaded by connected clients or parameters of local models of each edge node. Therefore, when grouping clients, client grouping across edge nodes can be achieved, thereby achieving more accurate grouping of each client in the federated learning system, which is more conducive to personalized model training and obtains personalized models with better performance.
[0009] In a possible implementation, the aforementioned first edge node groups at least one client according to the information of the client model uploaded by at least one client to obtain at least one group corresponding to the first edge node, which may include: first, the first edge node obtains the information of the client model uploaded by at least one client; then the first edge node groups at least one client according to the information of the client model uploaded by at least one client to obtain at least one end-side group; the first edge node determines the edge-side group of the first edge node according to the at least one end-side group.
[0010] In the implementation manner of the present application, clients can be grouped according to the information of the client models uploaded by each client. For example, clients with similar data distribution can be grouped together, so that when model aggregation is performed subsequently, models with similar distribution can be aggregated, and personalized models can be obtained by aggregation.
[0011] In a possible implementation, the aforementioned first edge node groups at least one client according to the information of the client model uploaded by at least one client to obtain at least one end-side group corresponding to the first edge node, which may include: the first edge node clusters the information of the client model uploaded by at least one client, and groups at least one client according to the clustering result to obtain at least one end-side group.
[0012] In an embodiment of the present application, when grouping clients, they can be grouped in a clustering manner. For example, clients corresponding to data in the same category can be grouped as a group of clients, or clients in each category whose distance from the cluster center point is less than a preset distance value can be grouped as a group of clients, etc., thereby grouping clients with closer data distribution into a group, realizing personalized model training within the group, and retaining as many personalized features as possible.
[0013] In a possible implementation, the aforementioned first edge node clusters the information of the client model uploaded by at least one client, groups the at least one client according to the clustering result, and obtains at least one end-side group, which may include: the first edge node clusters the information of the client model uploaded by at least one client to obtain a clustering result; the first edge node groups the at least one client according to the clustering result and the number of groups to obtain at least one end-side group; the first edge node evaluates the at least one end-side group to obtain an evaluation value of the at least one end-side group; if the evaluation value does not meet the preset evaluation conditions, the first edge node adjusts the number of groups, and re-groups the at least one client according to the adjusted number of groups.
[0014] In the implementation manner of the present application, when determining the number of clients in each group, the number of groups with the best evaluation results can be selected by evaluating the groups, thereby obtaining a more optimal grouping solution.
[0015] In a possible embodiment, the aforementioned first edge node determines the edge group of the first edge node based on at least one end side group, which may include: the first edge node obtains the similarity between the locally stored model information and the information of the client model uploaded by each client in at least one client, and the model information stored locally by the first node is the information of the edge node model obtained in the previous iteration process; the first edge node determines the edge group from at least one end side group based on the similarity between the locally stored model information and the information of the client model uploaded by each client, such as selecting a preset number of clients with the highest similarity or clients with similarity higher than the preset similarity as clients in the edge group.
[0016] In an embodiment of the present application, when an edge node selects clients to access a group, it can calculate the similarity between the clients in the group and the data stored by the edge node, such as the local model of the edge node, and then select the clients to access based on this similarity. For example, clients whose similarity to the local model of the edge node is higher than a preset similarity can be selected, thereby reducing the interference caused by the differences between the edge node model and the models of each client in model aggregation, and obtaining a personalized model that is more suitable for the clients in each group.
[0017] In a possible implementation, the aforementioned first edge node determines the edge group of the first edge node based on at least one end side group, and may also include: the first edge node also obtains communication resources corresponding to at least one client, and the communication resources include communication resources used by at least one client to communicate with the first edge node, such as the computing power of each client, the available bandwidth of the first edge node, the computing power of the first edge node, or the communication delay between the first edge node and the client; then the first edge node combines the communication resources corresponding to the clients in at least one group, and determines the edge group based on at least one end side group, such as selecting clients with lower latency, higher computing power or less bandwidth occupancy to access the edge group.
[0018] In the embodiments of the present application, the client accessing the edge group can be selected based on the communication resources when the edge node communicates with the client. Therefore, when selecting a client to access, the edge node can select a client that is compatible with the communication resources, thereby balancing the communication load of the edge node and improving the learning efficiency of federated learning.
[0019] In a possible embodiment, the aforementioned method may further include: if the first edge node determines that at least one end-side group also includes an end-side group corresponding to the second edge node, which can also be referred to as the first end-side group, such as the information of the client model uploaded by the client in one group of end-side groups has a high similarity with the edge node model of the second edge node, or the information of the client model uploaded by the client in the end-side group and the information of the client model uploaded by the client in the edge group of the second edge node belong to the same category (i.e., the clustered category), then the end-side group can be divided into the edge group corresponding to the second edge node. In this scenario, the first edge node sends the information of the client model uploaded by the client in the end-side group corresponding to the second edge node to the second edge node; when there are multiple clients divided into the edge group of the second edge node, the client model information uploaded by the multiple clients can be directly sent to the second edge node, or the first edge node can aggregate the client model information uploaded by the multiple clients and send the aggregated model information to the second edge node to reduce communication overhead.
[0020] In the implementation manner of the present application, grouping can be performed across edge nodes. For example, edge nodes can transmit their own stored local models. When each edge node groups the connected clients (or clients with which a connection has been established), it can also calculate the similarity between the client model information uploaded by the client and the local models of other edge nodes, and divide the clients with higher similarity to other edge nodes (such as similarity higher than a preset similarity value) into the edge groups of other edge nodes, thereby realizing client grouping across edge nodes. After grouping, the client model information uploaded by the clients divided into other edge nodes is sent to other edge nodes.
[0021] In a possible embodiment, the aforementioned first edge node sends information about the client model uploaded by the client in the end side group corresponding to the second edge node to the second edge node, which may include: the first edge node sends information about the client model uploaded by the client in the end side group corresponding to the second edge node to the second edge node through a direct connection established between the first edge node and the second edge node, and the direct connection is such as short-distance communication or wired connection; or, the first edge node sends information about the client model uploaded by the client in the end side group corresponding to the second edge node to the second edge node through a backbone network in a communication network, and the communication network is the network where the first edge node and the second edge node are located.
[0022] In the implementation manner of the present application, edge nodes can communicate directly with each other or through the backbone network. Therefore, the client model information uploaded by the client can be transmitted directly between the edge nodes, or the client model information uploaded by the client can be transmitted through the backbone network, thereby realizing data interaction between the edge nodes.
[0023] In a possible implementation, the aforementioned first edge node sending information of the edge node model to the server may include: the first edge node sending information of at least one network layer in the edge node model to the server, such as parameters or gradients of one or more network layers, etc., that is, when uploading the model at the edge node, it can be uploaded in layers, such as only uploading one or more layers, thereby reducing communication overhead.
[0024] In an embodiment of the present application, when the first edge node uploads the edge node model to the server, it can upload the parameters of all or part of the network layers of the edge node model, such as transmitting a network layer whose sensitivity during training is higher than a preset sensitivity value or whose change value exceeds a preset change value, thereby reducing communication overhead.
[0025] In a possible implementation, the aforementioned first edge node groups at least one client according to the information of the model uploaded by at least one client to obtain the edge group corresponding to the first edge node, and may also include: the first edge node obtains the similarity between the locally stored model information and the information of the model uploaded by at least one client, and the locally stored model information may be the information of the edge node model obtained in the last iteration process; based on the similarity between the locally stored model information and the information of the model uploaded by at least one client, at least one client is screened to obtain the edge group, such as selecting the N (positive integer) clients with the highest similarity as the clients in the edge group. Therefore, in the implementation of the present application, the edge node can directly filter out the clients of the edge group from one or more clients, so that the model that is more similar to the local model of the edge node can be determined more efficiently.
[0026] In one possible implementation, the client model stored locally on each client is used to perform at least one of a recommendation task, a classification task, a segmentation task, or an object detection task. This is equivalent to deploying the aggregated global model on each client. The global model obtained using the method provided herein is more compatible with the distribution of local data on each client, enabling personalized model training for each client. This allows the client to combine the aggregated global model to perform local, personalized reasoning tasks, such as recommendation, classification, segmentation, or object detection.
[0027] In a second aspect, the present application provides a federated learning method, which is applied to a federated learning system. The federated learning system includes a server, multiple edge nodes and multiple clients, and multiple clients access multiple edge nodes. The method includes: in one round of iteration, the server receives information of edge node models sent by multiple edge nodes respectively, or called edge node models, and the information of the edge node model includes each edge node in the multiple edge nodes aggregating information of client models uploaded by clients in at least one edge group, and at least one edge group is obtained by grouping at least one terminal according to data uploaded by at least one terminal accessing multiple edge nodes for each edge node; then the server aggregates the received edge node model information to obtain a global model; and sends the aggregated new global model to each edge node, so that each edge node sends the new global model to each client, so that each client updates the locally stored client model based on the new global model, and uses the locally stored training data to train the updated client model, obtains the trained model and uploads it to the edge node again for the next iteration.
[0028] In this embodiment, the server only needs to receive the aggregated models uploaded by edge nodes, eliminating the need to receive information uploaded by all clients, reducing communication overhead. Furthermore, edge nodes can communicate with each other, enabling cross-edge node client grouping. This allows for personalized client grouping, enabling better personalized model training and resulting in a more performant personalized model.
[0029] In one possible implementation, during one iteration, the method may further include: the server sending global model information to each edge node. Thus, the server may send a new global model to each edge node, thereby protecting and updating the global model of each edge node.
[0030] In a third aspect, the present application provides a federated learning device, which is applied to a federated learning system. The federated learning system includes a server, multiple edge nodes, and multiple clients. The multiple clients access the multiple edge nodes. The federated learning device is any one of the multiple boundary nodes. The device includes:
[0031] a transceiver module, configured to obtain data uploaded by at least one client, wherein the at least one client includes a client connected to multiple edge nodes, and the first edge node is any one of the multiple edge nodes;
[0032] a grouping module, configured to group at least one client according to information of a client model uploaded by at least one client, and obtain side groups corresponding to the federated learning device;
[0033] Aggregation module, used to aggregate the client model information uploaded by the clients in the edge group to obtain the edge node model information;
[0034] The transceiver module is also used to send edge node model information to the server, and the edge node model information is used by the server to aggregate and obtain the global model;
[0035] The transceiver module is also used to receive the aggregated global model sent by the server;
[0036] The transceiver module is also used to send the aggregated global model to each client connected to the first edge node, so that each client uses the aggregated global model to update the client model stored locally on each client, and uses the locally stored data to train the updated client model to obtain the trained client model. The information of the trained client model is used to upload it to the first edge node again.
[0037] In a possible implementation, the transceiver module is further configured to obtain information of a client model uploaded by at least one client;
[0038] a grouping module, configured to group at least one client according to information of a client model uploaded by at least one client, to obtain at least one terminal-side group;
[0039] The grouping module is configured to determine an edge-side grouping of the federated learning device according to at least one end-side grouping.
[0040] In a possible implementation, the grouping module is specifically configured to perform clustering based on information of a client model uploaded by at least one client, and group the at least one client based on the clustering result to obtain at least one terminal-side group.
[0041] In one possible implementation, the grouping module is specifically used to: cluster information of a client model uploaded by at least one client to obtain a clustering result; group at least one client according to the clustering result and the number of groups to obtain at least one end-side group; evaluate at least one end-side group to obtain an evaluation value of at least one end-side group; if the evaluation value does not meet a preset evaluation condition, adjust the number of groups, and re-group at least one client according to the adjusted number of groups.
[0042] In one possible implementation, the grouping module is specifically used to: obtain the similarity between the locally stored model information and the information of the client model uploaded by each client in at least one client, where the locally stored model information is the information of the edge node model obtained in the previous iteration process; based on the similarity between the locally stored model information and the information of the client model uploaded by each client, determine the edge-side grouping according to at least one end-side grouping.
[0043] In one possible implementation, the grouping module is specifically used to: obtain communication resources corresponding to at least one client, the communication resources including communication resources used when at least one client communicates with the first edge node; and determine the edge group corresponding to the first edge node from at least one end-side group in combination with the communication resources corresponding to the client in at least one group.
[0044] In a possible embodiment, the transceiver module is also used for: if the first edge node determines that at least one end side group also includes an end side group corresponding to the second edge node, which can also be called a first end side group, such as the information of the client model uploaded by the client in one group of end side groups has a high similarity with the edge node model of the second edge node, or the information of the client model uploaded by the client in the end side group and the information of the client model uploaded by the client in the edge side group of the second edge node belong to the same category (i.e., the clustered category), then the information of the client model uploaded by the client in the end side group corresponding to the second edge node is sent to the second edge node.
[0045] In one possible embodiment, the transceiver module is also used to: send information about the client model uploaded by the client in the end-side group corresponding to the second edge node to the second edge node through a connection established with the second edge node; or send information about the client model uploaded by the client in the end-side group corresponding to the second edge node to the second edge node through a backbone network.
[0046] In a possible implementation, the transceiver module is further configured to send information of at least one network layer in the edge node model to the server.
[0047] In one possible embodiment, the grouping module is also used for: the first edge node obtains the similarity between the locally stored model information and the model information uploaded by at least one client, and the locally stored model information may be the information of the edge node model obtained during the previous iteration; according to the similarity between the locally stored model information and the model information uploaded by at least one client, at least one client is screened to obtain an edge group, for example, selecting the N (positive integer) clients with the highest similarity as the clients in the edge group.
[0048] In one possible implementation, the client model stored locally on each client is used to perform at least one of a recommendation task, a classification task, a segmentation task, or an object detection task. This aggregated global model can be deployed on the client. The global model obtained using the method provided herein is more compatible with the distribution of local data on each client, enabling personalized model training for each client. This allows the client to combine the aggregated global model to perform local, personalized reasoning tasks, such as recommendation, classification, segmentation, or object detection.
[0049] In a fourth aspect, the present application provides a federated learning device, which is applied to a federated learning system. The federated learning system includes a server, multiple edge nodes, and multiple clients, where the multiple clients access the multiple edge nodes. The federated learning device includes the server, and the device includes:
[0050] a transceiver module configured to receive edge node model information respectively transmitted by a plurality of edge nodes, the edge node model information comprising information obtained by aggregating client model information uploaded by clients in at least one edge group by each edge node among the plurality of edge nodes, the at least one edge group being obtained by grouping at least one terminal according to data uploaded by at least one terminal connected to the plurality of edge nodes;
[0051] Aggregation module, used to aggregate the information of edge node models to obtain the global model;
[0052] The transceiver module is also used to aggregate the new global model to each edge node, so that each edge node can send the new global model to each client, so that each client can update the locally saved client model based on the new global model, and use the locally saved training data to train the updated client model, obtain the trained model and upload it to the edge node again for the next iteration.
[0053] In a possible implementation, the transceiver module is further configured to send global model information to each edge node.
[0054] In a fifth aspect, embodiments of the present application provide a federated learning device, comprising: a processor and a memory, wherein the processor and the memory are interconnected via a circuit, and the processor invokes program code in the memory to execute the processing-related functions of the federated learning method described in any one of the first aspects above. Optionally, the federated learning device may be a chip.
[0055] In a sixth aspect, embodiments of the present application provide a federated learning device, comprising: a processor and a memory, wherein the processor and the memory are interconnected via a circuit, and the processor invokes program code in the memory to execute the processing-related functions of the first method described in any one of the second aspects above. Optionally, the federated learning device may be a chip.
[0056] In the seventh aspect, an embodiment of the present application provides a federated learning system, comprising multiple servers, multiple edge nodes and at least one client, wherein the multiple clients access multiple edge nodes; any one of the multiple servers is used to implement the steps of the federated learning method shown in any one of the above-mentioned second aspects; the multiple edge nodes are used to implement the steps of the federated learning method shown in any one of the above-mentioned first aspects; and at least one client is used to use locally stored data for model training.
[0057] In the eighth aspect, an embodiment of the present application provides a digital processing chip or chip, the chip includes a processing unit and a communication interface, the processing unit obtains program instructions through the communication interface, the program instructions are executed by the processing unit, and the processing unit is used to perform processing-related functions in any optional implementation of the first or second aspect above.
[0058] In a ninth aspect, an embodiment of the present application provides a computer-readable storage medium comprising instructions, which, when executed on a computer, enables the computer to execute a method in any optional implementation of the first or second aspect above.
[0059] In the tenth aspect, an embodiment of the present application provides a computer program product comprising a computer program / instructions, which, when executed by a processor, enables the processor to execute the method in any optional implementation of the first aspect or the second aspect above. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] FIG1 is a schematic diagram of the architecture of a federated learning system provided in this application;
[0061] FIG2 is a schematic diagram of the architecture of another federated learning system provided in this application;
[0062] FIG3 is a schematic diagram of the architecture of another federated learning system provided in this application;
[0063] FIG4 is a flow chart of a federated learning method provided in this application;
[0064] FIG5 is a flow chart of another federated learning method provided in this application;
[0065] FIG6 is a flow chart of another federated learning method provided in this application;
[0066] FIG7 is a flow chart of another federated learning method provided in this application;
[0067] FIG8 is a flow chart of another federated learning method provided in this application;
[0068] FIG9 is a flow chart of another federated learning method provided in this application;
[0069] FIG10 is a flow chart of another federated learning method provided in this application;
[0070] FIG11 is a flow chart of another federated learning method provided in this application;
[0071] FIG12 is a flow chart of another federated learning method provided in this application;
[0072] FIG13 is a schematic diagram of the structure of a federated learning device provided by this application;
[0073] FIG14 is a schematic diagram of the structure of another federated learning device provided in this application;
[0074] FIG15 is a schematic diagram of the structure of another federated learning device provided in this application;
[0075] FIG16 is a schematic diagram of the structure of another federated learning device provided in this application. DETAILED DESCRIPTION
[0076] The following will describe the technical solutions in the embodiments of this application in conjunction with the drawings in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of this application.
[0077] The examples of this application are mainly used to train the machine learning models used in various application federated learning scenarios. The trained machine learning models can be applied to the above-mentioned various application fields to achieve classification, regression or other functions. The processing objects of the trained machine learning models can be image samples, discrete data samples, text samples or voice samples, etc., which are not listed here in detail. The machine learning model can be specifically expressed as a neural network, a linear model or other types of machine learning models, etc. Correspondingly, the multiple modules that make up the machine learning model can be specifically expressed as a neural network module, an existing model module or a module that constitutes other types of machine learning models, etc., which are not listed in detail here. In the subsequent embodiments, only the machine learning model expressed as a neural network is used as an example for explanation. When the machine learning model is expressed as other types other than a neural network, it can be understood by analogy, and will not be repeated in the embodiments of this application.
[0078] The embodiments of this application can be applied to the field of federated learning, enabling the collaborative training of neural networks through client and server collaboration. Therefore, they involve a large number of neural network-related applications, such as neural networks trained by clients during federated learning. To better understand the solutions of the embodiments of this application, the following first introduces the relevant terms and concepts of neural networks that may be involved in the embodiments of this application.
[0079] (1) Neural Network
[0080] A neural network can be composed of neural units. A neural unit can refer to an operation unit with xs and intercept 1 as input. The output of the operation unit can be shown as formula (1-1):
[0081] Where, s = 1, 2, ... n, n is a natural number greater than 1, W s is x s The weight of the neural unit, b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce nonlinear characteristics into the neural network to convert the input signal in the neural unit into the output signal. The output signal of the activation function can be used as the input of the next convolutional layer, and the activation function can be a sigmoid function. A neural network is a network formed by connecting multiple single neural units mentioned above, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be an area composed of several neural units.
[0082] (2) Loss function
[0083] In the process of training a neural network, because we hope that the output of the neural network is as close as possible to the value we really want to predict, we can compare the predicted value of the current network with the target value we really want, and then update the weight vector of each layer of the neural network according to the difference between the two (of course, there is usually an initialization process before the first update, that is, pre-configuring parameters for each layer in the deep neural network). For example, if the predicted value of the network is high, adjust the weight vector to make it predict lower, and continue to adjust until the deep neural network can predict the target value we really want or a value very close to the target value we really want. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function (loss function) or objective function (objective function), which are important equations used to measure the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference, so the training of the deep neural network becomes a process of minimizing this loss as much as possible. The loss function can usually include loss functions such as mean square error, cross entropy, logarithm, exponential, etc. For example, the mean square error can be used as a loss function, defined as The specific loss function can be selected according to the actual application scenario.
[0084] (3) Backpropagation algorithm
[0085] Neural networks can use the back propagation (BP) algorithm to correct the size of the parameters in the initial neural network model during training, reducing the reconstruction error loss of the neural network model. Specifically, the forward propagation of the input signal to the output generates error loss. This error loss information is then backpropagated to update the parameters in the initial neural network model, thereby converging the error loss. The BP algorithm is a backward propagation movement dominated by error loss, aiming to obtain the optimal parameters of the neural network model, such as the weight matrix.
[0086] In this application, when the client performs model training, the global model can be trained through a loss function or a BP algorithm to obtain a trained global model.
[0087] (4) Gradient
[0088] That is, the derivative vector of the loss function with respect to the parameters. It can be used to measure the change in the parameters of the model.
[0089] (5) Non-independent-identically-distributed (Non-IID)
[0090] In the context of federated learning, this refers to the different data distributions on each end-user. "Independent and identically distributed" means that the data distributions on each end-user are independent and consistent.
[0091] (6) Federated learning (FL)
[0092] A distributed machine learning algorithm uses multiple clients (such as mobile devices or edge servers) and servers to collaboratively train models and update algorithms without leaving the server's data domain, ultimately producing a trained global model. This means that during the machine learning process, each participant can leverage the other's data for joint modeling. Without sharing data resources, the parties can jointly train data and build a shared machine learning model without leaving the server's data domain.
[0093] Federated learning is a method for training neural network models without requiring data to leave user devices. Typically, at the beginning of training, a central server sends an initial model to each edge user, who then sends the model to each end-user. Subsequently, each end-user iterates the model for several epochs using local data and then feeds back model changes (such as gradients or model weights) to the edge server. The edge server takes a weighted average of the feedback gradients, uses the resulting average gradient to update the initial model, and distributes the updated model to each end-user, restarting the next small iteration. After several rounds of end-edge synchronization, the edge server feeds back changes (such as gradients or model weights) to the central server. The central server takes a weighted average of the feedback gradients, uses the resulting average gradient to update the initial model, and distributes the updated model to each edge user. The edge users then distribute the updated model to the end-user, restarting the next large iteration.
[0094] However, when user data is not independent and identically distributed (Non-IID), the gradient directions of each user's edge nodes after iteration vary greatly, resulting in the global model generated by the central server performing poorly on each edge node or even failing to converge. In real life, the geographical distance between the end nodes physically connected to the edge nodes is relatively small, and the local data collected is usually distributed similarly.
[0095] Some existing federated learning solutions can train two models simultaneously: a global model and a local model. After performing several epochs of local iterations, the end-user uploads the gradient updates obtained from the global model to the server for averaging to obtain a new global model. The end-user then downloads the global model from the server, performs a weighted sum of the global model and the local model to obtain a new local model, and prepares for the next round of local iterations. The trained local model is ultimately used as the personalized model. However, since all clients participate in the training of the global model, the personalized model is easily affected by the global model under non-IID or malicious attacks, resulting in poor performance.
[0096] In some existing solutions, the server clusters nodes with similarly distributed data and trains a personalized model for each group. However, all nodes in a group have identical model parameter values, resulting in suboptimal personalized models for each node. Furthermore, the high synchronization frequency causes the model on a node to diverge slightly before being averaged by the models of other nodes in the group. Therefore, once a node is grouped, it is difficult to assign it to another group.
[0097] Therefore, this application provides a federated learning method that can group clients through edge nodes to achieve better personalized model learning.
[0098] First, the federated learning system provided by this application is introduced.
[0099] Refer to Figure 1, which is a schematic diagram of the architecture of a federated learning system provided by the present application. The system (or cluster for short) may include multiple servers, and the multiple servers can establish connections with each other, that is, the servers can also communicate with each other. Each server can be connected to one or more clients through an intermediate device (also called an edge node or edge server, etc.), or each edge node can access one or more clients. The intermediate device can be a routing device between the server and the client, such as a base station, a router, a switch, etc. The client accesses the corresponding edge node, and the client can be deployed in various devices, such as a mobile terminal or a server, etc., such as client 1, client 2, ..., client N-1, and client N shown in Figure 1.
[0100] Specifically, servers or servers, edge nodes and clients can interact through a communication network of any communication mechanism / communication standard, and the communication network can be a wide area network, a local area network, a point-to-point connection, or any combination thereof. Specifically, the communication network may include a wireless network, a wired network, or a combination of a wireless network and a wired network. The wireless network includes but is not limited to: a fifth-generation mobile communication technology (5th-Generation, 5G) system, a long term evolution (long term evolution, LTE) system, a global system for mobile communication (global system for mobile communication, GSM) or a code division multiple access (code division multiple access, CDMA) network, a wideband code division multiple access (wideband code division multiple access, WCDMA) network, wireless fidelity (wireless fidelity, WiFi), Bluetooth (bluetooth), Zigbee protocol (Zigbee), radio frequency identification technology (radio frequency identification, RFID), long range (Lora) wireless communication, near field communication (NFC) any one or more combinations. The wired network may include a fiber optic communication network or a network composed of a coaxial cable, etc.
[0101] Accordingly, the edge node may include a device that is connected to multiple clients, such as the next generation node basestation (gNB) in the 5G system, the evolved node B (eNB) in the LTE system, the radio network controller (RNC), the node B (NB), the base station controller (BSC), the base transceiver station (BTS), the home evolved nodeB (or home node B, HNB), the base band unit (BBU), the transmission and receiving point (TRP), the transmitting point (TP), the small base station device (pico), the mobile switching center, or the network equipment in the future network. It is understandable that the present application does not limit the specific type of access network equipment. In systems using different wireless access technologies, the names of devices with access network equipment functions may be different.
[0102] Generally, the client can be deployed in various servers or terminals. The client mentioned below may also refer to a server or terminal on which the client software program is deployed. The terminal may include a mobile terminal or a fixedly installed terminal, etc. For example, the terminal may specifically include a mobile phone, a tablet, a personal computer (PC), a smart bracelet, a speaker, a television, a smart watch or other terminals.
[0103] Specifically, the federated learning system to which the federated learning method provided in this application can be applied may include a variety of topological relationships. For example, the federated learning system may include a two-layer or more-layer architecture. Some possible architectures are exemplarily introduced below.
[0104] 1. Three-tier architecture
[0105] As shown in Figure 2, this application provides a structural diagram of a federated learning system.
[0106] Among them, the federated learning system includes one or more cloud servers, one or more edge servers and one or more clients, forming a three-layer architecture of cloud server-edge server-client.
[0107] In this system, one or more edge servers are connected to the cloud server, and one or more clients are connected to the edge servers.
[0108] During federated learning, the cloud server sends its locally stored global model to the edge server, which then sends it to the connected client. The client then trains the received global model using locally stored training samples and feeds the trained global model back to the edge server. The edge server then updates the locally stored global model based on the trained global model and feeds the updated global model back to the cloud server, completing federated learning.
[0109] Two or three-layer architecture
[0110] As shown in Figure 3, this application provides a structural diagram of another federated learning system.
[0111] The federated learning system includes a three-tier architecture, with one tier comprising one or more cloud servers. Multiple edge servers form two or more tiers, such as one tier consisting of one or more upstream edge servers, each connected to one or more downstream edge servers. The final tier of edge servers, each connected to one or more clients, forms a tier.
[0112] During federated learning, the upstream cloud server sends the latest locally stored global model to the next-level edge server. The edge server then sends the global model down layer by layer until it reaches the client. After receiving the global model from the edge server, the client trains the received global model using locally stored training samples and feeds the trained global model back to the edge server in the previous layer. The edge server in the previous layer then updates the locally stored global model based on the received trained global model and uploads the updated global model to the edge server in the next layer. This continues until the second-level edge server uploads the updated global model to the cloud server. The cloud server then updates the local global model based on the received global model to obtain the final global model, completing federated learning.
[0113] The above describes the federated learning system provided in this application. Now, in conjunction with the federated learning system, the flow of the method provided in this application will be introduced.
[0114] Referring to FIG4 , a flowchart of a federated learning method provided in this application is described below.
[0115] First, the architecture of the federated learning system provided in this application can be seen in Figures 2 and 3 above. Here, for example, a three-tier architecture is used for this purpose. Federated learning typically involves multiple rounds of iterations, and this application embodiment uses one of these iterations as an example for this purpose.
[0116] For ease of distinction, the model obtained by server aggregation is called the global model, the model obtained by edge node aggregation is called the edge node model, and the model maintained and trained by the client is called the client model. They will not be repeated below.
[0117] 401. The server sends the global model to the edge node.
[0118] First, the server can send a global model to the edge node. This global model is the global model aggregated in the previous iteration or the initial model. If this iteration is the first, the global model is the initialized model to be trained. If this iteration is not the first, the global model can be updated based on the model uploaded by the edge node in the previous iteration.
[0119] In the federated learning system provided in this application, the number of edge nodes can be one or more. If there are multiple edge nodes, the cloud server can send the global model to each of the edge nodes. Furthermore, the cloud service can proactively send the global model to the edge nodes, or send the global model to each edge node at the request of the edge node. The specific adjustment can be made based on the actual application scenario and is not limited by this application.
[0120] Specifically, the server can send the structural parameters of the global model (such as the width, depth or convolution kernel size of the neural network) or initial weight parameters to the downstream device.
[0121] Optionally, the server can also send training configuration parameters to downstream devices, such as learning rate, number of epochs, or categories in security algorithms, so that the client that ultimately performs training can use the training configuration parameters to train the global model.
[0122] In the embodiments of the present application, the models mentioned, such as global models, edge node models, client models or local models, can specifically include convolutional neural networks (CNN), deep convolutional neural networks (DCNN), recurrent neural networks (RNN) and other neural networks. The model to be learned can be determined according to the actual application scenario, and the present application does not limit this. The aforementioned model can be specifically used to perform at least one of the recommendation task, classification task, segmentation task or target detection task. The specific application scenario of the model can be adjusted according to the actual application scenario, and the present application does not limit this.
[0123] 402. The edge node sends the model to the client.
[0124] The edge node updates the local edge node model of the edge node based on the received global model. Specifically, the locally stored edge node model can be directly replaced with the global model to obtain a new edge node local model; or, the received global model and the locally stored edge node model obtained in the previous iteration process can be aggregated, such as averaging or weighted fusion, to obtain a new local model.
[0125] After obtaining the global model, the edge node can directly send the received global model to the client, or send the updated local edge node model to the client.
[0126] 403. The client uses local data for training and uploads information about the trained client model to the edge node.
[0127] After receiving the model sent by the edge node, the client can update the local client model based on the received model, such as fusing the locally stored client model with the global model, or replacing the locally stored client model with the global model. The new local client model is trained using the locally stored training samples to obtain the trained client model.
[0128] The trained client model information is uploaded to the edge node. Specifically, the client model weight parameters, gradient parameters, etc. can be uploaded.
[0129] Accordingly, the local training data of each client can include data related to the inference tasks that the client model needs to perform. For example, if the client model is used to perform a recommendation task, the training data can include data related to the recommendation scenario collected locally by the client, such as data recommended to the user and the client's feedback on the recommended data; if the client model is used to perform a classification task, the training data can include classified data collected by the client; if the client model is used to perform a segmentation task, the training data can include segmented data collected by the client; if the client model is used to perform a target detection task, the training data can include target-labeled data collected by the client, etc. The specific selection can be based on the actual application scenario.
[0130] 404. The edge node groups the connected clients based on the client model information uploaded by the clients.
[0131] After receiving the trained model information uploaded by one or more connected clients, the edge node can group the one or more clients based on the trained model information uploaded by the one or more clients, and divide the one or more clients into at least one side group.
[0132] Furthermore, a federated learning system can include multiple edge nodes, which can transfer data between them. Therefore, when grouping clients, each edge node can also group clients connected to other edge nodes. This allows for the grouping of clients within the federated learning system, allowing clients with similar data distributions to be grouped together across edge nodes. This enables more optimized personalized model training, thereby improving personalized model performance.
[0133] Optionally, the edge node may first group the clients based on the client model information uploaded by the clients to obtain at least one end-side group, and then further determine the edge-side group corresponding to the edge node based on the at least one end-side group.
[0134] Optionally, when grouping clients, clustering can be performed based on the client model information uploaded by the clients, and then the clients can be grouped according to the clustering result to obtain at least one end-side group. For example, the client model information uploaded by at least one client can be clustered, and at least one client can be grouped according to the clustering result, such as grouping clients corresponding to model information with data in the same category into one group to obtain at least one end-side group. Therefore, in the embodiment of the present application, when grouping clients, the terminals can be grouped in combination with the data structure features of the client model information uploaded by the clients, so that clients with similar data structure features are grouped together.
[0135] Optionally, the number of each end-side grouping can be a preset number or a number determined according to the evaluation results. For example, after obtaining the clustering results, at least one client can be grouped based on the clustering results and the number of groups to obtain at least one end-side grouping. The at least one end-side grouping is evaluated to obtain an evaluation value for each end-side grouping. If the evaluation value does not meet the preset evaluation conditions, such as if the evaluation value is less than the preset value, the number of groups can be adjusted, and the at least one client can be regrouped according to the adjusted number of groups until an end-side grouping whose evaluation value meets the preset evaluation conditions is obtained. In the implementation manner of the present application, the grouping results can be evaluated to obtain a grouping result that meets the conditions.
[0136] In a possible scenario, the edge node can also obtain the similarity between the locally stored model information and the information of the client model uploaded by each client in at least one client, and the locally stored model information is the information of the edge node model obtained in the last iterative process; the first edge node determines the edge grouping based on at least one end-side grouping based on the similarity between the locally stored edge node model information and the client model information uploaded by each client, such as screening out clients in the end-side group whose model similarity with the edge node is less than a preset similarity, so as to avoid the influence of data with large differences on the edge node model data of the edge node. In an embodiment of the present application, the edge node can select clients in the edge grouping based on the similarity between its own stored edge node model and the client model uploaded by the client, so that clients with higher model similarity can be selected as the edge group corresponding to the edge node.
[0137] Optionally, the way in which the edge node groups the clients to obtain the edge grouping may also include: the edge node obtains the similarity between the locally stored model information and the information of the model uploaded by at least one client, and the locally stored model information may be the information of the edge node model obtained in the last iteration process; based on the similarity between the locally stored model information and the information of the model uploaded by at least one client, at least one client is screened to obtain the edge grouping, such as selecting the top N (positive integer) clients with the highest similarity as the clients in the edge grouping. Therefore, in the embodiment of the present application, the edge node can directly filter out the clients of the edge grouping from one or more clients, so that the model that is more similar to the local model of the edge node can be determined more efficiently.
[0138] In addition, in one possible scenario, the edge node can also obtain the communication resources corresponding to each client, that is, the resources available when each client communicates with the edge node, such as network delay or computing power. The edge node can combine the communication resources corresponding to each client to determine the clients in the edge group, so as to obtain clients with higher communication efficiency with the edge node as the edge group, and obtain the edge group that is adapted to the communication resources of the edge node, thereby improving the learning efficiency of federated learning.
[0139] In addition, in a possible scenario, one of the edge nodes (such as the first edge node for ease of distinction) can identify the similarity between the information of the client model of the client in each end-side group and the edge node model of other edge nodes (such as the second edge node for ease of distinction) or the client model of the client connected to the second edge node. If the similarity is higher than a preset value, the information of the client model uploaded by the client with a similarity higher than the preset value can be sent to the second edge node. Therefore, the edge nodes can communicate with each other, and the client model information uploaded by the client with similar data structure features can be sent to other edge nodes with similar data structure features, so that the model information between the edge nodes and the clients with similar data structure features can be aggregated.
[0140] In addition, in some possible scenarios, when model information needs to be transmitted between edge nodes, the model information to be transmitted can be directly transmitted through the connection established between the edge nodes, or the model information to be transmitted between the edge nodes can be transmitted through the backbone network in the communication network. Taking the first edge node and the second edge node as an example, the first edge node sends the information of the client model uploaded by the client in the end-side group corresponding to the second edge node to the second edge node through a direct connection established between the first edge node and the second edge node (such as short-distance communication or wired connection, etc.); or, the first edge node sends the information of the client model uploaded by the client in the end-side group corresponding to the second edge node to the second edge node through the backbone network of the communication network, and the communication network is the communication network to which the first edge node and the second edge node access.
[0141] 405. The edge node aggregates the client model information uploaded by the clients in the edge group to obtain information of the edge node model.
[0142] After the edge node groups the clients into edge groups, it can aggregate the client model information uploaded by the clients in the edge groups to obtain the edge node model information.
[0143] The number of edge node models can be one or more. For example, when the edge grouping includes multiple types of client groupings, i.e., end-side groupings, aggregation can be performed based on each type of end-side grouping to obtain multiple edge node models. When the edge grouping includes only one type of client grouping, i.e., end-side groupings, aggregation can be performed to obtain a single edge node model.
[0144] 406. The edge node uploads information of the edge node model to the server.
[0145] After the edge node obtains the information of the edge node model, it can upload the information of the edge node model to the server, which can specifically be the weight parameters or gradient parameters of the transmission model.
[0146] 407. The server aggregates the received information of the edge node models to obtain an aggregated global model.
[0147] After receiving information about one or more edge node models, the server can aggregate the received edge node models again to obtain an aggregated global model.
[0148] 408. The server sends the aggregated global model to the edge nodes.
[0149] 409. The edge node sends the aggregated complete situation model to the client.
[0150] For details of steps 408 to 409, please refer to the description of steps 401 to 402, which will not be repeated here.
[0151] After the server obtains the aggregated global model, it can proceed to the next round of iteration, that is, repeat step 401. Similar steps will not be repeated here.
[0152] In the implementation of the present application, the edge node can group one or more clients connected to it, and can group them based on the data structure characteristics of each client. Therefore, when performing model aggregation, the client model information uploaded by clients in the same group can be aggregated, thereby realizing personalized model training within each group, so that the final model performance is better. This is equivalent to achieving personalized training for the global model obtained by the method provided by the present application, which maximizes the retention of the personalized characteristics of each client. The aggregated global model can be deployed as a new client model in each client and can be used to perform local personalized reasoning tasks for each client, such as recommendation tasks, classification tasks, segmentation tasks, or target detection tasks adapted to each client, thereby realizing personalized client reasoning tasks.
[0153] The above describes the process of the federated learning method provided by this application. The following describes the process of the method provided by this application in combination with more specific application scenarios.
[0154] In conjunction with the process shown in FIG4 , the federated learning method provided in this application can be divided into multiple steps. Each step is introduced below.
[0155] 501. The server sends the global model.
[0156] As shown in Figure 5, if the current iteration is in the initial stage, the server can send down the initial global model and local models, where the local models are the initial models of each terminal. Specifically, the server can send down the structural parameters or weight parameters of the global model.
[0157] If the current iteration is not the first, the server can send only the global model or a new local model. The specific adjustment can be made according to the actual application scenario. The local model is the model that the client needs to update locally, or the structure in the global model that the client needs to update.
[0158] Typically, in a federated learning system, the client can train two models simultaneously: a global model and a local model. After the end-side user performs several epochs of local iteration, the gradient update amount obtained from the global model is uploaded to the server for averaging to obtain a new global model. In the subsequent iteration process, the end-side user receives the new global model, performs a weighted summation of the new global model and the local model to obtain a new local model, and prepares for the next round of local iteration. Finally, the trained local model is used as the local personalized model. The embodiment of this application is illustratively introduced by taking one of the iterations as an example, and will not be repeated below.
[0159] Furthermore, when data is transmitted between nodes in a federated learning system, such as between servers and edge nodes, between edge nodes, or between edge nodes and clients, to reduce communication overhead, the data to be transmitted can be quantized. This reduces the number of bits occupied by the data being transmitted, thereby reducing communication overhead. After receiving the quantized data, the receiver can dequantize it to restore the full-precision model information.
[0160] 502. The edge node sends the global model, and the client uploads the trained client model.
[0161] As shown in Figure 6, after receiving the global model sent by the server, the edge node can directly forward the global model to each client; or it can fuse the global model with the edge node model of the edge node and then send the fused model to each client.
[0162] After receiving the global model or edge node model sent by the edge node, the client can perform a weighted fusion of the received global model or edge node model with the local client model to obtain a new client model, train the new client model to obtain a trained client model, and upload the new trained client model to the edge node. Alternatively, the client can directly use the received global model or edge node model as the new client model, train the new client model to obtain a trained client model, and upload the new trained client model to the edge node.
[0163] When updating a model, each client segment can update all or part of the client model's structural parameters, thereby obtaining a personalized client model adapted to each client's local needs and uploading it to the edge node. When each client updates part of the client model's structure, the updated part of the structure can be determined by the server or the edge node.
[0164] 503. The edge nodes are grouped based on the data structure characteristics of the client.
[0165] As shown in Figure 7, after receiving the trained client models uploaded by connected clients, edge nodes can group the clients based on the client model information uploaded by each client. When grouping, edge nodes can communicate with each other and transmit information about the client models uploaded by connected clients, or features of the client models uploaded by each client. This allows for cross-edge-node client grouping, allowing more clients with similar data distributions to be grouped together. This improves the training effect of personalized models and prevents the adverse effects of client models with significantly different data on personalized models.
[0166] In the embodiments of the present application, the steps performed by one of the edge nodes are taken as an example for exemplary introduction. The edge node can be any node in the federated learning system, and will not be repeated below.
[0167] Optionally, clients with similar data structure characteristics can be grouped together. For example, the client model information uploaded by each client can be clustered, and clients corresponding to model information in the same category can be grouped together.
[0168] In addition, when determining the edge grouping, the clients in the edge grouping can also be determined by combining the similarity between the client model information uploaded by the clients in each group and the local edge node model of the edge node. For example, the cosine similarity between the client model information uploaded by each client and the edge node model of the edge node can be calculated, and clients with similarity higher than a preset similarity can be divided into the edge grouping, that is, the client is included in the edge group of the edge node. Therefore, when the edge node performs model aggregation, its local personalized model characteristics can be retained as much as possible, achieving personalized model training with better performance.
[0169] For example, when an edge node selects an end node to establish a logical connection, it uses cosine similarity to determine the difference between the end node data distribution and the model parameters stored locally on the edge node, and selects K (positive integer) clients with the smallest difference as the edge group, fully accessing the end node within limited bandwidth resources to serve personalized model training.
[0170] Optionally, when grouping edge nodes, the communication resources or computing power within the group can also be considered. For example, edge nodes can evaluate the comprehensive metrics of an end node by analyzing communication-related information such as the latency and bandwidth when the client uploads, as well as computing power information. This allows them to consider whether to connect to that end node for synchronization within the limited bandwidth and round time. This can, for example, prevent clients with high transmission latency or low computing power from accessing the network, improving the efficiency of federated learning.
[0171] When edge nodes are grouped, they can be grouped according to a preset number, or a specific number of groups can be selected according to actual application scenarios.
[0172] For example, the silhouette score can be used as an indicator to evaluate the clustering effect, which measures the similarity between data points within the same cluster and the dissimilarity between data points in different clusters. The steps of calculating the number of groups k using the silhouette score may include: selecting a range of k values, where k is the number of groups; performing kmeans clustering for each k value; calculating the silhouette score for each k; and selecting the k value corresponding to the largest silhouette score as the number of groups. Therefore, it can assist in obtaining the best grouping for the client when the distribution of each edge node is unknown or when it is difficult to ensure the security of all edge nodes. In this way, it can automatically help the user find the most suitable number of groups when the user is unsure of the number of groups, so as to find the relatively optimal grouping and improve the performance of the personalized model.
[0173] Of course, when grouping clients, in addition to using the silhouette score to evaluate each number of groups, other evaluation indicators can also be used, such as the average distance within the cluster or the average distance between clusters. The specific evaluation indicators can be selected according to the actual application scenario, and they will not be described here one by one.
[0174] Furthermore, in some scenarios, edge nodes can transmit data to each other, and client model information uploaded by clients connected to each edge node can be transmitted between edge nodes. Therefore, when grouping edge nodes, you can also group them across edge nodes. For example, you can group clients connected to the first edge node into the edge group of the second edge node. This allows clients with closer data distribution to be grouped together across edge nodes, which is more conducive to personalized model training.
[0175] 504. The edge node aggregates the client models uploaded by the clients in each group.
[0176] As shown in FIG8 , after receiving and grouping the client model information uploaded by the connected clients, the edge node can aggregate the client model information uploaded by the clients in each group to obtain one or more edge node models.
[0177] Usually, when performing model aggregation, you can refer to the weighted formula: λ.h1+(1-λ).h c , where λ is the weight of the edge node model parameter h1 local to the edge node, h c That is, the client model parameters uploaded by the client.
[0178] Among them, when multiple edge node models are obtained, if the clients are divided into multiple groups, an edge node model is obtained for each group. At this time, the multiple edge node models obtained can be aggregated again to obtain the final edge node model. Of course, multiple edge node models can also be directly uploaded to the server.
[0179] When the clients under the edge node are divided into only one group, only the client model parameters uploaded by the clients in this group can be aggregated.
[0180] Therefore, by grouping, we can adapt to scenarios with various data distribution structures and achieve more personalized model training.
[0181] Typically, edge nodes can upload only the parameters of one or more layers in the model after each grouping, reducing communication overhead while avoiding information leakage caused by direct data distribution. For example, an edge node can upload only the parameters of one or more network layers in the locally updated global model whose sensitivity exceeds a preset sensitivity or whose change exceeds a preset change value.
[0182] Furthermore, the data structure characteristics of clients connected to different edge nodes may be similar, so clients connected to different edge nodes can be grouped together. In this scenario, client model information uploaded by clients in different edge node groups can also be transmitted between edge nodes.
[0183] Optionally, when client model information uploaded by multiple clients needs to be transmitted between edge nodes, the sending side node can aggregate the client model information uploaded by the multiple clients and transmit it to the receiving side node, or it can directly transmit the client model information uploaded by multiple clients. The specific transmission method can be selected according to the actual application scenario, and this application does not limit this.
[0184] For example, the information of the transmission model between nodes in different scenarios can be shown in Figure 9 or Figure 10.
[0185] That is, optionally, step 5041 may be performed: client models uploaded by clients in cross-node groups are transmitted between edge nodes.
[0186] As shown in Figure 9, if a direct communication connection is established between edge nodes, the edge nodes can directly transmit the model information that needs to be transmitted. The direct communication connection can be established between edge nodes through short-range communication (such as Bluetooth or WiFi), or directly through a wired connection.
[0187] As shown in Figure 10, edge nodes can also transmit data through the backbone network. If edge nodes cannot communicate directly, the model information that needs to be transmitted between edge nodes can be transmitted through the backbone network. This backbone network can be the backbone network of the communication network that the edge nodes access.
[0188] In this scenario, during the next round of iteration, the second edge node can send a new global model to the client connected across the nodes, so that the client can use the new global model adapted to the local data distribution for downstream tasks.
[0189] 505. The edge node uploads the edge node model.
[0190] As shown in Figure 11, after the edge node aggregates the client models uploaded by the clients in each group to obtain one or more edge node models, it can upload information of the one or more edge node models to the server, such as the weight parameters or gradients of the edge node models.
[0191] In the implementation of the present application, clients training personalized models are grouped based on their differences, effectively avoiding interference from nodes with large differences or potential malicious attacks on personalized learning. Taking into account the impact of personalized model differences on model training, edge nodes with large differences are clustered and grouped according to their model parameters to eliminate interference from model training caused by edge nodes with large differences. The greater the difference in data distribution between clients, the more obvious the advantage of the grouped model; at the same time, the selection of client groups takes into account both bandwidth resources and distribution differences, so that edge nodes can select high-quality end nodes as much as possible, thereby improving the performance of personalized models.
[0192] 506. The server aggregates to obtain a new global model.
[0193] As shown in FIG12 , after the server receives one or more edge node models, it can update the local global model based on the received one or more edge node models to obtain an updated global model.
[0194] After obtaining the updated global model, the server can send the new global model to each edge node to enter the next round of iteration, that is, the server can continue to send the new global model to the edge node, that is, repeat step 501 and subsequent steps, so that each client can keep the global model updated synchronously and use the latest model for downstream tasks.
[0195] In the implementation of the present application, the server can send the global model to the edge node, and the edge node will group the client model information uploaded by each client based on the data structure characteristics, so that the client models uploaded by clients with closer data distribution can be aggregated later, and the personalized characteristics of clients with closer distribution can be retained as much as possible, so as to achieve a personalized model with better performance. This is equivalent to restricting the logical connection between clients and edge nodes that have not established a physical connection, ensuring that the distribution differences between the end nodes selected by the edge node are as small as possible on limited bandwidth resources, thereby improving the performance of the personalized model. In addition, the edge node can aggregate the client models uploaded by the connected clients, and only needs to upload the aggregated model to the server, without uploading the client models uploaded by all clients to the server. Therefore, the amount of data that needs to be transmitted to the server is greatly reduced, and the communication overhead and server load are greatly reduced. After the clients are grouped, the global model in the group is only synchronized within the group, further reducing the communication pressure on the cloud server.
[0196] The above describes the method flow provided by this application. The following describes the device for executing the above method.
[0197] Refer to Figure 13, which is a schematic diagram of the structure of a federated learning device provided by this application. The device can be used to perform the steps performed by the edge nodes in Figures 4 to 12 above. The device includes the following:
[0198] The transceiver module 1301 is configured to obtain data uploaded by at least one client, where the at least one client includes a client connected to multiple edge nodes in the federated learning system, and the federated learning device is any one of the multiple edge nodes;
[0199] A grouping module 1302 is configured to group at least one client according to information of a client model uploaded by at least one client, and obtain side groups corresponding to the federated learning device;
[0200] Aggregation module 1303, configured to aggregate information of client models uploaded by clients in the edge group to obtain information of edge node models;
[0201] The transceiver module 1301 is further configured to send information about edge node models to the server, and the server is configured to aggregate the edge node model information received from multiple edge nodes to obtain a global model.
[0202] The transceiver module 1301 is further configured to receive the aggregated global model sent by the server;
[0203] The transceiver module 1301 is also used to send the aggregated global model to each client connected to the first edge node, so that each client uses the aggregated global model to update the client model stored locally by each client, and uses the locally stored data to train the updated client model to obtain the trained client model. The information of the trained client model is used to upload it to the first edge node again.
[0204] In a possible implementation, the transceiver module 1301 is further configured to obtain information of a client model uploaded by at least one client;
[0205] A grouping module 1302 is configured to group at least one client according to information of a client model uploaded by at least one client to obtain at least one terminal-side group;
[0206] The grouping module 1302 is configured to determine an edge-side grouping of the federated learning device according to at least one end-side grouping.
[0207] In a possible implementation, the grouping module 1302 is specifically configured to cluster the information of the client model uploaded by at least one client, and group the at least one client according to the clustering result to obtain at least one terminal-side group.
[0208] In one possible implementation, the grouping module 1302 is specifically used to: cluster the information of the client model uploaded by at least one client to obtain a clustering result; group at least one client according to the clustering result and the number of groups to obtain at least one end-side group; evaluate at least one end-side group to obtain an evaluation value of at least one end-side group; if the evaluation value does not meet the preset evaluation conditions, adjust the number of groups, and re-group at least one client according to the adjusted number of groups.
[0209] In one possible implementation, the grouping module 1302 is specifically used to: obtain the similarity between the locally stored model information and the information of the client model uploaded by each client in at least one client, where the locally stored model information includes the information of the edge node model obtained in the last iteration process; based on the similarity between the locally stored model information and the information of the client model uploaded by each client, determine the edge-side grouping according to at least one end-side grouping.
[0210] In one possible implementation, the grouping module 1302 is specifically used to: obtain communication resources corresponding to at least one client, where the communication resources include communication resources used by at least one client to communicate with a federated learning device; and determine an edge-side group based on at least one end-side group in combination with the communication resources corresponding to the client in at least one group.
[0211] In a possible embodiment, the transceiver module 1301 is also used to send the information of the client model uploaded by the client in the end-side group corresponding to the second edge node to the second edge node if the federated learning device determines that at least one end-side group also includes an end-side group corresponding to the second edge node, which can also be referred to as the first end-side group for ease of distinction. For example, if the information of the client model uploaded by the client in one group of end-side groups has a high similarity with the edge node model of the second edge node, or the information of the client model uploaded by the client in the end-side group and the information of the client model uploaded by the client in the edge group of the second edge node belong to the same category (i.e., the clustered category).
[0212] In one possible implementation, the transceiver module 1301 is further used to: send information about the client model uploaded by the client in the end-side group corresponding to the second edge node to the second edge node through a connection established with the second edge node; or send information about the client model uploaded by the client in the end-side group corresponding to the second edge node to the second edge node through a backbone network.
[0213] In a possible implementation, the transceiver module 1301 is further configured to send information of at least one network layer in the edge node model to the server.
[0214] In one possible implementation, the grouping module 1302 is further used for: the first edge node obtains the similarity between the locally stored model information and the model information uploaded by at least one client, where the locally stored model information may be the information of the edge node model obtained during the previous iteration; based on the similarity between the locally stored model information and the model information uploaded by at least one client, screening at least one client to obtain an edge group, for example, selecting the N (positive integer) clients with the highest similarity as the clients in the edge group.
[0215] Referring to FIG14 , a schematic diagram of another federated learning apparatus provided in an embodiment of the present application is provided. The apparatus can be used to execute the steps performed by the server in FIG4 to FIG12 . The apparatus includes:
[0216] The transceiver module 1401 receives edge node model information respectively transmitted by multiple edge nodes, where the edge node model information includes information obtained by aggregating client model information uploaded by clients in at least one edge group by each edge node among the multiple edge nodes, where the at least one edge group is obtained by grouping at least one terminal according to data uploaded by at least one terminal connected to the multiple edge nodes.
[0217] Aggregation module 1402, for aggregating information of edge node models to obtain a global model;
[0218] The transceiver module 1401 is also used to aggregate the new global model to each edge node, so that each edge node can send the new global model to each client, so that each client can update the locally stored client model based on the new global model, and use the locally stored training data to train the updated client model, obtain the trained model and upload it to the edge node again for the next iteration.
[0219] In a possible implementation, the transceiver module 1401 is further configured to send global model information to each edge node.
[0220] As shown in FIG15 , a hardware structure diagram of a federated learning device 150 provided in an embodiment of the present application is shown. The federated learning device can be used to execute the steps executed by the edge nodes in FIG5 to FIG12 .
[0221] The federated learning apparatus 150 shown in FIG15 may include a processor 1501 , a memory 1502 , a communication interface 1503 , and a bus 1504 . The processor 1501 , the memory 1502 , and the communication interface 1503 may be connected via the bus 1504 .
[0222] The processor 1501 is the control center for generating the federated learning apparatus 150 and can be a general-purpose central processing unit (CPU) or other general-purpose processors, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0223] As an example, the processor 1501 may include one or more CPUs, such as CPU 0 and CPU 1 shown in FIG. 15 .
[0224] The memory 1502 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to these.
[0225] In one possible implementation, memory 1502 may exist independently of processor 1501. Memory 1502 may be connected to processor 1501 via bus 1504 and used to store data, instructions, or program code. When processor 1501 calls and executes the instructions or program code stored in memory 1502, the method provided in the embodiments of the present application can be implemented.
[0226] In another possible implementation, the memory 1502 may also be integrated with the processor 1501 .
[0227] Communication interface 1503 is used to connect federated learning apparatus 150 to other devices via a communication network, such as Ethernet, a radio access network (RAN), or a wireless local area network (WLAN). Communication interface 1503 may include a receiving unit for receiving data and a transmitting unit for sending data.
[0228] Bus 1504 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. This bus can be classified as an address bus, a data bus, a control bus, etc. For ease of illustration, FIG15 shows only one thick line, but this does not mean that there is only one bus or only one type of bus.
[0229] It should be noted that the structure shown in FIG15 does not constitute a limitation on the federated learning device 150. In addition to the components shown in FIG15, the federated learning device 150 may include more or fewer components than shown, or combine certain components, or arrange the components differently.
[0230] Optionally, the federated learning device shown in the aforementioned FIG15 is a chip.
[0231] As shown in FIG16 , it is a schematic diagram of the hardware structure of a federated learning device 160 provided in an embodiment of the present application. The federated learning device can be used to execute the steps executed by the server in FIG5 to FIG12 .
[0232] The federated learning apparatus 160 shown in FIG16 may include a processor 1601 , a memory 1602 , a communication interface 1603 , and a bus 1604 . The processor 1601 , the memory 1602 , and the communication interface 1603 may be connected via the bus 1604 .
[0233] The processor 1601 is the control center for generating the federated learning device 160 and can be a general-purpose central processing unit (CPU) or other general-purpose processors, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0234] As an example, the processor 1601 may include one or more CPUs, such as CPU 0 and CPU 1 shown in FIG. 16 .
[0235] The memory 1602 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to these.
[0236] In one possible implementation, memory 1602 may exist independently of processor 1601. Memory 1602 may be connected to processor 1601 via bus 1604 and used to store data, instructions, or program code. When processor 1601 calls and executes the instructions or program code stored in memory 1602, the method provided in the embodiments of the present application can be implemented.
[0237] In another possible implementation, the memory 1602 may also be integrated with the processor 1601 .
[0238] Communication interface 1603 is used to connect federated learning apparatus 160 to other devices via a communication network, such as Ethernet, a radio access network (RAN), or a wireless local area network (WLAN). Communication interface 1603 may include a receiving unit for receiving data and a transmitting unit for sending data.
[0239] Bus 1604 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. This bus can be classified as an address bus, a data bus, a control bus, etc. For ease of illustration, FIG16 shows only one thick line, but this does not mean that there is only one bus or only one type of bus.
[0240] It should be noted that the structure shown in FIG16 does not constitute a limitation on the federated learning device 160. In addition to the components shown in FIG16, the federated learning device 160 may include more or fewer components than shown, or combine certain components, or arrange the components differently.
[0241] Optionally, the federated learning device shown in the aforementioned FIG16 is a chip.
[0242] The present application also provides a digital processing chip. The digital processing chip integrates circuits and one or more interfaces for implementing the aforementioned processor 1501 or 1601, or the functions of processor 1501 or 1601. When the digital processing chip integrates a memory, the digital processing chip can perform the method steps of any one or more of the aforementioned embodiments. When the digital processing chip does not integrate a memory, it can be connected to an external memory via a communication interface. The digital processing chip implements the actions performed by the federated learning device in the aforementioned embodiments based on the program code stored in the external memory.
[0243] An embodiment of the present application also provides a computer program product, which, when executed on a computer, enables the computer to execute the steps of the method described in the embodiments shown in the aforementioned Figures 5 to 12.
[0244] The federated learning device provided in the embodiments of the present application may be a chip, which includes: a processing unit and a communication unit. The processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin, or a circuit. The processing unit may execute computer-executable instructions stored in the storage unit, and use the chip to perform the methods described in the embodiments shown in Figures 5 to 12. Optionally, the storage unit is a storage unit within the chip, such as a register, a cache, etc. The storage unit may also be a storage unit located outside the chip within the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0245] Specifically, the aforementioned processing unit or processor may be a central processing unit (CPU), a neural-network processing unit (NPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processing unit or processor may be used to execute the steps of the method corresponding to the aforementioned Figures 5 to 12.
[0246] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.
[0247] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general-purpose hardware, and of course can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits. However, for the present application, software program implementation is a better implementation method in most cases. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc., including a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0248] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0249] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website, a computer, a server, or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a server or a data center that includes one or more available media integrations. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0250] The terms "first," "second," "third," "fourth," and the like in the specification and claims of this application and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions, e.g., a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such process, method, product, or apparatus.
[0251] Finally, it should be noted that the above is only a specific implementation method of the present application, but the protection scope of the present application is not limited to this. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the protection scope of the present application.
Claims
1. A federated learning method, characterized in that: Applied to a federated learning system, the federated learning system includes a server, multiple edge nodes and multiple clients, the multiple edge nodes are connected to the server, the multiple clients are connected to the multiple edge nodes, the method includes multiple iterative processes, wherein one iterative process is: The first edge node obtains information of a client model uploaded by at least one client, wherein the at least one client includes a client connected to the plurality of edge nodes, and the first edge node is any one of the plurality of edge nodes; The first edge node groups the at least one client according to the information of the client model uploaded by the at least one client to obtain an edge group corresponding to the first edge node; The first edge node aggregates information of the client model uploaded by the clients in the edge group to obtain information of the edge node model; The first edge node sends the information of the edge node model to the server, and the server is used to obtain the information of the global model by aggregating the edge node model information uploaded by the multiple edge nodes; The first edge node receives the aggregated global model sent by the server; The first edge node sends the aggregated global model to each client connected to the first edge node, so that each client uses the aggregated global model to update the client model stored locally by each client, and uses the locally stored data to train the updated client model to obtain the trained client model, and the information of the trained client model is used to upload to the first edge node again.
2. The method according to claim 1, characterized in that The first edge node groups the at least one client according to information of the model uploaded by the at least one client to obtain an edge group corresponding to the first edge node, including: The first edge node groups the at least one client according to the information of the client model uploaded by the at least one client to obtain at least one terminal side group; The first edge node determines an edge side group corresponding to the first edge node according to the at least one end side group.
3. The method according to claim 2, characterized in that The first edge node groups the at least one client according to the information of the client model uploaded by the at least one client to obtain at least one end-side group corresponding to the first edge node, including: The first edge node clusters the information of the client model uploaded by the at least one client, and groups the at least one client according to the clustering result to obtain at least one end-side group.
4. The method according to claim 3, characterized in that The first edge node clusters the information of the client model uploaded by the at least one client, and groups the at least one client according to the clustering result to obtain at least one client-side group, including: The first edge node clusters the information of the client model uploaded by the at least one client to obtain a clustering result; The first edge node obtains the number of groups, where the number of groups includes the number of clients included in each end-side group; The first edge node groups the at least one client according to the clustering result and the number of groups to obtain at least one end-side group; The first edge node evaluates the at least one end-side group to obtain an evaluation value of the at least one end-side group; If the evaluation value does not meet the preset evaluation condition, the first edge node adjusts the number of groups and regroups the at least one client according to the adjusted number of groups.
5. The method according to any one of claims 2 to 4, characterized in that: The first edge node determines, according to the at least one end-side group, an edge-side group corresponding to the first edge node, including: The first edge node obtains the similarity between the locally stored model information and the information of the client model uploaded by each of the at least one client, wherein the locally stored model information is the information of the edge node model obtained in the last iteration process; The first edge node determines the edge group corresponding to the first edge node from the at least one end-side group based on the similarity between the locally stored model information and the information of the client model uploaded by each client.
6. The method according to any one of claims 2 to 5, characterized in that: The first edge node determines an edge side group of the first edge node according to the at least one end side group, further comprising: The first edge node further acquires communication resources corresponding to the at least one client, where the communication resources include communication resources used when the at least one client communicates with the first edge node; The first edge node determines the edge-side group according to the at least one end-side group in combination with the communication resources corresponding to the clients in the at least one group.
7. The method according to any one of claims 2 to 6, characterized in that: The method further comprises: If the first edge node determines that the at least one end side group also includes a first end side group corresponding to a second edge node, the first edge node sends information on a client model uploaded by a client in the end side group corresponding to the second edge node to the second edge node, and the second edge node is any one of the multiple edge nodes except the first edge node, and the information on the client model of the first end side group and the information on the client model of the edge side group of the second edge node belong to the same category, and the same category is the same clustering category after clustering the client model information.
8. The method according to claim 7, characterized in that The first edge node sends information of a client model uploaded by a client in an end-side group corresponding to the second edge node to the second edge node, including: The first edge node sends, to the second edge node, information of the client model uploaded by the client in the terminal side group corresponding to the second edge node through a direct connection established between the first edge node and the second edge node; Alternatively, the first edge node sends information of the client model uploaded by the client in the end-side group corresponding to the second edge node to the second edge node through a backbone network in a communication network, and the communication network is a network where the first edge node and the edge node are located.
9. The method according to any one of claims 1 to 8, characterized in that The first edge node sends information of the edge node model to the server, including: The first edge node sends information of at least one network layer in the edge node model to the server.
10. The method according to any one of claims 1 to 9, characterized in that The first edge node groups the at least one client according to information of the model uploaded by the at least one client to obtain an edge group corresponding to the first edge node, further comprising: The first edge node obtains the similarity between the locally stored model information and the model information uploaded by the at least one client, wherein the locally stored model information is the information of the edge node model obtained in the last iteration process; According to the similarity between the locally stored model information and the model information uploaded by the at least one client, the at least one client is screened to obtain the side grouping.
11. The method according to any one of claims 1 to 10, characterized in that The client model locally stored in each client is used to perform at least one of a recommendation task, a classification task, a segmentation task or a target detection task.
12. A federated learning method, characterized in that: Applied to a federated learning system, the federated learning system includes a server, multiple edge nodes, and multiple clients, each edge node is connected to the server, and the multiple clients access the multiple edge nodes, the method includes: The server receives information of edge node models respectively sent by the multiple edge nodes, wherein the information of the edge node model includes information of client models uploaded by clients in at least one edge group aggregating each of the multiple edge nodes, and the at least one edge group is obtained by grouping at least one terminal connected to the multiple edge nodes according to information of the client model uploaded by at least one terminal by each edge node; The server aggregates the information of the edge node models to obtain the information of the global model.
13. The method according to claim 12, characterized in that The method further comprises: The server sends the information of the global model to each edge node.
14. A federated learning device, characterized in that: Applied to a federated learning system, the federated learning system includes a server, multiple edge nodes and multiple clients, each edge node is connected to the server, the multiple clients access the multiple edge nodes, the federated learning device is any one of the multiple edge nodes, the federated learning includes multiple iterative processes, in one of the iterative processes, the device includes: a transceiver module, configured to obtain information of a model uploaded by at least one client, wherein the at least one client includes a client connected to the plurality of edge nodes, and the federated learning device is any one of the plurality of edge nodes; A grouping module, configured to group the at least one client according to information of the client model uploaded by the at least one client, to obtain a side group corresponding to the federated learning device; An aggregation module, used for aggregating information of client models uploaded by clients in the edge group to obtain information of edge node models; The transceiver module is further used to send the information of the edge node model to the server, and the server is used to obtain a global model by aggregating the information of the edge node models uploaded by the multiple edge nodes; The transceiver module is further used to receive the aggregated global model sent by the server; The transceiver module is also used to send the aggregated global model to each client connected to the first edge node, so that each client uses the aggregated global model to update the client model stored locally by each client, and uses the locally stored data to train the updated client model to obtain the trained client model, and the information of the trained client model is used to upload to the first edge node again.
15. The device according to claim 14, characterized in that The transceiver module is further used to obtain information of the client model uploaded by the at least one client; The grouping module is used to group the at least one client according to the information of the client model uploaded by the at least one client to obtain at least one terminal side group; The grouping module is used to determine the edge-side grouping of the federated learning device according to the at least one end-side grouping.
16. The device according to claim 15, characterized in that The grouping module is specifically used to cluster the information of the client model uploaded by the at least one client, and group the at least one client according to the clustering result to obtain at least one terminal side group.
17. The device according to claim 16, characterized in that The grouping module is specifically used for: Performing clustering according to the information of the client model uploaded by the at least one client to obtain a clustering result; Acquire the number of groups, where the number of groups includes the number of clients included in each end-side group; Grouping the at least one client according to the clustering result and the number of groups to obtain at least one terminal-side group; Evaluate the at least one end-side group to obtain an evaluation value of the at least one end-side group; If the evaluation value does not meet the preset evaluation condition, the number of groups is adjusted, and the at least one client is regrouped according to the adjusted number of groups.
18. The device according to any one of claims 15 to 17, characterized in that The grouping module is specifically used for: Obtaining similarity between locally stored model information and information of a client model uploaded by each of the at least one client, wherein the locally stored model information is information of an edge node model obtained in a previous iteration process; Based on the similarity between the locally stored model information and the information of the client model uploaded by each client, the edge group corresponding to the federated learning device is determined from the at least one edge group.
19. The device according to any one of claims 15 to 18, characterized in that The grouping module is specifically used for: Acquire communication resources corresponding to the at least one client, wherein the communication resources include communication resources used by the at least one client to communicate with the federated learning device; The edge-side group is determined according to the at least one end-side group in combination with the communication resources corresponding to the client in the at least one group.
20. The device according to any one of claims 15 to 19, characterized in that The transceiver module is also used to send information on a client model uploaded by a client in the end side group corresponding to a second edge node to the second edge node if the federated learning device determines that the at least one end side group also includes a first end side group corresponding to a second edge node, wherein the second edge node is any one of the multiple edge nodes except the federated learning device, and the information on the client model in the first end side group and the information on the client model in the edge side group of the second edge node belong to the same category, and the same category is the same clustering category after clustering the information on the client model.
21. The device according to claim 20, characterized in that The transceiver module is also used for: Sending information of the client model uploaded by the client in the terminal side group corresponding to the second edge node to the second edge node through the direct connection established with the second edge node; Alternatively, information on the client model uploaded by the client in the end-side group corresponding to the second edge node is sent to the second edge node via a backbone network in a communication network, where the communication network is the network where the federated learning device and the second edge node are located.
22. The device according to claims 14-21, characterized in that The transceiver module is further used to send information of at least one network layer in the edge node model to the server.
23. The device according to any one of claims 14 to 22, characterized in that The grouping module is further used for: Obtaining similarity between locally stored model information and information of a model uploaded by at least one client, wherein the locally stored model information is information of an edge node model obtained in a previous iteration process; According to the similarity between the locally stored model information and the model information uploaded by the at least one client, the at least one client is screened to obtain the side grouping.
24. The device according to any one of claims 14 to 22, characterized in that The client model locally stored in each client is used to perform at least one of a recommendation task, a classification task, a segmentation task or a target detection task.
25. A federated learning device, characterized in that: Applied to a federated learning system, the federated learning system includes a server, multiple edge nodes and multiple clients, each edge node is connected to the server, the multiple clients access the multiple edge nodes, the federated learning device is the server, and the device includes: a transceiver module, receiving information of edge node models respectively sent by the plurality of edge nodes, wherein the information of the edge node model includes information of a client model uploaded by a client in at least one edge group aggregating information of the client model uploaded by each edge node of the plurality of edge nodes, wherein the at least one edge group is obtained by grouping at least one terminal connected to the plurality of edge nodes according to information of the client model uploaded by the at least one terminal by each edge node; The aggregation module is used to aggregate the information of the edge node model to obtain a global model.
26. The device according to claim 25, characterized in that The device also includes: The transceiver module is also used to send the information of the global model to each edge node.
27. A federated learning system, characterized in that: comprising a plurality of servers, a plurality of edge nodes and at least one client, wherein the plurality of clients access the plurality of edge nodes; Any one of the multiple servers is used to implement the steps of any one of the methods of claims 12 to 13; The plurality of edge nodes are used to implement the steps of the method described in any one of claims 1 to 11; The at least one client is used to perform model training using locally stored data.
28. A federated learning device, characterized in that: The method comprises one or more processors, wherein the one or more processors are coupled to a memory, wherein the memory stores a program, and when the program instructions stored in the memory are executed by the one or more processors, the steps of the method according to any one of claims 1 to 11 are implemented.
29. A federated learning device, characterized in that: The method comprises one or more processors, wherein the one or more processors are coupled to a memory, wherein the memory stores a program, and when the program instructions stored in the memory are executed by the one or more processors, the steps of the method according to claim 12 or 13 are implemented.
30. A computer-readable storage medium, characterized in that: The invention comprises a program which, when executed by a processing unit, performs the method according to any one of claims 1 to 13.
31. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.
Citation Information
Patent Citations
Federal learning method, device and system
CN113469367A
Method and system for hierarchical federal learning under end-side cloud architecture and incomplete information
CN113992692A
Multilayer center asynchronous federal learning method applied to Internet of Vehicles
CN114827198A
Cloud edge-end collaborative personalized federal learning method, system, device and medium
CN115688913A
System for secure federated learning
US20200285980A1
Cited By
Federal learning client selection method based on Lyapunov optimization in mobile edge computing network
CN121984878A