A General Federated Learning Implementation Method Applicable to Heterogeneous Networks
By adopting the federated learning method of central node allocation tasks and child node adaptive training in a heterogeneous network environment, the problem of federated learning in the existing technology is difficult to apply under heterogeneous networks, and the effective participation of weak computing nodes and efficient construction of global models is achieved.
Patent Information
- Application Number
- CN202111582604.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-12-22
AI Technical Summary
The existing federated learning methods are difficult to effectively apply in heterogeneous network environments with weak computing power and resource-constrained resources, especially in scenarios of mobile terminals and heterogeneous nodes. The limitations of computing resources and communication resources make it difficult to synchronize and train.
A general federated learning implementation method suitable for heterogeneous networks is proposed. The basic federated learning tasks are assigned through the central node, and the child nodes are allowed to perform adaptive training and privacy protection. The updated information is weighted and aggregated by the global optimizer to obtain a new round of global model parameters.
This method enables weak computing nodes and heterogeneous nodes to adaptively participate in federated learning tasks, make full use of the data and computing resources of each node, construct a global model with superior generalization performance, and keep the system running stable when resources are tight.
Smart Images

Figure CN114254761B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to federated learning, and particularly to a general federated learning implementation method applicable to heterogeneous networks. Background Art
[0002] Data-driven artificial intelligence has become a pillar technology in today's big data information era with the wide application of deep learning, and has gradually penetrated into all aspects of social life. However, the existence of data islands hinders the further penetration of this technology into sensitive fields, such as the financial and medical fields, as well as some fields involving enterprise-level commercial data. As a privacy-aware distributed training solution, federated learning has received extensive attention from the academic and industrial communities in recent years, and there are currently some enterprise-level implementation applications that can support the federated learning training tasks of data centers.
[0003] However, existing federated learning methods have serious technical bottlenecks in mobile devices with limited computing and communication resources, and more generally in heterogeneous nodes:
[0004] Existing privacy protection technologies, such as secure multi-party computing and differential privacy, often require huge computing resources, and general federated learning methods require a relatively frequent parameter exchange process, which also brings unbearable communication overhead to nodes with weak computing power and communication resources. These factors make existing federated learning algorithms unable to be well applied in mobile devices;
[0005] Existing federated learning methods, especially horizontal federated learning methods, often only target homogeneous nodes, that is, nodes with similar computing resources and network resources, and each node adopts the same privacy protection strategy and training strategy, which ignores the fact that nodes with training requirements in reality are often heterogeneous in terms of hardware resources, restricting the application prospects of federated learning. Summary of the Invention
[0006] (1) Technical Problems to be Solved
[0007] In view of the above-mentioned disadvantages of the prior art, the present invention provides a general federated learning implementation method applicable to heterogeneous networks, which can effectively overcome the defect that existing federated learning algorithms are difficult to be implemented in weak nodes and heterogeneous network environments.
[0008] (2) Technical Solutions
[0009] To achieve the above object, the present invention is realized through the following technical solutions:
[0010] A general federated learning implementation method applicable to heterogeneous networks includes the following steps:
[0011] S1. Establish a file for the resource description and data description of child nodes;
[0012] S2. The central node distributes basic federated learning tasks to other heterogeneous asynchronous child nodes and sets a deadline for each child node to execute the tasks;
[0013] S3. The child nodes perform adaptive training and send the updated information to the central node together;
[0014] S4. The central node uses a global optimizer to perform weighted aggregation on the uploaded updated information to obtain a new round of global model parameters.
[0015] Preferably, in S1, a file is established for the resource description and data description of the child nodes, including:
[0016] Establishing a file for resource description, that is, establishing a file for the available computing resources and network status at the current clock; establishing a file for data description, that is, storing and establishing a file for the source, label, and structural metadata of the local data;
[0017] Among them, the data description includes two aspects:
[0018] Data usage record. In some cases, local data may participate in training multiple times or only once, which will affect the update error of the local model parameters of the child nodes;
[0019] Data attribute description. The child nodes need to fully describe the finer-grained attributes of the local data so as to quickly screen the data during training.
[0020] Preferably, in S2, the central node distributes basic federated learning tasks to heterogeneous asynchronous child nodes and gives a deadline for each child node to execute the tasks, including:
[0021] The central node sets a basic federated learning task and sends it to each child node together with the initial global model parameters. Considering the high asynchrony of the system, the central node gives a deadline for each child node to execute the tasks and excludes abnormal nodes by sending a confirmation time to each child node.
[0022] Preferably, the content sent by the central node is within the range of the system communication load, and at the same time, the central node monitors the network status of each child node through the deadline mechanism.
[0023] Preferably, in S3, the child nodes perform adaptive training and send the updated information to the central node together, including:
[0024] The child nodes screen the qualified sample data locally according to the training task requirements and the initial global model parameters;
[0025] The child node reasonably arranges the number and frequency of local iteration samples and a suitable privacy protection strategy according to the locally available computing resources, and obtains updated local model parameters;
[0026] The child node sends the updated local model parameters, the sample information and iteration hyperparameters participating in the iteration, and the privacy protection strategy adopted to the central node.
[0027] Preferably, the child node reasonably arranges the number and frequency of local iteration samples and a suitable privacy protection strategy according to the locally available computing resources, including:
[0028] In the case of extremely scarce locally available computing resources, only a small number of samples are selected, and the updated information is directly transmitted without adopting any privacy protection strategy;
[0029] In the case of abundant locally available computing resources, the differential privacy strategy is adopted, and the updated information is sent to the central node at the end of each local iteration;
[0030] In the case of poor local network status, the updated information is compressed to reduce communication overhead.
[0031] Preferably, before the child node performs adaptive training in S3, it includes:
[0032] The central scheduler schedules computing resources and data resources for each child node according to the training task requirements, and adjusts the amount of transmitted data and communication frequency between each child node and the central node.
[0033] Preferably, in S4, the central node uses a global optimizer to perform weighted aggregation on the uploaded updated information to obtain a new round of global model parameters, including:
[0034] The global model performs weighted aggregation through a global optimizer according to the updated local model parameters, training sample conditions, and privacy protection strategies adopted by each child node to obtain a new round of global model parameters.
[0035] Preferably, the global optimizer has two goals:
[0036] The weighted loss of the global model is minimized, and the global weighted loss function is:
[0037]
[0038] where, L optimize is the global weighted loss function, f is the global model, w is the weight of the global model, and the weight of node i is based on the number of training samples, the local data bias parameter δ i and the data accuracy loss γ caused by the local privacy protection strategy iDetermine, f i (w, X i ) = l(w, X i ) is the learning task for each child node, and f i is the local model parameter of node i, and X i is the training sample set of node i for the current iteration;
[0039] Meanwhile, the global optimizer also needs to optimize the learning tasks assigned in the next round according to the current network topology and heterogeneous network. The specific basis is as follows:
[0040]
[0041] Among them, M and N respectively represent the number of edges and nodes in the current network topology, and C i,compute represents the computational overhead of training and privacy protection at node i, and C j,communicate represents the communication overhead of the j-th edge, which depends on the amount of data transmitted and the communication frequency in the adopted communication protocol.
[0042] Preferably, before the central node uses the global optimizer to perform weighted aggregation on the uploaded update information in S4, it includes:
[0043] The central node counts the upload situation of the update information of each child node after the deadline of one round, and discards the child nodes that do not upload the update information in time.
[0044] (III) Beneficial effects
[0045] Compared with the prior art, a general federated learning implementation method applicable to heterogeneous networks provided by the present invention has the following beneficial effects:
[0046] 1) Utilize the high degree of autonomy of child nodes to adopt targeted training strategies and privacy protection strategies for each node, enabling weak computing power nodes to also participate in the federated learning task;
[0047] 2) Enable heterogeneous nodes, such as mobile terminals and data servers, to participate in the same federated learning task together, so that the computing power and data of heterogeneous devices that are difficult to synchronize can be fully utilized;
[0048] 3) Since the autonomy of nodes during training is fully mobilized, each node can adopt different training strategies and privacy protection strategies, thereby being able to fully utilize resources such as data and computing power of each node to construct a global model with excellent generalization performance, so that node data physically isolated in space can participate in the global model training;
[0049] 4) Under the dynamic decision-making of the central scheduler and the global optimizer, the entire system can still operate when resources are tight, without problems such as synchronization difficulties and unstable training round times due to network or computing resource limitations. Description of the Drawings
[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0051] Figure 1 It is a flowchart of the present invention;
[0052] Figure 2 It is a schematic diagram of the communication link in a certain round of iteration of federated learning in the present invention. Detailed Implementation Modes
[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0054] A general method for implementing federated learning applicable to heterogeneous networks, as Figure 1 shown, includes the following steps:
[0055] S1. Establish files for the resource descriptions and data descriptions of the child nodes;
[0056] S2. The central node distributes basic federated learning tasks to other heterogeneous asynchronous child nodes and sets a deadline for each child node to execute the tasks;
[0057] S3. The child nodes perform adaptive training and send the updated information to the central node together;
[0058] S4. The central node uses the global optimizer to perform weighted aggregation on the uploaded updated information to obtain a new round of global model parameters.
[0059] ① Establishing files for the resource descriptions and data descriptions of the child nodes includes:
[0060] Resource description archiving, i.e., archiving the available computing resources (CPU, GPU, etc.) and network status (network bandwidth and latency, etc.) at the current clock; data description archiving, i.e., storing and archiving the source, tags, and structural metadata of local data.
[0061] Among them, data description includes two aspects:
[0062] Data usage records, local data may participate in training multiple times or only once in some cases, which will affect the local model parameter update error of the child node. Therefore, this usage record is recorded locally.
[0063] Data attribute description, the child node needs to fully describe the local data with respect to finer-grained attributes for quick data screening during training.
[0064] In most cases, the child node will randomly sample and train locally according to the data usage. However, in some other application scenarios, such as having a special preference for data with special sources or special attributes (such as preferring data on the user's mobile phone or data collected during a certain time period), this part of the preference will be described in the basic federated learning task initially sent by the central node to each child node. To meet this requirement, it is necessary to archive the data attribute description in the data description.
[0065] ② The central node assigns basic federated learning tasks to heterogeneous and asynchronous child nodes and gives a deadline for each child node to execute the task, including:
[0066] The central node sets a basic federated learning task (data label and source requirements, training objectives, and hyperparameters, etc.) and sends it together with the initial global model parameters to each child node. Considering the high asynchrony of the system, the central node gives a deadline for each child node to execute the task and excludes abnormal nodes (offline or computationally hardware-abnormal child nodes) by sending a confirmation time to each child node.
[0067] Among them, the content sent by the central node is within the system communication load range, and at the same time, the central node monitors the network status of each child node through the deadline mechanism.
[0068] ③ Before the child node performs adaptive training, the central scheduler schedules the computing resources and data resources for each child node according to the training task requirements and adjusts the amount of transmitted data and communication frequency between each child node and the central node.
[0069] The child node performs adaptive training and sends the update information to the central node together, including:
[0070] The child node filters the eligible sample data locally according to the training task requirements and the initial global model parameters;
[0071] The child node reasonably arranges the number and times of samples for local iteration, as well as the suitable privacy protection strategy according to the local available computing resources, and obtains the updated local model parameters;
[0072] The child node sends the updated local model parameters, the sample information and iteration hyperparameters participating in the iteration, and the adopted privacy protection strategy to the central node together.
[0073] Among them, the child node reasonably arranges the number and times of samples for local iteration, as well as the suitable privacy protection strategy according to the local available computing resources, including:
[0074] In the case of extremely scarce local available computing resources, only a small number of samples are selected, and the updated information is directly transmitted without any privacy protection strategy;
[0075] In the case of abundant local available computing resources, the differential privacy strategy is adopted, and the updated information is sent to the central node at the end of each local iteration;
[0076] In the case of poor local network status, the updated information is compressed to reduce the communication overhead.
[0077] In the technical solution of this application, targeted training strategies and privacy protection strategies are adopted for each node by using the high degree of autonomy of the child node, enabling weak computing power nodes to also participate in the federated learning task. The high degree of autonomy of the child node is the key to the effectiveness of the heterogeneous training strategy.
[0078] The local adaptive training work needs to be quickly fed back to the central node. Therefore, it is necessary to schedule the training and privacy protection tasks according to the local resource description in the most recent time window, and this local task scheduling should be lightweight. Since some privacy protection strategies will cause damage to the accuracy of the clear text of the updated information, the updated information must be sent to the central node at the end of each local iteration.
[0079] ④ Before the central node uses the global optimizer to perform weighted aggregation on the uploaded updated information, after the arrival of the deadline for one round, the central node counts the upload situation of the updated information of each child node, and discards the child nodes that have not uploaded the updated information in time.
[0080] The central node uses the global optimizer to perform weighted aggregation on the uploaded updated information to obtain the new round of global model parameters, including:
[0081] The global model performs weighted aggregation through a global optimizer based on the updated local model parameters uploaded by each child node, the training sample situation (quantity and label-related information), and the privacy protection strategy adopted (lossy or lossless data precision), to obtain a new round of global model parameters.
[0082] The global optimizer has two objectives:
[0083] a. The weighted loss of the global model is minimized, and the global weighted loss function is:
[0084]
[0085] where L optimize is the global weighted loss function, f is the global model, w is the weight of the global model, and the weight of node i is determined according to the number of training samples, the local data bias parameter δ i , the data precision loss γ caused by the local privacy protection strategy i , f i (w, X i ) = l(w, X i ) is the learning task of each child node, f i is the local model parameter of node i, and X i is the training sample set of node i in the current iteration;
[0086] b. The global optimizer also needs to optimize the learning tasks assigned in the next round according to the current network topology and heterogeneous network. The specific basis is:
[0087]
[0088] where M and N respectively represent the number of edges and nodes in the current network topology, C i,compute represents the computational overhead of training and privacy protection at node i, and C j,communicate represents the communication overhead of the j-th edge, which depends on the amount of data transmitted and the communication frequency in the adopted communication protocol.
[0089] In the technical solution of this application, under the dynamic decision of the central scheduler and the global optimizer, the entire system can still operate when resources are scarce, and problems such as synchronization difficulties and unstable training round times caused by network or computational resource limitations will not occur.
[0090] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A general federated learning implementation method applicable to heterogeneous networks, characterized in that: It includes the following steps: S1. Establish files for the resource description and data description of the child nodes; S2. The central node distributes basic federated learning tasks to other heterogeneous asynchronous child nodes and sets a deadline for each child node to execute the tasks; S3. The child nodes perform adaptive training and send the updated information to the central node together; S4. The central node uses a global optimizer to perform weighted aggregation on the uploaded updated information to obtain a new round of global model parameters; In S3, the child nodes perform adaptive training and send the updated information to the central node together, including: The child nodes locally screen the eligible sample data according to the training task requirements and the initial global model parameters; The child nodes reasonably arrange the number and times of locally iterated samples and a suitable privacy protection strategy according to the locally available computing resources to obtain updated local model parameters; The child nodes send the updated local model parameters, the sample information and iteration hyperparameters participating in the iteration, and the adopted privacy protection strategy to the central node together; The child nodes reasonably arrange the number and times of locally iterated samples and a suitable privacy protection strategy according to the locally available computing resources, including: In the case of extremely scarce locally available computing resources, only select a small number of samples and directly transmit the updated information without adopting any privacy protection strategy; In the case of abundant locally available computing resources, adopt the differential privacy strategy and send the updated information to the central node together at the end of each local iteration; In the case of poor local network status, compress the updated information to reduce the communication overhead.
2. The general federated learning implementation method applicable to heterogeneous networks according to claim 1, characterized in that: In S1, establishing files for the resource description and data description of the child nodes includes: Establishing a file for resource description, that is, establishing a file for the available computing resources and network status at the current clock; establishing a file for data description, that is, storing and establishing a file for the source, label, and structural metadata of the local data; Among them, the data description includes two aspects: Data usage record. In some cases, the local data may participate in training multiple times or only once, which will affect the update error of the local model parameters of the child nodes; Data attribute description. The child nodes need to fully describe the local data about finer-grained attributes in order to quickly screen the data during training.
3. The general federated learning implementation method applicable to heterogeneous networks according to claim 1, characterized in that: In S2, the central node distributes basic federated learning tasks to heterogeneous asynchronous child nodes and gives a deadline for each child node to execute the tasks, including: The central node sets a basic federated learning task and distributes it to each child node together with the initial global model parameters. At the same time, considering the high asynchrony of the system, the central node gives a deadline for each child node to execute the tasks and excludes abnormal nodes by sending a confirmation time to each child node.
4. The general federated learning implementation method applicable to heterogeneous networks according to claim 3, characterized in that: The content sent by the central node is within the system communication load range, and at the same time, the central node monitors the network status of each child node through a deadline mechanism.
5. The general federated learning implementation method applicable to heterogeneous networks according to claim 1, characterized in that: Before the child nodes perform adaptive training in S3, it includes: The central scheduler schedules computing resources and data resources for each child node according to the training task requirements, and adjusts the amount of transmitted data and communication frequency between each child node and the central node.
6. The general federated learning implementation method applicable to heterogeneous networks according to claim 1, characterized in that: In S4, the central node uses a global optimizer to perform weighted aggregation on the uploaded update information to obtain a new round of global model parameters, including: The global model performs weighted aggregation through a global optimizer based on the updated local model parameters, training sample conditions, and privacy protection strategies uploaded by each child node to obtain a new round of global model parameters.
7. The general federated learning implementation method applicable to heterogeneous networks according to claim 6, characterized in that: The global optimizer has two objectives: The weighted loss of the global model is minimized, and the global weighted loss function is: Among them, L optimize is the global weighted loss function, f is the global model, w is the weight of the global model, and the weight of node i is determined according to the number of training samples and the local data bias parameter δ i , the data precision loss γ caused by the local privacy protection strategy i ; f i (w, X i ) = l(w, X i ) is the learning task of each child node, fi is the local model parameter of node i, and X i is the training sample set of node i in the current iteration; At the same time, the global optimizer also needs to optimize the learning tasks assigned in the next round according to the current network topology and heterogeneous network, and the specific basis is: where M and N represent the number of edges and nodes in the current network topology, respectively, and C i,compute represents the computational overhead of training and privacy protection at node i, and C j,communicate represents the communication overhead of the j-th edge, which depends on the amount of data transmitted and the communication frequency in the adopted communication protocol.
8. The general federated learning implementation method applicable to heterogeneous networks according to claim 6 or 7, characterized in that: Before the central node in S4 uses a global optimizer to perform weighted aggregation on the uploaded update information, it includes: After a round of deadline arrives, the central node counts the upload situation of the update information of each child node, and discards the child nodes that have not uploaded the update information in time.
Citation Information
Patent Citations
Data source selection method for multi-source heterogeneous data fusion
CN111309755A
Transverse federation learning system optimization method, device and equipment and readable storage medium
CN111522669A
Hierarchical federated learning method and device based on asynchronous communication, terminal equipment and storage medium
CN112532451A
Federal learning training method based on training stage perception strategy
CN113723619A