Federated Learning Method

Through optimal device set selection and structured update compression algorithm, the high communication consumption problem caused by device differences in traditional federated learning is solved, and efficient and fast federated learning training is achieved, while protecting private data.

CN115099424BActive Publication Date: 2025-07-18UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210799970.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-06
Publication Date
2025-07-18
Estimated Expiration
2042-07-06

AI Technical Summary

Technical Problem

Traditional federated learning methods do not consider device variability when selecting devices, resulting in high communication consumption and training time too long, and there is a risk of privacy data leakage.

Method used

Through the optimal device set selection algorithm and structured update compression algorithm, the device with the largest contribution factor is selected for training, and the low-rank matrix decomposition and weighted average of the difference parameter matrix are performed during the communication process to reduce communication consumption.

Benefits of technology

Taking into account device differences, the communication consumption of federated learning is reduced, training efficiency is improved, and private data is protected, which is suitable for fast model training in actual scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115099424B_ABST
    Figure CN115099424B_ABST
Patent Text Reader

Abstract

The present invention discloses a federated learning method. To solve the technical problem of high communication resource consumption in federated learning, the present invention can find an optimal device set to participate in the current round of training in each round of training under the conditions that user devices have different computing capabilities, different local data, and different communication qualities, thus solving the problem of high communication consumption in traditional federated learning. The present invention completes the federated learning training more efficiently and quickly at the cost of slightly sacrificing the model accuracy. The present invention is applicable to the field of artificial intelligence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a federated learning method, and more particularly to a highly available and low communication consumption federated learning method. Background Art

[0002] The advent of the Internet era has provided convenient conditions for the collection of big data. As the amount of data grows larger, it also provides great assistance to fields such as machine learning and artificial intelligence as the raw materials for machine learning. With the protection of data privacy by people, the problem of "data islands" has become increasingly serious.

[0003] Federated learning was first proposed by Google, and then the concept of "federated machine learning" was clearly put forward at the federated learning technology and data privacy protection conference. It is a machine learning technology used to collaboratively train a global model on distributed devices or mobile devices storing local data.

[0004] At the same time, countries are strengthening the protection of data security and privacy. Traditional machine learning requires user devices to upload local data to the central server for training, which may lead to the leakage of personal privacy. Compared with traditional centralized learning, federated learning does not require data to be aggregated together, which not only reduces the transmission cost between devices but also greatly protects the privacy of data.

[0005] As a good solution to this problem, federated learning has increasingly attracted the attention of industry insiders. However, with the gradual maturity of technology, many new problems have emerged, such as the differences in user devices, the need to occupy a large amount of communication resources during the aggregation stage, and the long training time.

[0006] The bottleneck problem in the development of federated learning is the communication consumption problem. On the one hand, after each round of training, a large number of devices send their local updates to the central server. On the other hand, due to the existence of heterogeneous networks, different devices have different communication qualities, which also poses a great obstacle to synchronous federated learning.

[0007] In addition, the traditional federated learning method uses random selection when selecting devices, which means that the central server considers that all devices have the same priority when being selected. However, in actual learning scenarios, the data quality and communication environment of different devices are different, and traditional federated learning does not adapt to the real scenario.

[0008] If a method can be found that can significantly reduce the communication consumption of federated learning without significantly affecting the model accuracy on the premise of considering the differences in devices, it will undoubtedly bring a huge impetus to the further development and practical application of federated learning.

[0009] In view of this, there is an urgent need in the art for a federated learning method with high availability and low communication consumption. Summary of the Invention

[0010] To solve or alleviate some or all of the above technical problems, the present invention is implemented through the following technical solutions:

[0011] A federated learning method is applied to a system including a number of local devices and a central server. The federated learning method includes the following steps:

[0012] Step 1: The local device uploads its own data volume size, computing power, and channel quality to the central server;

[0013] Step 2: The central server initializes the global model parameters;

[0014] Step 3: The central server selects the optimal device set participating in the current round of training;

[0015] Step 4: The central server distributes the global model parameters to the local devices in all the optimal device sets participating in the current round of training;

[0016] Step 5: The devices in all the optimal device sets participating in the current round of training use local data for training;

[0017] Step 6: The devices in all the optimal device sets participating in the current round of training compress the local model parameters and upload them to the central server; the central server averages and aggregates all the received model parameters, that is, averages all the updated parameters uploaded by the devices as the global model parameters for the next round;

[0018] Step 7: Determine whether the global model has converged. If so, end the learning; if not, repeat steps 3-6.

[0019] In a certain embodiment, in step 3, the central server uses the optimal device set selection algorithm according to the time window and the total communication bandwidth resources to meet the maximization objective and selects the optimal device set participating in the current round of training; where in the t-th round, the contribution factor of local device i is specifically: where α ∈ (0, 1) is used to balance the effects of data volume and computing power on the contribution factor, N is the total number of local devices, represents whether it is selected into the optimal device set in the current round; λ i is the computing power of the local device, is the data volume size of the local device.

[0020] In a certain embodiment, after training is completed, the local model parameters are compressed by a structured update compression algorithm, and the compressed and re-encoded local model parameters are uploaded to the central server.

[0021] In a certain embodiment, the difference parameter matrix between the local model parameters and the global model parameters is uploaded to the central server.

[0022] In a certain embodiment, the structured update compression algorithm is specifically as follows:

[0023] (a) In the t-th round of learning, the central server generates a random seed and shares it with the local device i; then the pseudo-random generation algorithm pse() calculates on the central server and the local device i

[0024] (b) The local device i uses local data for training by the stochastic gradient descent method to obtain the difference parameter matrix Then it calculates the emission matrix data P through low-rank matrix factorization i t , and sends it to the central server;

[0025] (c) The central server receives the emission matrix data P i t , and calculates the difference parameter matrix

[0026] In a certain embodiment, the central server performs a weighted average innovation aggregation method on all received model parameters.

[0027] Some or all embodiments of the present invention have the following beneficial technical effects:

[0028] 1) The present invention takes into account the differences between devices in federated learning, improves the original device random selection algorithm by the selection of the optimal device set, and provides a new idea for the further development of federated learning;

[0029] 2) By the selection of the optimal device set, the device set is selected by the dynamic programming method, which ensures the rapid convergence of the global model and does not cause the leakage of local private data; in addition, the structured update algorithm is used to compress the uploaded data, reducing the communication consumption;

[0030] 3) The present invention takes into account the complex and changeable network channel conditions in the real scenario, is highly efficient and available, and greatly reduces the communication overhead, and can be implemented in industrial scenarios.

[0031] 4) In the case where user devices have different computing capabilities, different local data, and different communication qualities, an optimal device set can be found in each round of training to participate in the current round of training, solving the problem of high communication consumption in traditional learning. At the cost of slightly sacrificing model accuracy, the present invention can complete federated learning training more efficiently and quickly.

[0032] More beneficial effects will be further introduced in the preferred embodiments.

[0033] The technical solutions / features disclosed above are intended to summarize the technical solutions and technical features described in the specific implementation part. Therefore, the scope recorded may not be exactly the same. However, these new technical solutions disclosed in this part also belong to a part of the numerous technical solutions disclosed in the present invention document. The technical features disclosed in this part, the technical features disclosed in the subsequent specific implementation part, and the content in the drawings that is not clearly described in the specification are disclosed in combination with each other in a reasonable way to disclose more technical solutions.

[0034] The technical solutions formed by combining all the technical features disclosed at any position of the present invention are used to support the generalization of technical solutions, the modification of patent documents, and the disclosure of technical solutions. Brief Description of the Drawings

[0035] Figure 1 is a schematic diagram of the federated learning method proposed by the present invention;

[0036] Figure 2 is a flowchart of the federated learning method proposed by the present invention. Detailed Description of the Preferred Embodiments

[0037] Since it is impossible to exhaustively describe all alternative solutions, the key points of the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. For other technical solutions and details not disclosed in detail below, generally they are all technical objectives or technical features that can be achieved by conventional means in the art. Due to space limitations, the present invention will not introduce them in detail.

[0038] Unless it means division, " / " at any position in the present invention represents logical "or". The serial numbers such as "first" and "second" at any position in the present invention are only used as distinguishing marks in description, and do not imply an absolute order in time or space, nor do they imply that the terms with such serial numbers are necessarily different from the same terms with other attributives.

[0039] The present invention will describe the key points for combining various different specific embodiments. These key points will be combined into various methods and products. In the present invention, even if only the key points described when introducing the method / product solution are mentioned, it means that the corresponding product / method solution also clearly includes this technical feature.

[0040] When it is described in the present invention that there exists or includes a certain step, module, or feature at any position, it does not imply that this existence is the only exclusive existence. Those skilled in the art can completely obtain other embodiments by supplementing other technical means according to the technical solutions disclosed in the present invention. Based on the key points described in the specific embodiments of the present invention, those skilled in the art can completely apply means such as replacement, deletion, addition, combination, and order swapping to certain technical features to obtain a technical solution that still follows the inventive concept of the present invention. These solutions that do not deviate from the inventive concept of the present invention are also within the protection scope of the present invention.

[0041] In view of the foregoing problems and scenarios, the present invention is premised on data being stored locally on the user device, and multiple parties cooperate together to greatly reduce the communication duration and improve the learning efficiency while slightly sacrificing the learning accuracy.

[0042] Such as Figure 2 FIG. is a specific flowchart of a federated learning method with high availability and low communication consumption. This federated learning method is applied to a system including several local devices and a central server, and specifically includes the following steps:

[0043] Step 1: The local device uploads its own data volume size, computing power, and channel quality to the central server.

[0044] All participating devices (i.e., local devices, hereinafter referred to as the local machine) upload their own information, including the local computing power, local data volume size, and the channel size allocated to the local machine, to the central server in advance through the channel. Usually, a local device refers to an edge node with computing power and storage capacity.

[0045] Specifically, in step 1, in the t-th (positive integer) round of training, the local information uploaded by local device i to the central server is: the local computing power is λ i , the local data volume size is The channel quality of the local device (i.e., communication quality information) is manifested as the allocated bandwidth in the system being These information are originally publicly available information and not private data owned by the local.

[0046] In step 1, the local data owned by the local device is private information, but it will not be leaked throughout the process. The information uploaded in step 1 is publicly available information, which will not leak local privacy data and does not require uploading additional information.

[0047] Step 2: The central server initializes the global model parameters.

[0048] The central server summarizes the information and stores it in the form of a multi-tuple; for example, it is

[0049] The differences among devices are reflected in the amount of data and local computing power. In actual scenarios, the more local data a device has and the stronger its local computing power, the greater its contribution to the global model of federated learning should be. The differences among devices can distinguish different devices. Considering all devices as having the same priority in traditional federated learning does not conform to the actual scenario either.

[0050] Now, define the contribution of local device i as the contribution factor Specifically: where α ∈ (0, 1) is used to balance the effects of data volume and computing power on the contribution factor; devices with larger contribution factors should have higher priorities when selecting devices.

[0051] The amount of data uploaded by each local device is b ∈ 0(|W|×(H(ΔW)+η))), where |W| represents the size of the parameter matrix, H(ΔW) represents the entropy increase of parameter updates in the upload phase, and η represents the inadequacy of network coding. The entropy increase is reflected in the difference between the local model parameters and the global model parameters uploaded in this round. If the gap is larger, the rank of its model parameter matrix is higher, and the amount of data uploaded after compression will be larger. The inadequacy of coding is reflected in the choice of coding scheme, and an efficient coding scheme will reduce the amount of data uploaded.

[0052] The upload delay of each device is:

[0053]

[0054] where k is the compression ratio of the model and SNR represents the signal-to-noise ratio.

[0055] Step 3: The central server selects the optimal set of devices participating in this round of training.

[0056] Based on the time window and the total communication bandwidth resources, the central server uses the optimal device set selection algorithm to the objective function to select the optimal set of devices participating in this round of training (learning).

[0057] The central server uses the optimal device set selection algorithm to select devices to participate in this round of training. Preferably, the best set of devices should meet the requirement of maximizing the sum of the contribution factors of all local devices under the constraints of the time window T and the total communication bandwidth resource C. That is, under the satisfaction of:

[0058]

[0059]

[0060] where N is the total number of local devices, The meaning of this is whether device i is selected to enter the optimal device set in the t-th round.

[0061] To solve the above 0-1 programming problem, the present invention proposes a bivariate dynamic programming algorithm. First, define I {i} as a device selection set, which contains a total of i devices numbered from 1 to i. Define as the optimal contribution value when the device set is I {i} Then, the state transition equation is for the i-th device. If C i > C, or there is an equation:

[0062]

[0063] In other cases, there is an equation:

[0064]

[0065] Finally, after multiple rounds of iteration, the optimal device set can be found.

[0066] Step 4: The central server distributes the global model parameters to the local devices in all the optimal device sets participating in this round of training.

[0067] Step 5: The devices in all the optimal device sets participating in this round of training use local data for training.

[0068] For example, the devices in all the optimal device sets participating in this round of training adopt the stochastic gradient descent method. Any reasonable training method in the present invention is feasible, and the present invention is not limited to a specific means.

[0069] Step 6: The devices in all the optimal device sets participating in this round of training compress the local model parameters and upload them to the central server; the central server averages and aggregates all the received model parameters, that is, averages all the updated parameters uploaded by the devices as the global model parameters for the next round.

[0070] After the training is completed, the local model parameters are compressed by the structured update compression algorithm, and the compressed and re-encoded local model parameters are uploaded to the central server.

[0071] Preferably, when the local device uploads the local model parameters (also called the model matrix) to the central server, it does not need to upload the entire model parameters, but only needs to upload the difference parameter matrix between the local model parameters and the global model parameters; this difference parameter matrix is a low-rank matrix. In this case, the model parameters received by the central server are this difference parameter matrix.

[0072] Preferably, the local model parameters are compressed by the structured update compression algorithm, and the compression method used is the structured update compression method. Suppose the local difference parameter matrix is As described above, it is a low-rank matrix. For simplicity of representation, its rank is described as k. Therefore, the parameter matrix can be described as the product of two matrices, that is where is the reconstruction matrix, is the emission matrix. The reconstruction matrix data is randomly generated by the central server through a pseudo-random function at the beginning of the t-th round and sent to the local device i.

[0073] The specific steps of applying the aforementioned compression method to federated learning are described as follows:

[0074] (a) In the t-th round of training, the central server generates a random seed and shares it with the local device i. Then the pseudo-random generation algorithm pse() calculates on the central server and the local device i For security reasons, a random seed needs to be regenerated for each round of learning;

[0075] (b) The local device i uses the local data for training by the stochastic gradient descent method to obtain the difference parameter Then it calculates the emission matrix data P through low-rank matrix factorization i t and sends it to the central server;

[0076] (c) The central server receives the emission matrix data P i t , calculates to obtain the difference parameter Using this compression scheme, the local device saves k / m of the upload data volume during the model upload phase.

[0077] The central server performs a weighted average aggregation method on all the received model parameters.

[0078] Step 7: Determine whether the global model has converged. If so, end the learning; if not, repeat steps 3-6.

[0079] Although the present invention has been described with reference to specific features and embodiments of the present invention, various modifications, combinations, and substitutions can still be made without departing from the present invention. The protection scope of the present invention is intended to be not limited to the specific embodiments of the processes, machines, manufactures, compositions of matter, devices, methods, and steps described in the specification, and these methods and modules may also be implemented in one or more related, interdependent, cooperating, front / back-stage products and methods.

[0080] Therefore, the description and the drawings should simply be regarded as introductions to partial embodiments of the technical solutions defined by the appended claims. Thus, the appended claims should be interpreted according to the principle of the most reasonable interpretation, and are intended to cover all modifications, variations, combinations or equivalents within the scope of the disclosure of the present invention as much as possible, while also avoiding unreasonable interpretation methods.

[0081] In order to achieve better technical effects or for the needs of certain applications, those skilled in the art may make further improvements to the technical solutions based on the present invention. However, even if this part of the improvement / design is creative or / and progressive, as long as it relies on the technical concept of the present invention and covers the technical features defined by the claims, this technical solution should also fall within the protection scope of the present invention.

[0082] There may be alternative technical features for some of the technical features mentioned in the appended claims, or the order of certain technical processes or the order of material organization can be reorganized. After those of ordinary skill in the art know the present invention, it is easy to think of these replacement means, or change the order of technical processes or the order of material organization, and then use basically the same means to solve basically the same technical problems and achieve basically the same technical effects. Therefore, even if the above means or / and order are clearly defined in the claims, these modifications, changes and replacements should fall within the protection scope of the claims according to the doctrine of equivalents.

[0083] Combined with the method steps or modules described in the embodiments disclosed herein, they can be implemented in hardware, software, or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the steps and components of each embodiment have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application or design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered outside the scope claimed by the present invention.

Claims

1. A federated learning method is applied to a system composed of a number of local devices and a central server, characterized in that, The federated learning method includes the following steps: Step 1: The local device uploads its own data volume size, computing power, and channel quality to the central server; Step 2: The central server initializes the global model parameters; Step 3: The central server selects the optimal device set for this round of training; the central server uses the optimal device set selection algorithm according to the time window and the total communication bandwidth resources to select the optimal device set for this round of training to meet the goal of maximizing MAX( ); among them, in the t-th round, the contribution factor of the local device i is specifically: , where is used to balance the effects of the data volume and computing power on the contribution factor, N is the total number of local devices, represents whether it is selected to enter the optimal device set in this round; is the computing power of the local device, is the data volume size of the local device; Step 4: The central server distributes the global model parameters to the local devices in the set of optimal devices participating in this round of training; Step 5: The devices in the set of optimal devices participating in this round of training use local data for training; Step 6: The devices in the set of optimal devices participating in this round of training are compressed through the structured update compression algorithm. By generating a random seed and using the pseudo-random generation algorithm, a new seed is regenerated in each round. The local device calculates the difference parameter matrix through the stochastic gradient descent method, compresses the local model parameters, and uploads them to the central server; The central server performs averaging aggregation on all received model parameters, that is, averages the update parameters uploaded by all devices as the global model parameters for the next round; Step 7: Determine whether the global model has converged. If so, end the learning; if not, repeat steps 3-6.

2. The federated learning method according to claim 1, wherein: The specific structured update compression algorithm is as follows: (a) In the t-th round of learning, the central server generates a random seed and shares it with the local device i; then the pseudo-random generation algorithm pse () calculates on the central server and the local device i ; (b) The local device i uses the local data for training by the stochastic gradient descent method to obtain the difference parameter matrix ; Then, it is calculated through low-rank matrix factorization to obtain the emission matrix data , and it is sent to the central server; (c) The central server receives the radiation matrix data , and calculates the difference parameter matrix .

3. The federated learning method according to claim 1, wherein: The central server performs weighted average aggregation on all received model parameters.