Communication optimization federated learning method based on self-guiding evolutionary strategy

By introducing a self-guided evolution strategy in federated learning, the communication between users and servers is optimized, and the vector group is evaluated using fitness values ​​and pseudo-stochastic gradients, the problem of inefficiency in traditional federated learning is solved, and more efficient model training convergence and performance improvement is achieved.

CN119940478AActive Publication Date: 2025-05-06JINAN UNIVERSITY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510007584.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

Traditional centralized learning methods have challenges in data privacy and security, especially in the healthcare and finance industries, where inefficient communications of federated learning are its main bottleneck.

Method used

A federated learning method for communication optimization based on a self-guided evolution strategy is proposed. Through communication optimization between the user and the server, the user replaces the local gradient update with the fitness value and uploads it to the server. The server replaces the global gradient with the global fitness to each user. The pseudo-stochastic gradient evaluation vector group is generated under the guidance of the global gradient in the previous rounds, improving communication efficiency and model performance.

Benefits of technology

It significantly reduces communication overhead, improves the convergence efficiency and model performance of federated learning, and improves the effectiveness of search direction by adaptively using historical estimation gradients, and accelerates the training convergence of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940478A_ABST
    Figure CN119940478A_ABST
Patent Text Reader

Abstract

The invention discloses a communication optimization federated learning method based on a self-guiding evolutionary strategy, which belongs to the technical field of privacy computing, and comprises the following steps of: dividing a real number space into a main space and an orthogonal complement space from a real number space by using global model gradient vectors of the first few rounds by a user, and respectively extracting pseudo-random evaluation vectors in the main space and the orthogonal complement space; the method comprises the following steps: converting a high-dimensional model gradient vector into fitness values of a plurality of pseudo-random evaluation vectors, and sending the fitness values to a server; and the server aggregates the local fitness value and calculates a global gradient vector by using the evaluation vector, and sends the global fitness value to the user, and the user and the server calculate the global gradient vector of the current round based on the global fitness value and the pseudo-random evaluation vector, and update the global model. According to the method, the communication overhead can be remarkably reduced, and the effectiveness of the search direction can be improved by adaptively utilizing the historical estimation gradient, so that convergence is accelerated, and the model performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of privacy computing technology, and in particular relates to a communication optimization federated learning method based on a self-guided evolutionary strategy. Background Art

[0002] With the development of big data and cloud computing technologies, machine learning plays an increasingly important role in fields such as autonomous driving, speech recognition, image classification, and disease detection. In order to better extract valuable information from massive data, machine learning systems are often deployed in architectures containing thousands of processors. However, traditional centralized learning methods face challenges in data privacy and security, especially in the healthcare and financial industries. Federated Learning (FL), as a distributed learning framework, allows multiple devices (called users) to jointly train a model without sharing raw data, aiming to solve the above problems. However, one of the main bottlenecks of FL is inefficient communication, especially when dealing with deep learning models with a large number of parameters. Therefore, studying how to reduce communication costs has become an important topic in the field of federated learning.

[0003] In order to alleviate the communication bottleneck in federated learning, researchers have proposed a variety of technical means. These technologies mainly include model compression, knowledge distillation, and client update sampling. Although the existing technologies for optimizing federated learning communication have their own advantages, they also expose many limitations and challenges. Existing methods either require complex preprocessing steps, or sacrifice the accuracy and robustness of the model, or are difficult to implement in a large-scale distributed environment. Therefore, it is particularly important to find a new way to significantly reduce communication requirements while maintaining good model performance. Summary of the invention

[0004] In order to solve the above technical problems, the present invention proposes a communication optimization federated learning method based on a self-guided evolutionary strategy to solve the problems existing in the above prior art.

[0005] To achieve the above object, the present invention provides a communication optimization federated learning method based on a self-guided evolutionary strategy, comprising:

[0006] Step 1: The server initializes global model parameters and random number seeds and sends them to each user;

[0007] Step 2: The user performs training through local data to obtain a local gradient vector; a subspace is divided from the real number space based on the global gradient vector of a preset number of rounds and the random number seed, and a number of pseudo-random evaluation vectors are extracted from the subspace; the local fitness value of each pseudo-random vector to the local gradient vector is calculated and fed back to the server;

[0008] Step 3: The server calculates the global fitness value based on the local fitness value fed back by each user and sends it to each user;

[0009] Step 4: The user and the server calculate the global gradient vector of this round based on the global fitness value and the pseudo-random evaluation vector, and update the global model;

[0010] Repeat steps 2 to 4 to continuously update the global model.

[0011] Optionally, the process of obtaining the pseudo-random evaluation vector in step 2 includes:

[0012] The user trains the local model through local data to obtain a local gradient vector; if the number of training rounds of the local model is less than or equal to the preset number of rounds, a pseudo-random gradient evaluation vector is generated based on the random number seed; if the number of training rounds of the local model is greater than the preset number of rounds, a subspace is divided from the real number space based on the global gradient vector and the random number seed, and a number of pseudo-random evaluation vectors are extracted from the subspace.

[0013] Optionally, in the step 2, the pseudo-random gradient evaluation vector generated based on the random number seed obeys a normal distribution.

[0014] Optionally, in step 2, a subspace is divided from the real number space based on the global gradient vector and the random number seed, and a process of extracting a plurality of pseudo-random evaluation vectors from the subspace includes:

[0015] The matrix composed of the global gradient vectors of the first few rounds of a preset number of rounds is subjected to matrix decomposition and divided into two parts: a main space and an orthogonal complement space. The main space and the orthogonal complement space are sampled according to a preset sampling ratio to obtain a number of pseudo-random evaluation vectors.

[0016] Optionally, in step 2, the formula for calculating the local fitness value of each pseudo-random vector to the local gradient vector is as follows:

[0017]

[0018] in, represents the local gradient vector obtained by user u in the tth round of training, represents the local fitness value of the i-th pseudo-random gradient vector, and m is the number of pseudo-random gradient vectors.

[0019] Optionally, in step 3, the formula for calculating the global fitness value is as follows:

[0020]

[0021] Among them, w uis the weight of user u, U is the total number of users, f i t is the global fitness value.

[0022] Optionally, in step four, the server generates a pseudo-random evaluation vector set using the method of step two, the server and the user respectively calculate the global gradient of this round and update the model parameters, and the user saves the global gradient for the next round of training.

[0023] Optionally, the formula for calculating the global gradient and model parameters of this round is as follows:

[0024]

[0025] θ t+1 =θ t -η*g t

[0026] Among them, g t is the global gradient obtained by aggregation in the tth round, η is the learning rate, θ t+1 is the t+1th round global model, m is the number of pseudo-random gradient evaluation vectors, A vector of pseudo-random gradient evaluations.

[0027] Compared with the prior art, the present invention has the following advantages and technical effects:

[0028] The user replaces the local gradient update with the fitness value and uploads it to the server, and the server also replaces the global gradient with the global fitness and sends it to each user. Compared with the high-dimensional gradient vector, the transmission of the fitness value reduces a lot of communication overhead; at the same time, since the pseudo-random gradient evaluation vector group is generated under the guidance of the global gradient of the previous rounds, it has a certain self-guidance characteristic, which makes the pseudo-random gradient evaluation vector group higher in quality, and the convergence efficiency of federated learning is greatly improved compared with the traditional evolutionary strategy. The present invention can not only significantly reduce the communication overhead, but also improve the effectiveness of the search direction by adaptively utilizing the historical estimated gradient, thereby accelerating convergence and improving model performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0030] Figure 1 The present invention is a flowchart of a communication optimization federated learning method based on a self-guided evolutionary strategy according to an embodiment of the present invention. DETAILED DESCRIPTION

[0031] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0032] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0033] Embodiment 1

[0034] like Figure 1 As shown, this embodiment provides a communication optimization federated learning method based on a self-guided evolutionary strategy, including:

[0035] The present invention involves two different parties: users and servers, which are described as follows:

[0036] User: There are multiple users, whose functions are: provide a local private data set that has completed training to implement local model training; use a shared random number seed and the global gradient vector of the previous k rounds to generate multiple pseudo-random gradient evaluation vectors, and evaluate the fitness of the local gradient vector and these evaluation vectors, and submit the fitness to the server; after receiving the global fitness from the server, calculate the global gradient and update the model parameters. Users may be disconnected.

[0037] Server: There is a server whose functions are: initialize global model parameters and random number seeds and send them to each user; after receiving the local fitness of each user, aggregate them into global fitness. At the same time, use the shared random number seed and the global gradient vector of the previous k rounds to generate multiple pseudo-random gradient evaluation vectors, and use the global fitness calculation to generate the global gradient and update the global model parameters; send the global fitness to each user.

[0038] Initialization step: The server initializes the global model parameters and random number seeds and sends them to each user. After the initialization is completed, the global model is continuously updated by looping the following steps, which are described in detail as follows:

[0039] Step (i), the user uses local data training to obtain a local gradient vector, and then uses a random number seed and the global gradient vector of the previous k rounds to divide the entire real number space into a principal component space and its orthogonal complement space; extracts pseudo-random evaluation vectors in the principal space and the orthogonal complement space according to a certain ratio, and calculates the fitness value of each evaluation vector to the local gradient vector, and the user sends the local fitness value to the server;

[0040] Step (ii), after receiving the local fitness value from each user, the server aggregates it into a global fitness value and sends the global fitness value to each user;

[0041] Step (iii): the user and the server calculate the global gradient vector of this round according to the global fitness value and the pseudo-random evaluation vector, and update the global model.

[0042] In this embodiment, after the user completes this round of training locally, a self-guided pseudo-random gradient evaluation vector group is generated using a shared random number seed and the global gradient of the previous k rounds, and the local gradient vector obtained in this round of training is converted into a fitness value for the evaluation vector group. This greatly reduces the communication overhead between the user and the server, and retains more information than the unguided gradient evaluation vector group, which can ensure faster convergence of the global model training.

[0043] As a preferred implementation, the initialization step specifically includes:

[0044] The server selects the appropriate machine learning / deep learning model according to specific needs and initializes the global model parameters θ 0 , select a random number seed s, and set {θ 0 ,s} is sent to each user.

[0045] As a preferred implementation, step (1) specifically includes:

[0046] In round t, user u uses local training data to train the local model Get the local gradient vector

[0047] If t≤k, user u uses a random number seed to generate a normal distribution The pseudo-random gradient evaluation vector set If t>k, the user performs PCA decomposition on the first k global gradient vectors and takes the first d principal components to form the main space Take the remaining principal components to form an orthogonal complementary space Where d can be determined by taking the logarithm log2() of the dimension of the gradient vector; and sampling in these two subspaces according to the proportion α:

[0048] Where i = 1, ..., m;

[0049] Among them, the eigenvectors (principal components) corresponding to the first d eigenvalues ​​are taken according to the size of the eigenvalue.

[0050] In practical applications, in order to balance the bootstrap sampling and the extended sampling, the probability α is usually set to 0.5 or 0.3; secondly, in practical applications, the gradient vector usually has a dimension of more than hundreds of thousands, and it is unrealistic for ordinary users to store a matrix of hundreds of thousands by hundreds of thousands of dimensions. Therefore, the principal components from d+1 to 2d can be taken to form a pseudo-orthogonal complementary space.

[0051] User u calculates the fitness of each pseudo-random gradient vector and sends it to the server:

[0052]

[0053] in, is the pseudo-random gradient evaluation vector, represents the local gradient vector obtained by user u in the tth round of training, represents the local fitness value of the i-th pseudo-random gradient vector, and m is the number of pseudo-random gradient vectors.

[0054] As a preferred implementation, step (ii) specifically includes:

[0055] The server receives the fitness value vector of each user After that, calculate the global fitness: where w u is the weight of user u;

[0056] The server will set the global fitness value Each user.

[0057] As a preferred implementation, step (iii) specifically includes:

[0058] The server uses the random number seed to generate the same pseudo-random evaluation vector set as the user according to the method described in step (1)

[0059] The server and the user independently calculate the global gradient of this round and update the model parameters:

[0060]

[0061] θ t+1 =θ t -η*g t ;

[0062] The user saves the global gradient g t For future training rounds; where g t is the global gradient obtained by aggregation in the tth round, η is the learning rate, θ t+1 is the t+1th round global model, m is the number of pseudo-random gradient evaluation vectors, A vector of pseudo-random gradient evaluations.

[0063] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A communication optimization federated learning method based on a self-guided evolutionary strategy, characterized in that: The following steps are involved: Step 1: The server initializes global model parameters and random number seeds and sends them to each user; Step 2: The user performs training through local data to obtain a local gradient vector; a subspace is divided from the real number space based on the global gradient vector of a preset number of rounds and the random number seed, and a number of pseudo-random evaluation vectors are extracted from the subspace; Calculate the local fitness value of each pseudo-random vector to the local gradient vector, and feed it back to the server; Step 3: The server calculates the global fitness value based on the local fitness value fed back by each user and sends it to each user; Step 4: The user and the server calculate the global gradient vector of this round based on the global fitness value and the pseudo-random evaluation vector, and update the global model; Repeat steps 2 to 4 to continuously update the global model.

2. The communication optimization federated learning method based on self-guided evolutionary strategy according to claim 1 is characterized in that: The process of obtaining the pseudo-random evaluation vector in step 2 includes: The user trains the local model through local data to obtain a local gradient vector; if the number of training rounds of the local model is less than or equal to the preset number of rounds, a pseudo-random gradient evaluation vector is generated based on the random number seed; if the number of training rounds of the local model is greater than the preset number of rounds, a subspace is divided from the real number space based on the global gradient vector and the random number seed, and a number of pseudo-random evaluation vectors are extracted from the subspace.

3. The communication optimization federated learning method based on self-guided evolution strategy according to claim 2 is characterized in that: In the step 2, the pseudo-random gradient evaluation vector generated based on the random number seed obeys a normal distribution.

4. The communication optimization federated learning method based on self-guided evolution strategy according to claim 2 is characterized in that: In the step 2, a subspace is divided from the real number space based on the global gradient vector and the random number seed, and a process of extracting a plurality of pseudo-random evaluation vectors from the subspace includes: The matrix composed of the global gradient vectors of the first few rounds of a preset number of rounds is subjected to matrix decomposition and divided into two parts: a main space and an orthogonal complement space. The main space and the orthogonal complement space are sampled according to a preset sampling ratio to obtain a number of pseudo-random evaluation vectors.

5. The communication optimization federated learning method based on self-guided evolution strategy according to claim 1 is characterized in that: In the step 2, the formula for calculating the local fitness value of each pseudo-random vector to the local gradient vector is as follows: in, is the pseudo-random gradient evaluation vector, represents the local gradient vector obtained by user u in the tth round of training, represents the local fitness value of the i-th pseudo-random gradient vector, and m is the number of pseudo-random gradient vectors.

6. The communication optimization federated learning method based on self-guided evolution strategy according to claim 5 is characterized in that: In step 3, the formula for calculating the global fitness value is as follows: Among them, w u is the weight of user u, U is the total number of users, f i t is the global fitness value.

7. The communication optimization federated learning method based on self-guided evolution strategy according to claim 6 is characterized in that: In step 4, the server generates a pseudo-random evaluation vector set using the method of step 2. The server and the user respectively calculate the global gradient of this round and update the model parameters. The user saves the global gradient for the next round of training.

8. The communication optimization federated learning method based on self-guided evolution strategy according to claim 7 is characterized in that: The formula for calculating the global gradient and model parameters of this round is as follows: i t+1 =θ t -h*g t Among them, g t is the global gradient obtained by aggregation in the tth round, η is the learning rate, θ t+1 is the t+1th round global model, m is the number of pseudo-random gradient evaluation vectors, A vector of pseudo-random gradient evaluations.

Citation Information

Patent Citations

  • Federal learning optimization method and device, electronic equipment and storage medium

    CN115618960A

  • Federal learning method and system

    CN116861239A

  • Federal learning convergence acceleration optimization method and system based on corrected gradient descent

    CN117436513A

  • Communication efficient federal learning method based on gradient cosine similarity and differential privacy

    CN119089980A

  • Methods, apparatuses, and systems for multi-party collaborative model updating for privacy protection

    US20240112091A1