A method for realizing efficient multi-server federated learning through over-the-air computation
By using in-phase digital aerial computing technology to achieve gradient superposition in multi-server federated learning, the problems of slow convergence speed and high cost in traditional methods are solved, and efficient user activity and low-cost learning results are achieved.
Patent Information
- Application Number
- CN202310943612.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-28
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2043-07-28
AI Technical Summary
Traditional over-the-air computing technology has a slow convergence speed in federated learning, and each user needs to have an RF linear power amplifier, resulting in high system cost and low user activity.
The method employs in-phase digital over-the-air computing, which utilizes multi-server joint learning and gradient vector superposition via in-phase OTA channels to achieve unbiased estimation and eliminates the need for high-fidelity RF linear power amplifiers. It also employs random transmission methods and time-division duplex technology to reduce uplink bandwidth requirements.
It achieves efficient multi-server federated learning, reduces system costs, improves user activity and learning efficiency, and avoids the problems of bandwidth expansion and high costs in traditional methods.
Smart Images

Figure CN116962406B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of wireless communication, and particularly relates to a method for realizing efficient federated learning of multiple servers through over-the-air computation. BACKGROUND
[0002] The development of Internet of Things (IoT) leads to an explosive growth of edge devices, which generate a large amount of data with immeasurable potential value. The federated learning method using over-the-air computation technology can complete data processing while avoiding privacy risks and reducing the demand for the number of uplink channels. However, the traditional over-the-air computation technology has a slow convergence speed, and each user needs to have a radio frequency linear power amplifier in the transmission system, which greatly increases the system cost.
[0003] Federated learning allows the use of distributed data sets to train machine learning models, which can be used to protect data privacy and reduce transmission burden. Meanwhile, using the superposition characteristics of wireless (or wireless) communication channels, uplink data superposition can be realized through over-the-air (OTA) computation technology. SUMMARY
[0004] Therefore, the purpose of the application is to provide a method for realizing efficient federated learning of multiple servers using in-phase over-the-air computation. Specifically, a digital over-the-air (DOTA) computation method is applied to the transmission of uplink signals of multiple server federated learning. Combining the characteristics of federated learning that allows the use of distributed data sets to train machine learning models, user privacy protection and reduction of uplink channel bandwidth can be realized.
[0005] To achieve the above purpose, the application provides the following technical solutions:
[0006] A method for realizing efficient federated learning of multiple servers through over-the-air computation, comprising the following steps:
[0007] S1: Each user user respectively holds a part of the data set dataset and can receive the federated learning model and the corresponding parameter quantity from different server access points AP;
[0008] S2: Each user respectively uses different federated learning models to perform optimization operation on the local data set, thereby updating the state parameter vector of each federated learning model;
[0009] S3: The user converts the parameter vector of the updated different federated learning models into a gradient vector, and sends the gradient vector to the corresponding AP;
[0010] S4: Through over-the-air computation technology, the transmission signals of all users can realize gradient superposition and be converted into M gradient superposition vectors, where M is the number of federated learning models;
[0011] S5: M co-signal is superimposed in a channel by co-signal OTA aggregation technology;
[0012] S6: Each server makes unbiased estimation of the corresponding sending signal according to the channel characteristics and the received signal, and updates the federal learning model parameters according to the result, and enters a new round of iteration.
[0013] Further, M servers are associated with their own APs, and each server has a federal learning model, and the cost function of server m is represented by C m , where K represents the number of users, D k represents the user k of the data set, the size of the data set of each user |Dk| is the same, and the optimal parameter vector of the federal learning model m:
[0014]
[0015] w m ∈R L , where R represents a vector space, L is the length of the parameter vector, and K >> M, that is, the number of users is much larger than the number of APs.
[0016] Further, for federal learning, time is divided into different time slots, and each time slot is divided into two sub-slots: the first sub-slot transmits from the AP to the user, and the second sub-slot transmits from the user to the AP in a time division duplex manner;
[0017] In the first sub-slot, the AP m transmits its orthogonal pilot signal and the updated parameter vector, which is represented by
[0018] w(t)m at time slot t, the following gradient vector of the federal learning model m is obtained at the user k:
[0019]
[0020] where represents the gradient vector of the function f(x), which will be transmitted in the second sub-slot;
[0021] If the AP m can receive the gradient vector of the user without any distortion and error, the parameter vector of the model m is updated as follows:
[0022]
[0023] where is the aggregation vector, and μ is the step size or learning rate.
[0024] Further, h k,m (t) represents the channel coefficient between user k and server m at time slot t; the channel coefficient is a symmetric Rayleigh fading, that is:
[0025]
[0026] Then, user k estimates h k,m (t) at AP m in the second sub-slot is denoted as:
[0027]
[0028] where x k (t) is the signal sequence transmitted by user k, n m (t) ~ CN(0, N0I) is the additive white Gaussian noise at AP m.
[0029] Further, the uplink channel is divided into M orthogonal channels, each server has a dedicated uplink channel, and the upload is performed through OTA; user k transmits the local gradient vector of federated learning model m through channel m as follows:
[0030]
[0031] where, And Pm(k) denotes the transmit power of user k to channel m; at AP m, the signal received through channel m becomes:
[0032]
[0033] Further, the phase of the channel fading coefficient is independent and uniformly distributed, then the correlation of the channel coefficient is zero, that is:
[0034]
[0035] can be used to simplify the M orthogonal channels into one channel;
[0036] Each user transmits the superposition of M local gradient vectors through a shared uplink channel, where the signal to be transmitted by user k becomes:
[0037]
[0038] Using the random transmission method instead of the codebook, the transmission probability of the first bit vector of the signal transmitted by user k to server m is as follows:
[0039] p k;m,l = min{1, γ|[g k;m ] l |}
[0040] Where γ>0 is a control parameter;
[0041] The first element of the signal sequence transmitted by user k becomes
[0042]
[0043] where U k;m,l is a uniform random variable; let z k;m = [z k;m,l …x k;m,L ] T If γ is small enough, then
[0044]
[0045] z k;m is an unbiased estimate of g k;m The resulting method of generating z k;m is called soft-gain random transmission, which is regarded as a kind of random quantization;
[0046] Let κ m denote the set of users who choose model m, denote the union of users in κ m who can successfully transmit signals, denote the complement of κ , then the AP m received signal is denoted as:
[0047]
[0048]
[0049] If each user k uniformly randomly chooses a federated learning model, for a given federated learning model m, the probability of user k choosing the federated learning model m is 1 / M, r m The average value of r
[0050]
[0051] r m (t) is an unbiased estimate of g
[0052] If Km is randomly selected at each iteration, the update result obtained by using the federated learning model selection will become an SGD method, and the update rule of the SGD method at the AP m becomes:
[0053]
[0054] The uplink bandwidth of each iteration is reduced from KM to M, and then to 1;
[0055] z k;m,l ∈{-1, 0, +1}, transmission power is independent of g k;m ; if set all transmission power is same, P k = P max , i.e.
[0056] The beneficial effects of the present application are that the present application realizes a same-phase air computing method with efficient communication upload in multi-server federated learning, which does not need bandwidth expansion of the uplink channel and can realize unbiased estimation of aggregation on each server. At the same time, the user end transmission power fluctuation range is small, and high-fidelity radio frequency linear power amplifiers are not needed, which greatly reduces the cost. In addition, the proposed digital air computing technology can overcome the shortcomings of large transmission power fluctuation range and low number of active users of a single model of analog air computing. Simulation results show that with the increase of the number of participating users, the digital air computing method is superior to the analog air computing in terms of federated learning learning efficiency and user activity.
[0057] Other advantages, objects, and features of the present application will be set forth in the following specification and will become apparent to those skilled in the art from the practice of the present application. The objects and other advantages of the present application will be realized and attained by the structure particularly pointed out in the written description and claims thereof as well as the appended drawings. BRIEF DESCRIPTION OF DRAWINGS
[0058] In order to make the purposes, technical solutions and beneficial effects of the present application clearer, the present application provides the following drawings for illustration:
[0059] Figure 1 The system model diagram of the present application;
[0060] Figure 2 The uplink bandwidth change diagram of each iteration. DETAILED DESCRIPTION
[0061] The purpose of the present application is to use the same-phase digital air computing technology to realize channel superposition for the large amount of uplink data channel demand brought by multi-server federated learning, and the specific system model is as follows: Figure 1As shown, firstly, each user respectively holds a part of the dataset, and can receive the federated learning model and the corresponding parameter quantity from different server access points (APs); secondly, each user respectively adopts different models to perform optimization operation on the local dataset, so as to update the state parameter vector of each model; then, the user converts the parameter vector of the updated different models into a gradient vector, and sends the gradient vector to the corresponding AP; thereafter, through the over-the-air computation technology, the sending signals of all users can realize gradient superposition, and are converted into M gradient superposition vectors of the model, that is, the access point; then, through the same direction OTA aggregation technology, the M same direction signals are superposed in one channel; finally, each server can perform unbiased estimation on the corresponding sending signal according to the channel characteristics and the received signal, and update the model parameter according to the result, and enter a new round of iteration. Through the use of the digital over-the-air computation technology, the present application avoids the increase of the uplink channel number with the increase of the number M of access points. In addition, the present application also avoids the problems of the need for expensive radio frequency (RF) linear power amplifiers and the low user activity in the initial iteration operation in the existing analog over-the-air computation technology.
[0062] 1) Multi-server transmission model based on over-the-air computation
[0063] Suppose there are M servers associated with their own APs. Each server has a model, and the cost function of server m is denoted by Cm (thus, as mentioned before, the model, the server and the aforementioned AP are interchangeable). For each model, the distributed dataset at the user can be used for training. Let K denote the number of users, and Dk denote the dataset of user k. In addition, for simplicity, the size of each user's dataset Dk is the same. Then, the optimal parameter vector of model m is given by:
[0064]
[0065] Suppose w m ∈R L , where R represents a vector space, and L is the length of the parameter vector. It is assumed that K >> M, that is, the number of users is much larger than the number of aps.
[0066] For federated learning, time is divided into different slots, and each slot is further divided into two sub-slots: the first sub-slot transmits from the ap to the user, and the second sub-slot transmits from the user to the ap in a time division duplex manner (TDD). In the first sub-slot, each AP, such as AP m, transmits its orthogonal pilot signal and the updated parameter vector, denoted by w(t)m, in slot t. Then, at user k, the following gradient vector of model m can be obtained:
[0067]
[0068] Here denotes the gradient vector of the function f(x) which will be transmitted in the second sub-slot.
[0069] If the AP m can receive the gradient vector of the user without any distortion and error, the parameter vector of the model m can be updated as follows:
[0070]
[0071] where is the aggregated vector and μ is the step size or learning rate.
[0072] 2) Co-simulated OTA transmission model
[0073] Assume that h k;m (t) denotes the channel coefficient between user k and server (AP) m at time slot t. For the channel coefficient, we assume a symmetric Rayleigh fading, i.e.,
[0074]
[0075] Then, user k can estimate h k;m (t) when ap transmits their pilot in one sub-slot. So the received signal at AP m between the second sub-slot can be represented as:
[0076]
[0077] where x k (t) is the signal sequence transmitted by user k, n m (t) ~ CN(0, N0I) is the background noise at AP m.
[0078] Assume that the uplink channel is divided into M orthogonal channels, so that each AP / server has a dedicated uplink channel, which can be uploaded through OTA. Then, user k transmits the local gradient vector of model m through channel m as follows:
[0079]
[0080] where, and denotes the transmit power of user k to channel m. At AP m, the received signal through channel m becomes
[0081]
[0082] Assume that the phase of the channel fading coefficient is independent and uniformly distributed, then the correlation of the channel coefficient is zero, i.e.:
[0083]
[0084] This can be used to simplify M forward traffic channels into a single channel. In this case, each user transmits a superposition of M local gradient vectors through a shared uplink channel. The resulting method is called in-phase aggregation, where the signal to be transmitted by user k is:
[0085]
[0086] Using a random transmission method instead of a codebook, the transmission probability of the first bit vector of the signal transmitted by user k to server m is as follows:
[0087] p k;m,l =min{1,γ|[g k;m ] l |} (2.7)
[0088] Where γ > 0 is the control parameter.
[0089] The first element of the signal sequence transmitted by user k becomes
[0090]
[0091] Among them, U k;m,l ~U(0,1) is a uniform random variable. Let z k;m =[z k;m,1 …x k;m,L ] T If γ is small enough, we can prove
[0092]
[0093] This means z k;m It is g k;m The unbiased estimate of the (proportion). The resulting generated z k;m The method is called soft-gain random transfer, which can also be seen as a form of random quantization.
[0094] If let Let κ m Let m represent the set of users who choose model m. Indicates κ m A collection of users who were able to successfully send signals. express The complement of APm can then be expressed as:
[0095]
[0096] Since the average value of the second term (RHS) on the right side of equation (2.10) is 0 as shown in equation (2.5), and n m If the average value of (t) is also zero, then
[0097]
[0098] Thus, if each user k chooses a model uniformly at random, the probability that user k chooses model m for a given model m is 1 / M. It follows that the average value of r m becomes
[0099]
[0100] Thus, r m (t) in equation (2.10) can become a (proportional) unbiased estimate of .
[0101] If Km is chosen randomly at each iteration, the update resulting from the (random) model selection becomes the SGD method, whose update rule at APm becomes:
[0102]
[0103] It can be seen from Figure 2 that the uplink bandwidth at each iteration has been reduced from KM to M, and then to 1.
[0104] At the same time, since z k;m,l ∈ {-1, 0, +1}, the transmission power is independent of g k;m . If we set all transmission powers to be the same, P k = P max , i.e. This means that the transmission power is fixed at the time of transmission. Therefore, the transmit power is determined only by the CSI, which means that this method eliminates the disadvantage of analog OTAs that require radio frequency linear power amplifiers. In addition, because the maximum transmission power of the signal is independent of the vector value, this also means that this method eliminates the disadvantage of analog OTAs that cause the user activity to decrease when the vector value is too large at the beginning of the iteration.
[0105] Finally, it should be pointed out that the above preferred embodiments are only used to illustrate the technical solutions of the present application and not to limit it. Although the present application has been described in detail through the above preferred embodiments, those skilled in the art should understand that various modifications can be made in form and details without departing from the scope defined by the claims of the present application.
Claims
1. A method for implementing multi-server efficient federated learning through over-the-air computation, the method comprising: Comprise the following steps: S1: Each user user respectively holds a part of the data set dataset, and can receive the federated learning model and the corresponding parameter quantity from different server access points AP; S2: Each user respectively adopts different federated learning models to optimize the operation on the local data set, so as to update the state parameter vector of each federated learning model; S3: The user converts the parameter vector of the updated different federated learning models into a gradient vector, and sends the gradient vector to the corresponding AP; S4: Through the over-the-air computing technology, the sending signals of all users can realize gradient superposition and are converted into M gradient superposition vectors, M being the number of federated learning models; S5: Through the same direction OTA aggregation technology, M same direction signals are superimposed in one channel; S6: Each server makes an unbiased estimation on the corresponding sending signal according to the channel characteristics and the received signal, and updates the federated learning model parameters according to the result, and enters a new round of iteration; M servers are associated with their own APs, each server has a federated learning model, the cost function of server m is denoted by C m , where K denotes the number of users, D k denotes the data set of user k, the size of each user's data set |Dk| is the same, and the optimal parameter vector of federated learning model m: w m ∈R L where R represents a vector space, L is the length of the parameter vector, and K » M, i.e., the number of users is much larger than the number of APs. For federated learning, time is divided into different time slots, and each time slot is divided into two sub-slots: the first sub-slot transmits from the AP to the user, and the second sub-slot transmits from the user to the AP in the form of time division duplex; In the first sub-slot, APm transmits its orthogonal pilot signal and updated parameter vector, denoted as w(t)m in time slot t, and the following gradient vector of federated learning model m is obtained at user k: If APm can receive the gradient vector of the user without any distortion and error, the parameter vector of model m is updated as follows: wherein denotes the gradient vector of the function f(x) which will be transmitted in the second sub-slot; The uplink channel is divided into M orthogonal channels, each server has a dedicated uplink channel, and the upload is performed through OTA; The user k transmits the local gradient vector of the federated learning model m through channel m as follows: wherein is the aggregation vector, μ is the step size or learning rate; h k;m (t) denotes the channel coefficient between user k and server m at time slot t; the channel coefficients are symmetric Rayleigh fading, i.e.: Then, user k estimates h when the AP transmits their pilot in a sub-slot k;m (t); the signal received at the AP m between the second sub-slots is represented as: where x k (t) is the signal sequence transmitted by user k, n m (t) ~ CN(0, N0I) is the background noise at the APm; If the phase of the channel fading coefficient is independent and uniformly distributed, the correlation of the channel coefficient is zero, that is: wherein And denotes the transmit power of user k to channel m; at the AP m, the signal received through channel m becomes: It can be used to simplify M orthogonal channels into one channel; Each user transmits the superposition of M local gradient vectors through a shared uplink channel, where the signal to be transmitted by user k becomes: Adopting a random transmission method instead of a codebook, the transmission probability of the lth bit vector of the signal transmitted by user k to server m is as follows: Where γ>0 is a control parameter; p k;m,l = min{1, γ | [g k;m ] l |} The lth element of the signal sequence transmitted by user k becomes If Km is randomly selected in each iteration, the update result obtained by using the federated learning model selection will become the SGD method, and the update rule of the SGD method at APm becomes: where U k;m,l ~ U(0,1) is a uniform random variable; let z k;m = [z k;m,1 ...x k;m,L ] T if γ is small enough, then z k;m is an unbiased estimate of g k;m , the resulting method of generating z k;m is called soft-gain random transmission, which is a kind of random quantization; Let κ m Let m represent the set of users who choose model m. Indicates κ m A collection of users who were able to successfully send signals. express The complement of APm is then represented as: If each user k uniformly randomly selects one federated learning model, for a given federated learning model m, the probability that user k selects federated learning model m is 1 / M, r m The average value is: r m (t) is an unbiased estimate of an unbiased estimate of The uplink bandwidth of each iteration is reduced from KM to M, and then to 1; z k;m,l ∈ {-1,0, +1}, transmission power is independent of g k;m ; if set all transmission power is the same, all users k P k = P max , i.e.