Device Transmission Power Configuration Method for Accelerating the Convergence Rate of Distributed Learning Frameworks
By amplifying the distortion of the AirComp gradient aggregate signal in a distributed learning framework and reducing the transmission power of the equipment, the problem of excessive transmission power of the equipment in the prior art is solved, and the effect of accelerating the convergence speed and saving energy consumption is achieved.
Patent Information
- Application Number
- CN202510354094.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-25
AI Technical Summary
The existing equipment transmission power configuration method based on distortion suppression criteria causes the device transmission power to be too high in the distributed learning framework, and cannot effectively utilize the convergence acceleration effect brought by the distortion of the gradient aggregated signal based on AirComp, and does not conform to the actual situation of insufficient equipment transmission power and insufficient energy budget in wireless networks.
By amplifying the distortion of the gradient aggregated signal based on AirComp, the device transmission power is reduced, and the distortion is used to accelerate the convergence of the distributed learning framework. The specific method includes reducing the device power proportional factor and synchronously reducing the base station normalization factor while satisfying that the ratio of the device power proportional factor and the base station normalization factor is greater than 1, so as to increase the mean square error, thereby amplifying the over-air calculation distortion.
It has achieved the acceleration of the convergence speed of the distributed learning framework, reduced the transmission power and communication energy consumption of the equipment, and met the actual needs of the equipment energy budget in the wireless network.
Smart Images

Figure CN119892180B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of communication technologies, and particularly relates to a method for configuring the transmission power of a device to accelerate the convergence speed of a distributed learning framework. Background Art
[0002] A distributed learning framework in a wireless network includes a base station and multiple devices. The base station has a global model, and each device has a local dataset and a copy of the global model. The process of the distributed learning framework is divided into multiple rounds. In each round, each device uses the data in the local dataset and the copy of the global model to calculate a local gradient. Then, each device uploads the local gradient to the base station for aggregation to generate a global gradient. Next, the base station uses the aggregated global gradient to update the global model. Finally, the base station sends the updated global model to each device to update the copy of the global model of each device. The above process is repeated for multiple rounds until the global model of the base station converges, that is, the performance of the global model converges to a certain predetermined standard or reaches the maximum allowed number of rounds.
[0003] Multiple transmitters using over-the-air computation (AirComp) technology concurrently transmit signals on the same time-frequency resources, and utilize the superposition characteristics of the wireless channel to complete signal-level calculations, enabling the receiver to directly receive the calculation results of the signals transmitted by each transmitter. In a distributed learning framework, this technology is often used to aggregate local gradients from each device at the base station.
[0004] The gradient descent algorithm is a basic algorithm for training an artificial intelligence (AI) model. In this algorithm, first, a gradient is calculated using data samples and the AI model. Subsequently, the parameters of the AI model to be updated are subtracted by the product of the learning rate and the gradient to obtain the updated AI model. In a distributed learning framework, this algorithm is used to update the global model at the base station.
[0005] The closest prior art to the present invention is the device transmission power configuration method based on the distortion suppression criterion (i.e., aiming to minimize the mean square error (MSE) of the AirComp aggregated gradient signal) (refer to Paper 1 (How to Balance Accuracy and Integrity for Reconfigurable Intelligent Surface-Aided Over-the-Air Federated Learning, authors: Zheng Jingheng; Tian Hui; Ni Wanli; Ni Wei; Zhang Ping, IEEE Transactions on Wireless Communications (Volume: 21, Issue: 12, December 12, 2022), Pages: 10964-10980, Publication Date: July 13, 2022), URL: https: / / ieeexplore.ieee.org / document / 9829190) and Paper 2 (Federated Learning in Multi-RIS-Aided Systems, authors: Ni Wanli; Liu Yuanwei; Yang Zhaohui; Tian Hui; Shen Xuemin, IEEE Internet of Things Journal (Volume: 9, Issue: 9, June 15, 2022), Pages: 9608-9624, Publication Date: November 24, 2021), URL: https: / / ieeexplore.ieee.org / document / 9626135)). In the distributed learning framework applying AirComp, this method takes the local gradient aggregation of AirComp with distortion suppression as the criterion and believes that reducing the distortion of the aggregated local gradient can ensure the convergence of the distributed learning framework. Specifically, most of the existing device transmission power configuration methods based on the distortion suppression criterion require each device to use a high transmission power to offset the influence of the differences in the uplink channels of each device and suppress the received noise, so as to suppress the distortion of the gradient aggregation signal based on AirComp, thereby minimizing the distortion and error of the aggregated gradient signal received at the base station. By analyzing this method, it is found that to suppress distortion, it is necessary to minimize the MSE. This means that the value of the normalization factor of the base station needs to be increased as much as possible, and according to the configuration formula of the device transmission power, the increase of the base station normalization factor will lead to an increase in the device transmission power. This shows that the existing device transmission power configuration method based on the distortion suppression criterion requires the device to increase the transmission power to minimize the MSE.
[0006] Some existing research works have shown that the local gradients in the distributed learning framework are inherently resilient to damage for model training and do not need to be transmitted overly precisely. There is also research indicating that when using the gradient descent algorithm, imposing specific artificial interference on the gradients can actually accelerate the convergence rate of the AI model training process. Based on this finding, since the distortion of the gradient aggregation signal based on AirComp is also an interference to the global gradient, it can be inferred that artificially amplifying or leveraging the distortion of the gradient aggregation signal based on AirComp can help accelerate the convergence rate of the distributed learning framework. Therefore, a major drawback of the existing device transmission power configuration method based on the distortion suppression criterion is that it is too conservative, that is, it does not fully utilize the convergence acceleration effect brought by the distortion of the gradient aggregation signal based on AirComp. At the same time, the existing device transmission power configuration method based on the distortion suppression criterion causes the problem of excessive device transmission power, which does not conform to the actual situation of insufficient device transmission power and energy budget in wireless networks and is difficult to implement for applications. Summary of the Invention
[0007] The present invention is directed to a distributed learning framework that utilizes over-the-air computation (AirComp) technology for local gradient aggregation, aiming to provide a device transmission power configuration method for accelerating the convergence rate of the distributed learning framework to solve the technical problems existing in the existing device transmission power configuration method based on the distortion suppression criterion.
[0008] The present invention amplifies the distortion of the gradient aggregation signal based on AirComp to accelerate the convergence rate of the distributed learning framework. By reducing the transmission power of each device in the distributed learning framework, it is possible to save device transmission power and reduce the communication energy consumption of the distributed learning framework while amplifying the over-the-air computation distortion.
[0009] The technical solutions adopted by the present invention to solve the technical problems are as follows:
[0010] A device transmission power configuration method for accelerating the convergence rate of a distributed learning framework provided by the present invention includes the following steps:
[0011] Step S1: Establish a distributed learning framework;
[0012] The distributed learning framework includes: K devices and 1 base station;
[0013] Step S2: When entering each training round, each device obtains its channel gain vector to the base station, the receiving beamforming configuration scheme of the base station, and the local gradient to be transmitted;
[0014] Step S3: Each device sets its own transmission power with the criterion of amplifying the air-computation distortion, and concurrently transmits the local gradients to the base station at the set transmission power on the same time-frequency resources.
[0015] Step S4: The base station aggregates the local gradients of each device based on the air-computation technology to obtain the global gradient of the current training round, and updates the global model using the global gradient.
[0016] Further, the calculation formula for the local gradient is:
[0017]
[0018] where, represents the local gradient of device k in the t-th training round, represents the gradient operator, represents the n-th data sample in the local dataset of device k corresponding to the loss function, represents the global model updated in the previous training round broadcast by the base station in the t-th training round, represents the local dataset of device k in the t-th training round.
[0019] Further, the calculation formula for the transmission power is:
[0020]
[0021] where, represents the transmission power of device k, represents the receive beamforming configuration scheme of the base station, represents the channel gain vector from device k to the base station, represents a non-negative device power ratio factor, represents the conjugate transpose.
[0022] Further, the air-computation distortion is measured by the mean square error of the global gradient; the mean square error of the air-computation is calculated as:
[0023]
[0024] where, represents a non-negative device power ratio factor, represents the base station normalization factor, represents the noise intensity.
[0025] Further, the criterion for amplifying the air-computation distortion is as follows: on the premise that the ratio of the device power scaling factor to the base station normalization factor is greater than 1, reduce the device power scaling factor and simultaneously reduce the base station normalization factor, so that the MSE increases, thereby amplifying the air-computation distortion.
[0026] Further, the calculation formula for the global gradient is:
[0027]
[0028] where, represents the global gradient aggregated by the base station in the t-th training round, represents the received noise of the base station, represents the base station normalization factor, represents the received beamforming configuration scheme of the base station, represents the channel gain vector from device k to the base station, represents the conjugate transpose, represents the transmit power of device k, represents the local gradient of device k in the t-th training round.
[0029] Further, based on the configuration scheme of the transmit power of device k, the global gradient aggregated by the base station in the t-th training round is further expressed as:
[0030]
[0031] where, represents the non-negative device power scaling factor.
[0032] Further, the base station updates the global model using the gradient descent formula:
[0033]
[0034] where, represents the global model updated in the previous training round broadcast by the base station in the t-th training round, represents the global model updated in the previous training round broadcast by the base station in the (t - 1)-th training round, represents the learning rate, represents the global gradient aggregated by the base station in the t-th training round.
[0035] Further, the convergence rate of the distributed learning framework is defined as the expectation of the decrease value of the global loss function between two adjacent training rounds; the mathematical expression of the convergence rate of the distributed learning framework is:
[0036]
[0037] Among them, represents the global loss function of the t-th training round, represents the global loss function of the (t - 1)-th training round, represents the global model updated in the previous training round broadcast by the base station in the t-th training round, represents the global model updated in the previous training round broadcast by the base station in the (t - 1)-th training round.
[0038] Furthermore, the mathematical expression of the global loss function of the t-th training round is:
[0039]
[0040] Among them, represents the loss function corresponding to the n-th data sample in the local dataset of device k and D represents the number of data samples in the local dataset.
[0041] The beneficial effects of the present invention are:
[0042] 1. Compared with the existing device transmission power configuration method based on the distortion suppression criterion, the device transmission power configuration method proposed by the present invention artificially amplifies and utilizes the distortion of the gradient aggregation signal based on AirComp to accelerate the convergence speed of the distributed learning framework, and reduces the transmission power of the device, saving the communication energy consumption of the device.
[0043] 2. The device transmission power configuration method proposed by the present invention actively reduces the transmission power of the device to amplify the distortion of the gradient aggregation signal based on AirComp, which can save the communication energy consumption of the device. Description of the Drawings
[0044] Figure 1 is a schematic diagram of a distributed learning framework in a wireless network.
[0045] Figure 2 is a schematic diagram of the local gradient signals input by each device to the channel in the distributed learning framework in a wireless network, where the abscissa is time and the ordinate is the signal amplitude.
[0046] Figure 3 is a comparison schematic diagram of the global gradient signal aggregated through air computing output by the channel and the local gradient signals input by each device to the channel in the distributed learning framework in a wireless network. Among them, the solid line represents the global gradient signal aggregated through air computing output by the channel, and the dashed line represents the local gradient signals input by each device to the channel. The abscissa is time and the ordinate is the signal amplitude.
[0047] Figure 4 Flowchart of a method for configuring the transmission power of a device to accelerate the convergence speed of a distributed learning framework provided by the present invention. Detailed implementation manners
[0048] The present invention will be further described in detail below with reference to the accompanying drawings.
[0049] As Figure 4 shown, a method for configuring the transmission power of a device to accelerate the convergence speed of a distributed learning framework provided by the present invention has the following specific implementation process:
[0050] Step S1: Establish a distributed learning framework;
[0051] As Figure 1 shown, the distributed learning framework in the wireless network mainly includes: K devices (1,..., k,..., K) and 1 base station. Among them, at the current t-th training round, each device respectively has a local dataset ,... , and the global model updated in the previous training round broadcast by the base station ; each local dataset includes data samples, which are used to calculate the local gradients ,... of each device.
[0052] Among them, a training round includes a local training period and a communication round. Calculating the local gradients of each device is completed within the local training period, while uploading the local gradients of each device to the base station is completed within the communication round.
[0053] In the distributed learning framework, each device concurrently uploads its local gradient on the same time-frequency resource, and based on the air computing technology, the aggregation of local gradients is completed at the base station to generate a global gradient. Among them, the air computing technology utilizes the superposition characteristic of the wireless channel, and by enabling each device to concurrently transmit the local gradient on the same time-frequency resource, the local gradients of the devices are aggregated at the base station to generate a global gradient.
[0054] It can be understood that due to the influence of wireless channel fading and noise, the global gradient generated by aggregation based on air computing will be distorted. As Figure 2 and Figure 3 shown, the output signal of the wireless channel has a large distortion in amplitude relative to the input signal. In the traditional distributed learning framework that uses air computing technology to aggregate local gradients, a criterion for suppressing air computing distortion is used for system design. In the present invention, however, the air computing distortion is deliberately amplified and utilized to act on the globally aggregated gradient, thereby achieving the effect of accelerating the convergence speed of the distributed learning framework. This is the main feature that significantly differentiates the present invention from other inventions related to distributed learning frameworks.
[0055] Among them, the convergence rate of the distributed learning framework is defined as the expected value of the decrease in the global loss function between two adjacent training rounds. Specifically, the global loss function is defined as follows:
[0056]
[0057] Among them, represents the nth data sample in the local dataset of device k corresponding loss function; correspondingly, the convergence rate of the distributed learning framework can be defined as:
[0058]
[0059] Among them, represents the global loss function of the tth training round, represents the global loss function of the (t - 1)th training round, represents the global model updated in the previous training round broadcast by the base station in the tth training round; represents the global model updated in the previous training round broadcast by the base station in the (t - 1)th training round.
[0060] As Figure 1 shown, after the base station uses air computing technology to aggregate local gradients to generate global gradients and then updates the global model using the gradient descent formula. Among them, the gradient descent formula is specifically:
[0061]
[0062] Among them, represents the learning rate.
[0063] Step S2: When entering each training round, each device obtains its channel gain vector to the base station, the receiving beamforming configuration scheme of the base station, and the local gradient to be transmitted.
[0064] A training round of the distributed learning framework includes a local training period and a communication round. Among them, each device calculates the local gradient once using the local dataset within a local training period; subsequently, within a communication round, each device uploads the local gradient to the base station for aggregation to generate the global gradient.
[0065] Among them, the time lengths of the local training period and the communication round are related to the communication and computing capabilities of each device and need to be set according to the actual usage scenario. Regarding this, the present invention does not make specific limitations.
[0066] Among them, when calculating the local gradient using the local dataset, all data samples in the local dataset can be used, or only some data samples can be used. Specifically, the local gradient of device k is defined as:
[0067]
[0068] where represents the gradient operator, represents the local dataset of device k in the t-th training round.
[0069] The base station can inform each device of its channel gain vector to the base station and the receiving beamforming configuration scheme of the base station in various ways. For example, the base station can broadcast the information to each device, or use the polling method to send the information to each device. The present invention does not make specific limitations in this regard.
[0070] Step S3: Each device takes the amplification of the air computing distortion as the criterion, sets its own transmission power, and concurrently sends the local gradient to the base station using the set transmission power on the same time-frequency resources.
[0071] The present invention innovates the distortion suppression criterion generally followed in the model aggregation stage of the existing distributed learning framework into a distortion utilization criterion, and provides a new design criterion for the distributed learning framework.
[0072] Based on the obtained channel gain vector from each device to the base station and the receiving beamforming configuration scheme of the base station, each device configures the transmission power according to the criterion of amplifying the air computing distortion. Specifically, let represent the transmission power of device k, then the mathematical expression of the configuration scheme of the transmission power of device k is:
[0073]
[0074] where represents the receiving beamforming configuration scheme of the base station, represents the channel gain vector from device k to the base station, represents a non-negative device power ratio factor, represents the conjugate transpose.
[0075] The air computing distortion can be measured by the mean square error (MSE) of the global gradient. It can be understood that the larger the MSE, the larger the air computing distortion. Specifically, using the symbol to represent the mean square error of the air computing, then The specific calculation formula of is as follows:
[0076]
[0077] Among them, represents the base station normalization factor, represents the noise intensity.
[0078] In the present invention, on the one hand, the MSE increases with the increase of the ratio of the device power scaling factor to the base station normalization factor . Thus, it can be understood that the criterion for amplifying the air computing distortion is: on the premise of satisfying , reduce the device power scaling factor and synchronously reduce the base station normalization factor , so that the MSE increases, thereby amplifying the air computing distortion. On the other hand, on the premise of ensuring , reducing the device power scaling factor can reduce the transmission power of device k, thereby saving the device transmission power and the communication energy consumption of the distributed learning framework while amplifying the air computing distortion.
[0079] After obtaining the device transmission power configuration scheme, each device sets its transmission power according to the device transmission power configuration scheme and concurrently sends the local gradient to the base station on the preset same time-frequency resources.
[0080] Step S4: The base station aggregates the local gradients of each device based on the air computing technology to obtain the global gradient of the current training round, and finally updates the global model using the global gradient.
[0081] Due to the superposition characteristic of the wireless channel, the base station takes the aggregation result of the local gradients uploaded by each device as the global gradient. Specifically, the global gradient aggregated by the base station can be expressed as:
[0082]
[0083] Among them, represents the received noise of the base station.
[0084] Furthermore, based on the configuration scheme of the transmission power of device k, the global gradient can be further expressed as:
[0085]
[0086] Based on the above formula, it can be understood that the criterion for amplifying the air computing distortion makes the ratio of the device power scaling factor to the base station normalization factor increase, so that the amplitude of the global gradient increases, resulting in a larger distortion, and thus can accelerate the convergence speed of the distributed learning framework.
[0087] The base station updates the global model using the global gradient generated by aggregation. Specifically, the base station can update the global model using the following gradient descent formula:
[0088]
[0089] where represents the learning rate.
[0090] Subsequently, the base station broadcasts the updated global model to each device for the next training round of the distributed learning framework.
[0091] The present invention innovates the distortion suppression criterion generally followed in the model aggregation stage of the existing distributed learning framework into a distortion utilization criterion. At the same time, based on the inherent resilience of the local gradient to model training, the present invention artificially amplifies and utilizes the distortion of the gradient aggregation signal based on AirComp by reducing the transmission power of each device in the distributed learning framework, accelerating the convergence speed of the distributed learning framework, reducing the device transmission power and communication energy consumption.
[0092] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A device transmission power configuration method for accelerating the convergence speed of a distributed learning framework, characterized in that: The following steps are involved: Step S1: Establish a distributed learning framework; The distributed learning framework includes: K devices and 1 base station; Step S2: When entering each training round, each device obtains its channel gain vector to the base station, the base station's receive beamforming configuration scheme and the local gradient to be sent; Step S3: Each device sets its own transmit power based on the amplified air-calculated distortion as a criterion, and concurrently sends a local gradient to the base station using the set transmit power on the same time-frequency resource; The criterion for amplifying the air calculation distortion is: under the premise that the ratio of the device power proportional factor to the base station normalization factor is greater than 1, the device power proportional factor is reduced and the base station normalization factor is simultaneously reduced, so that the MSE increases, thereby amplifying the air calculation distortion; Step S4: The base station aggregates the local gradients of each device based on air computing technology to obtain the global gradient of the current training round, and uses the global gradient to update the global model.
2. The device transmission power configuration method for accelerating the convergence speed of a distributed learning framework according to claim 1 is characterized in that: The calculation formula of the local gradient is: , in, represents the local gradient of device k in the tth training round, represents the gradient operator, Represents the nth data sample in the local data set of device k The corresponding loss function is, represents the global model updated in the previous training round broadcasted by the base station in the tth training round, Represents the local dataset of device k in the tth training round.
3. The device transmission power configuration method for accelerating the convergence speed of a distributed learning framework according to claim 1 is characterized in that: The calculation formula of the transmission power is: , in, represents the transmit power of device k, represents the receive beamforming configuration scheme of the base station, represents the channel gain vector from device k to the base station, represents the non-negative device power scaling factor, represents the conjugate transpose.
4. The device transmission power configuration method for accelerating the convergence speed of a distributed learning framework according to claim 1, characterized in that: The distortion of the aerial calculation is measured by the mean square error of the global gradient; the mean square error of the aerial calculation The calculation formula is: , in, represents the non-negative device power scaling factor, represents the base station normalization factor, Indicates the noise intensity.
5. The device transmission power configuration method for accelerating the convergence speed of a distributed learning framework according to claim 1, characterized in that: The calculation formula of the global gradient is: , in, represents the global gradient of base station aggregation in the tth training round, represents the receiving noise of the base station, represents the base station normalization factor, represents the receive beamforming configuration scheme of the base station, represents the channel gain vector from device k to the base station, represents the conjugate transpose, represents the transmit power of device k, represents the local gradient of device k in the tth training epoch.
6. The device transmission power configuration method for accelerating the convergence speed of a distributed learning framework according to claim 5, characterized in that: Based on the transmit power of device k The configuration scheme of the global gradient of the base station aggregation in the tth training round Further expressed as: , in, Represents the non-negative device power scaling factor.
7. The device transmission power configuration method for accelerating the convergence speed of a distributed learning framework according to claim 1, characterized in that: The base station updates the global model using the gradient descent formula: , in, represents the global model updated in the previous training round broadcasted by the base station in the tth training round, represents the global model updated in the previous training round broadcasted by the base station in the t-1th training round, represents the learning rate, represents the global gradient of base station aggregation in the tth training round.
8. The device transmission power configuration method for accelerating the convergence speed of a distributed learning framework according to claim 1, characterized in that: The convergence speed of the distributed learning framework is defined as the expected decrease value of the global loss function between two adjacent training rounds; the mathematical expression of the convergence speed of the distributed learning framework is: , in, represents the global loss function of the tth training round, represents the global loss function of the t-1th training round, represents the global model updated in the previous training round broadcasted by the base station in the tth training round, Represents the global model updated in the previous training round broadcasted by the base station in the t-1th training round.
9. The device transmission power configuration method for accelerating the convergence speed of a distributed learning framework according to claim 8, characterized in that: The mathematical expression of the global loss function of the tth training round is: , in, Represents the nth data sample in the local data set of device k The corresponding loss function, D represents the number of data samples in the local dataset.
Citation Information
Patent Citations
Working method of digital pre-distortion system based on iterative learning control and principal curve analysis
CN113612455A
Dynamic power control method and system for resisting multi-user parameter biased aggregation in federated learning
US11956726B1