Aerial federated learning method based on rare 1-bit quantization, and a server and device therefor.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- CHUNG ANG UNIV IND ACADEMIC COOP FOUND
- Filing Date
- 2025-04-18
- Publication Date
- 2026-07-31
AI Technical Summary
【0018】 本発明の希少1ビット量子化に基づく空中連合学習方法並びにそのためのサーバ及びデバイスによれば、希少化に基づく圧縮率の動的制御に基づいて共同の最適電力及びエラー制御を行う新たな圧縮及び伝送が可能であり、これにより通信コストを低減することができる。
Smart Images

Figure 0007898208000188 
Figure 0007898208000189 
Figure 0007898208000190
Abstract
Claims
1. A federated Learning Over-the-Air (FLOA) method for performing distributed machine learning in a wireless network space based on rare 1-bit quantization by a computer-equipped device, The process involves receiving a global parametric vector as a parameter for processing data stored on the device in each communication round for aggregating the local gradients of the learning model from the server, Based on the global parametric vector, a first-order approximation algorithm is applied using the local data set held by the device to derive a local gradient vector, and the local gradient vector is compressed by 1-bit quantization for each layer and then scarceed. The process includes the step of transmitting the reduced 1-bit quantized local gradient vector to the server via an uplink channel using an analog method. The server aggregates the rare 1-bit quantized local gradient vectors transmitted from each device, updates the global parameter vector by repeating a federated learning process by each device until the convergence condition is met or the maximum communication round is reached, and then broadcasts the results, characterized in that it is an aerial federated learning method based on rare 1-bit quantization.
2. The step of compressing each layer by 1-bit quantization and then reducing its value is as follows: The step of calculating the size scaling elements of the local gradient vectors for each layer, A method for aerial federation learning based on rare one-bit quantization according to claim 1, comprising the step of deriving the rare one-bit quantization local gradient vector using the calculated size scaling elements and layer-specific scarcity masking indicators.
3. The layer-specific scarcity masking indicator is an indicator that determines whether a layer is transmitted, and is transmitted from the server in each communication round, as described in claim 2, for an aerial federated learning method based on scarcity 1-bit quantization.
4. The airborne federated learning method based on rare 1-bit quantization according to claim 2, characterized in that layers for which the scarcity masking indicator is 0 are excluded from transmission.
5. The aerial federated learning method based on rare 1-bit quantization according to claim 1, characterized in that the aforementioned layer is a layer constituting a deep learning model.
6. The method for learning an aerial federation based on rare 1-bit quantization according to claim 1, characterized in that deriving the local gradient vector includes calculating a local gradient vector corrected using the previous round compression error vector.
7. Updating the aforementioned global parametric vector means The analog signal received via the uplink is corrected by an amplitude scaling element and transmission power to reconstruct the gradient vector. The reconstructed gradient vectors are aggregated, and the global gradient vector is calculated and aggregated. The aerial federation learning method based on rare 1-bit quantization according to claim 1, characterized by including updating the global parametric vector using the aggregated global gradient vector.
8. The aerial federated learning method based on rare 1-bit quantization according to claim 7, further comprising the step of the server transmitting the reconstructed gradient vector to the device.
9. A federated Learning Over-the-Air (FLOA) method that performs distributed machine learning in a wireless network space based on rare 1-bit quantization by a server equipped with a computer, (a) In each communication round, based on the power loss function and constraints of each device, a step is to determine and transmit to each device a layer-specific amplitude scaling element and a scarce masking indicator, which is a binary value indicating whether or not a layer is transmitted. (b) A step of receiving a signal from each device which is a 1-bit quantized local gradient vector that has been compressed and reduced by 1-bit quantization for each layer, derived using the local data set held by each device, (c) The step of correcting the received signal with an amplitude scaling element and transmission power to reconstruct the gradient vector, (d) A step of aggregating the reconstructed gradient vectors and calculating and aggregating the global gradient vector, (e) a step of updating a global parametric vector using the aggregated global gradient vector, characterized in that it is an aerial federation learning method based on rare 1-bit quantization.
10. The aerial federation learning method based on rare 1-bit quantization according to claim 9, characterized in that step (a) is performed in parallel to determine the amplitude scaling element and the binary rare masking indicator, and the binary rare masking indicator is determined after the amplitude scaling element is determined.
11. The airborne federated learning method based on rare 1-bit quantization according to claim 9, characterized in that the rare masking indicator is an indicator that determines whether or not the layer is transmitted.
12. A computer-readable recording medium having a program code for performing the method described in any one of claims 1 to 11.
13. Communications Department and, Memory to store at least one instruction word, A device comprising a processor that executes instruction words stored in the memory, When the instruction word is executed, the processor In each communication round for aggregating the local gradients of the learning model, the server receives a global parametric vector as a parameter for processing the data stored on the device. Based on the global parameter vector, a first-order approximation algorithm is applied using the local data set held by the device to derive a local gradient vector, and the local gradient vector is compressed by 1-bit quantization for each layer and then scarce. The system is configured to transmit the reduced 1-bit quantized local gradient vector to the server via the uplink channel using an analog method. The server is characterized by aggregating the rare 1-bit quantized local gradient vectors transmitted from each device, updating the global parameter vector by repeating a federated learning process by each device until the convergence condition is met or the maximum communication round is reached, and then broadcasting the result.
14. The aforementioned processor, After compressing each of the aforementioned layers by 1-bit quantization and then reducing its value, Calculate the size and scaling elements of the local gradient vectors for each layer. The device according to claim 13, characterized in that it is configured to derive the rare 1-bit quantized local gradient vector using the calculated size scaling elements and layer-specific scarcity masking indicators.
15. Communications Department and, Memory to store at least one instruction word, A server comprising a processor that executes instruction words stored in the memory, When the instruction word is executed, the processor (a) In each communication round, based on the power loss function and constraints of each device, layer-specific amplitude scaling elements and scarce masking indicators, which are binary values indicating whether or not a layer is transmitted, are determined and transmitted to each device. (b) From each device, receive a signal which is a 1-bit quantized local gradient vector, which is a local gradient vector derived using the local data set held by each device and compressed and reduced by 1-bit quantization for each layer. (c) The received signal is corrected by an amplitude scaling element and transmission power to reconstruct the gradient vector, (d) The reconstructed gradient vectors are aggregated, and the global gradient vector is calculated and aggregated, (e) A server characterized by being configured to update the global parametric vector using the aggregated global gradient vector.