Frequency difference compressor based on precision perception, gradient compression method, equipment and medium
Through the multi-level gradient compression method of the frequency difference compressor, the gradient compression process is dynamically optimized by the spatial redundancy, frequency domain sparseness and timing consistency of the gradient, which solves the communication bottleneck problem of gradient transmission in distributed deep learning, and realizes efficient gradient information retention and model performance maintenance.
Patent Information
- Application Number
- CN202510741210.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The existing distributed deep learning gradient compression method has too much communication overhead during high-dimensional feature gradient transmission, so it is impossible to retain gradient information as much as possible, especially in low bandwidth or high-latency network environments in federated learning and edge computing scenarios.
The frequency difference compressor based on accuracy perception is adopted, including information bottleneck sparse unit, frequency domain compression unit, dynamic quantization unit and differential coding unit. By removing non-critical gradients, frequency domain transformation, adaptive quantization and differential coding, the balance between communication data volume and information accuracy is optimized.
It significantly reduces the communication overhead in personalized federated learning, maintains model accuracy and convergence performance, is suitable for resource-constrained environments such as low bandwidth and high latency, and has good adaptability and universality.
Smart Images

Figure CN120258058A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of data compression and distributed deep learning. Specifically, it relates to a frequency difference compressor and gradient compression method, device, and medium based on precision perception. Background Art
[0002] Distributed deep learning significantly improves the efficiency of model training and privacy protection capabilities by distributing computing tasks to multiple client devices, and is widely used in scenarios such as federated learning and edge computing. However, in distributed deep learning, high-dimensional feature gradients need to be frequently transmitted between the client and the server, and communication overhead has become the main bottleneck restricting its performance, especially more prominent in low-bandwidth or high-latency network environments. For example, when training a deep neural network, the gradient tensor usually contains millions or even tens of millions of floating-point parameters, resulting in a huge amount of data for a single communication, significantly increasing latency and energy consumption. In addition, client devices such as mobile devices or Internet of Things nodes often have limited computing and storage resources, further exacerbating the need for efficient communication.
[0003] To alleviate this problem, a series of gradient compression methods have been proposed in the prior art to reduce the amount of data transmission. These methods include gradient sparsification, quantization, low-rank decomposition, and differential coding. Gradient sparsification methods, such as Top-k sparsification, retain the top k% of elements by sorting according to the absolute value of the gradient, or randomly sparsify to discard elements with a fixed probability to reduce the dimension; gradient quantization techniques, such as SignSGD, only transmit the gradient sign and cooperate with a scaling factor to approximate the original value, while QSGD uses a fixed bit width to discretize the gradient into a finite set of values; low-rank decomposition methods, such as PowerSGD, approximate the gradient tensor as the product of low-dimensional matrices through singular value decomposition to reduce the communication volume; differential coding methods, such as EF-SGD, utilize the temporal correlation of the gradient, only transmit the difference between the current round and the previous round, and combine an error feedback mechanism to compensate for compression losses. However, although the prior art has achieved certain results in reducing the communication volume, there are still significant limitations, and it is difficult to balance a high compression rate and model performance. For example, Top-k or random sparsification methods lack adaptability, may discard key small-magnitude gradients or introduce too much noise, resulting in a decrease in the convergence speed or impaired performance; fixed-bit-width quantization cannot adapt to changes in the gradient distribution, and low-bit quantization performs poorly when high precision is required in the early stage of training, while high-bit quantization has insufficient compression rate; low-rank decomposition has a large computational overhead, is not suitable for resource-constrained devices, and low-rank approximation may not be able to capture the complex structure of the gradient, affecting convergence. In addition, the existing methods lack systematic multi-level integration, and the trade-off between the compression rate and information retention is not ideal, making it difficult to meet the needs of diverse distributed learning scenarios.
[0004] Therefore, this application is specifically proposed. Summary of the Invention
[0005] The present invention aims to provide a frequency difference compressor based on precision perception, as well as a gradient compression method, device, and medium, to solve the problem of excessive communication overhead and inability to retain gradient information as much as possible during the transmission of high-dimensional feature gradients in existing distributed deep learning gradient compression methods. Especially in the scenarios of federated learning and edge computing, the communication bottleneck in low-bandwidth or high-latency network environments is particularly prominent.
[0006] To solve the above technical problems, the present invention is realized through the following technical solutions: A frequency difference compressor based on precision perception, used for data compression and transmission between a server and a client to reduce the communication overhead in personalized federated learning, includes: An information bottleneck sparsification unit, which is used to remove non-critical gradient redundancy components in the original gradient data through a dynamic threshold according to the obtained original gradient data sent by the server, and retain critical gradients to generate sparse gradient data; A frequency domain compression unit, which is used to convert the sparse gradient data into a frequency domain representation through discrete cosine transform to obtain frequency domain coefficients, and only retain the first N high-energy frequency domain coefficients according to a preset compression rate to generate a meta data packet containing the first N high-energy frequency domain coefficients and their positions and quantities; A dynamic quantization unit, which is used to adaptively map the high-precision floating-point gradient data of the meta data packet to a low-bit integer representation to generate dynamic quantization gradient data, so as to optimize the balance between communication data volume and information precision; A differential coding unit, which is used to only retain the gradient change part, that is, the differential gradient, in the dynamic quantization gradient data according to the temporal correlation of gradients between consecutive training rounds, and transmit it to the client to complete data compression and transmission.
[0007] Preferably, the information bottleneck sparsification unit is specifically used for: Obtain the gradient statistical characteristics of the original gradient data, including the absolute value mean, second norm, and current training loss of gradient elements; Calculate a dynamic threshold based on the information bottleneck principle to distinguish critical gradients and non-critical gradients in the original gradient data, ensuring that more information is retained in the early stage of training and more strict compression is performed in the later stage; Compare the gradient elements of the original gradient data with the dynamic threshold, retain the gradient elements with absolute values higher than the dynamic threshold and generate a position mask to obtain sparse gradients, which is convenient for the client to restore the gradient structure during decompression.
[0008] Preferably, the dynamic threshold is calculated by comprehensively considering the gradient amplitude, current training loss, and distribution characteristics of the original gradient data; the expression of the dynamic threshold is: ; Where is the dynamic threshold; is the trade-off coefficient; represents the two-norm; represents the gradient element; , represents the average absolute value of the gradient elements; , represents the dimension of the data; is the number of samples; is the channel; is the height; is the width; is the current loss; is the exponential coefficient to prevent logarithmic divergence.
[0009] Preferably, the expression of the sparse gradient is: ; wherein, represents the sparse gradient; 1(·) is the indicator function.
[0010] Preferably, the frequency domain compression unit is specifically configured to: flatten the gradient tensor of the sparse gradient data into a one-dimensional vector; perform an orthogonal discrete cosine transform on the flattened one-dimensional vector to generate frequency domain coefficients; wherein, the frequency domain coefficients are concentrated in the low-frequency region; select the first N high-energy coefficients from the frequency domain coefficients according to a preset compression ratio, and set the remaining coefficients to zero to obtain compressed coefficients; transmit the positions and the number of the first N high-energy coefficients, and the compressed coefficients as a meta data packet to the client, so that the client performs an inverse discrete cosine transform to recover an approximate gradient and reshape it into the original tensor shape.
[0011] Preferably, the dynamic quantization unit is specifically configured to: calculate the statistics of the gradient according to the gradient distribution characteristics of the meta data packet; wherein, the statistics include the minimum value, the maximum value and the standard deviation of the gradient; dynamically determine the quantization bit width based on the statistics, for selecting a high bit width to retain the accuracy in the early stage of training and reducing the bit width in the later stage of training to improve the compression ratio, and the expression is: ; wherein, b is the quantization bit width; , , respectively represent the maximum value, the minimum value and the variance of the gradient; is the exponential coefficient to prevent division by zero; represents the gradient element; represents floor function; Calculate a scaling factor according to the statistic and the quantization bit width to map the gradient to an integer range, and round it to generate a quantized gradient in integer form. The expression is: ; ; where s is the scaling factor; represents the quantized gradient; represents the rounding operation; represents a gradient element; Transmit the quantized gradient, the scaling factor, and the zero point information as dynamic quantized gradient data, so that the client can receive the dynamic quantized gradient data and perform a linear transformation to restore it to a high-precision floating-point gradient. The expression is: ; where, is the restored high-precision floating-point gradient; If it is signed gradient quantization, the zero point information needs to be transmitted, that is, the data corresponding to the integer representing 0 in the quantized gradient, so that the client can correctly decode it; if it is unsigned quantization and 0 is the minimum value, the zero point information is the minimum value of the gradient.
[0012] Preferably, the differential encoding unit is specifically used for: Obtain the gradient data of the previous round of training stored on the server side and initialize it; Calculate the difference between the gradient of the current training round and the gradient of the previous training round to generate a differential gradient. The expression is: ; where, represents the differential gradient of the t-th round; represents the gradient of the t-th round of training, that is, the gradient of the current training round; is the gradient of the previous training round; if the current training round is the first round, that is, t = 1, then is 0, that is, the complete gradient is transmitted in the first round, and the differential gradient is transmitted in subsequent rounds; Transmit the differential gradient to the client to combine the gradient data of the previous round stored on the client side, restore the complete gradient data of the current training round, and update the storage to prepare for the differential encoding calculation of the next round.
[0013] The present invention also provides a gradient compression method, including: Obtain the original gradient data sent by the server; Remove the non-critical gradient redundancy components in the original gradient data through a dynamic threshold, and retain the key gradients to generate sparse gradient data; Convert the sparse gradient data into a frequency-domain representation through discrete cosine transform to obtain frequency-domain coefficients, and only retain the first N high-energy frequency-domain coefficients according to a preset compression rate to generate a meta-data packet containing the first N high-energy frequency-domain coefficients and their positions and quantities; Adaptively map the high-precision floating-point gradient data of the meta-data packet to a low-bit integer representation to generate dynamically quantized gradient data, so as to optimize the balance between communication data volume and information accuracy; According to the temporal correlation of gradients between consecutive training rounds, only retain the changing part of the gradients in the dynamically quantized gradient data, that is, the differential gradients, and transmit them to the client to complete data compression and transmission.
[0014] The present invention also provides a gradient compression device, including a processor and a memory. The memory stores a computer program, and the computer program can be executed by the processor to implement a gradient compression method as described above.
[0015] The present invention also provides a computer-readable storage medium. Computer-readable instructions are stored on the computer-readable storage medium. When the computer-readable instructions are executed by the processor of the device where the computer-readable storage medium is located, a gradient compression method as described above is implemented.
[0016] In summary, compared with the prior art, the present invention has the following beneficial effects: The present invention sequentially processes the original gradient through sparsification, frequency-domain transformation, dynamic quantization, and differential coding. Each stage optimizes different characteristics of the gradient to compress the gradient with the maximum compression rate and retain the key information of the gradient as much as possible. Finally, a highly compressed representation is generated for transmission between the client and the server. The client decompresses the data through the reverse process to restore the approximate gradient to update the model, ensuring low computational overhead in the compression and decompression processes and adapting to resource-constrained distributed environments at the same time.
[0017] The present invention makes full use of the spatial redundancy, frequency-domain sparsity, and temporal consistency of gradients, sets dynamic thresholds to retain key information, dynamically determines the quantization bit width to improve the compression rate while retaining the accuracy, and realizes the dual improvement of communication efficiency and model performance through the collaborative work of each stage, significantly reducing the communication overhead in personalized federated learning, having good adaptability and generality, being particularly suitable for resource-constrained environments such as low bandwidth and high latency, and maintaining the model accuracy and convergence performance at the same time.
[0018] The present invention provides an efficient and general gradient compression framework, which solves the trade-off problem between high compression rate and information retention in the prior art. Description of the Drawings
[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0020] Figure 1 Schematic structural diagram of personalized federated learning using a frequency difference compressor based on precision perception provided for Embodiment 1.
[0021] Figure 2 Schematic flowchart of a gradient compression method provided for Embodiment 2.
[0022] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Specific Embodiments
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0024] Embodiment 1 In the fields of machine learning and deep learning, gradients are important information used to guide the update of model parameters. A gradient is usually a multi-dimensional tensor (such as a matrix or an array of higher dimensions), where each element represents the rate of change or sensitivity of the model parameters in a certain direction. Specifically, the gradient elements represent the change trend of the model parameters in the loss function space, that is, the degree of influence of small parameter changes on the loss function value.
[0025] When training a neural network, the gradients of the model parameters are determined by calculating the loss function and adjusting the model parameters to minimize the loss function. The gradient elements can be real numbers, usually represented in floating-point form, and their magnitudes and signs (positive or negative) indicate the direction and magnitude of parameter adjustment.
[0026] In this embodiment, the gradient element refers to each real number element in the gradient tensor calculated during the model training process. These elements are the basis for subsequent steps such as sparsification, frequency domain transformation, quantization coding, and differential processing.
[0027] As Figure 1 shown, a frequency difference compressor based on precision awareness is used for data compression and transmission between a server and n clients to reduce the communication overhead in personalized federated learning. It includes an information bottleneck sparsification unit, a frequency domain compression unit, a dynamic quantization unit, and a differential coding unit. The obtained original feature / gradient data is efficiently compressed through the collaborative work of the information bottleneck sparsification unit, the frequency domain compression unit, the dynamic quantization unit, and the differential coding unit in sequence.
[0028] The information bottleneck sparsification unit is used to remove the non-critical gradient redundancy components in the original gradient data according to the obtained original gradient data sent by the server through a dynamic threshold, and retain the critical gradients to generate sparse gradient data. Specifically, it includes the following steps: First, obtain the gradient statistical characteristics of the original gradient data, including the absolute value mean, the second norm, and the current training loss of the gradient elements; Next, calculate the dynamic threshold based on the information bottleneck principle to distinguish the critical gradients and non-critical gradients in the original gradient data, ensuring that more information is retained in the early training and more strict compression is performed in the later stage; the dynamic threshold is calculated by comprehensively considering the gradient magnitude, the current training loss, and the distribution characteristics of the original gradient data; the expression of the dynamic threshold is: ; Among them, is the dynamic threshold; is the trade-off coefficient; represents the second norm; represents the gradient element; , represents the absolute value mean of the gradient elements; , represents the dimension of the gradient data, is the number of samples, that is, the number of samples in one training; is the channel; is the height; is the width; is the current loss; is the exponential coefficient, let , which is used to prevent logarithmic divergence, and other exponential values can also be set, which are not limited here.
[0029] Finally, compare the gradient elements of the original gradient data with the dynamic threshold, retain the gradient elements whose absolute values are higher than the dynamic threshold, generate a position mask, obtain the sparse gradient, and record the position mask of the non-zero elements to facilitate the client to restore the gradient structure during decompression. The expression of the sparse gradient is: ; where represents the sparse gradient; 1(·) is the indicator function.
[0030] The frequency domain compression unit is used to convert the sparse gradient data into a frequency domain representation through discrete cosine transform, obtain the frequency domain coefficients, and retain only the first N high-energy frequency domain coefficients according to the preset compression rate, and generate a meta-data packet containing the first N high-energy frequency domain coefficients and their positions and quantities.
[0031] Specifically, it includes the following steps: First, flatten the gradient tensor of the sparse gradient data into a one-dimensional vector for easy transformation processing. For example, for the gradient tensor (B = 256, C = 512, H = 7, W = 7), the dimension after flattening is d = 256 × 512 × 7 × 7 = 6,422,528.
[0032] Next, perform an orthogonal discrete cosine transform on the flattened one-dimensional vector to generate the frequency domain coefficients; among them, the frequency domain coefficients are usually concentrated in the low-frequency region, which is expressed as: ; where is the frequency domain coefficient, DCT is the orthogonal discrete cosine transform operation; is the flattened one-dimensional vector; norm = ortho represents the setting of the orthogonal normalization method for the discrete cosine transform (DCT); represents real numbers in d dimensions.
[0033] Then, according to the preset compression rate, select the first N high-energy coefficients from the frequency domain coefficients, and set the remaining coefficients to zero to obtain the compressed coefficients, which is expressed as: ; where is the i-th frequency domain coefficient; is the selected high-energy coefficient; N is the number of retained high-energy coefficients, , is the compression rate, for example, let r = 0.1; represents rounding down.
[0034] Finally, transmit the positions and quantities of the top N high-energy coefficients, along with the compression coefficient, to the client as a meta-data packet. After receiving the meta-data packet, the client performs an inverse discrete cosine transform to restore the approximate gradient and reshape it into the original tensor shape.
[0035] A dynamic quantization unit is used to adaptively map the high-precision floating-point gradient data of the meta-data packet to a low-bit integer representation, generating dynamic quantization gradient data to optimize the balance between communication data volume and information accuracy.
[0036] Specifically, it includes the following steps: First, calculate the statistics of the gradient according to the gradient distribution characteristics of the meta-data packet; among them, the statistics include the minimum value, maximum value, and standard deviation of the gradient; these indicators reflect the dynamic range and concentration of the gradient.
[0037] Next, based on the statistics, dynamically determine the quantization bit width. The quantization bit width range is usually between 2 and 8 bits, which is used to select a high bit width to retain accuracy in the early stage of training and reduce the bit width in the later stage of training to improve the compression rate. The expression is: ; where b is the quantization bit width; 、 、 respectively represent the maximum value, minimum value, and variance of the gradient; is the exponential coefficient. Let , to prevent division by zero, and other exponential values can also be set, which are not limited here; Then, calculate the scaling factor according to the statistics and the quantization bit width to map the gradient to the integer range and round it to generate the quantized gradient in integer form. The expression is: ; ; where s is the scaling factor; represents the quantized gradient; represents the rounding operation; represents the gradient element; Finally, transmit the quantized gradient, the scaling factor, and the zero-point information as dynamic quantization gradient data, so that the client can receive the dynamic quantization gradient data and perform a linear transformation to restore it to a high-precision floating-point gradient, significantly reducing the data transmission volume. The expression: ; where is the restored high-precision floating-point gradient; For signed gradient quantization, zero-point information needs to be transmitted, that is, the data corresponding to the integer representing 0 in the quantized gradient, so that the client can decode correctly; for unsigned quantization and 0 being the minimum value, the zero-point information is the minimum value of the gradient.
[0038] The differential encoding unit is used to only retain the part of the gradient that changes, that is, the differential gradient, in the dynamically quantized gradient data according to the temporal correlation of the gradients between consecutive training rounds, and transmit it to the client to complete data compression and transmission.
[0039] Specifically, it includes the following steps: First, obtain and initialize the gradient data of the previous round of training stored on the server side, such as initializing it to a zero tensor; Next, calculate the difference between the gradient of the current training round and the gradient of the previous training round to generate a differential gradient. In the first round, since there is no previous round of data, the complete gradient is directly transmitted, which is expressed as: ; Among them, represents the differential gradient of the t-th round; represents the gradient of the t-th round of training, that is, the gradient of the current training round; is the gradient of the previous training round; if the current training round is the first round, that is, t = 1, then is 0, that is, the complete gradient is transmitted in the first round, and the differential gradient is transmitted in subsequent rounds.
[0040] Finally, transmit the differential gradient to the client to combine with the gradient data of the previous round stored locally by the client to restore the complete gradient data of the current training round. After the transmission is completed, the server and the client synchronously update the stored gradients to prepare for the next round of differential encoding calculation. This method has a significant effect when the model is approaching convergence because the gradient change tends to be stable and the communication volume is greatly reduced.
[0041] In Figure 1 , in addition to being able to apply the frequency difference compressor of the present invention to the server side to send the original gradient data, which is compressed by the frequency difference compressor to obtain compressed gradient data and transmitted to each client for model parameter update, the original feature data sent by the client can also be compressed by the frequency difference compressor to obtain compressed feature data and transmitted to the server side, and after being reversely decompressed by the server side, it is used for model training.
[0042] The process of compressing the feature data by the frequency difference compressor is the same as the process of gradient compression. It goes through four stages: information bottleneck sparsification, frequency domain compression, dynamic quantization, and differential encoding in sequence. Each stage is optimized for different characteristics of the features to compress the features with the maximum compression rate and retain key information as much as possible, and finally generate a highly compressed feature representation and transmit it to the server.
[0043] In summary, compared with the prior art, the present invention has the following beneficial effects: By integrating four techniques: information bottleneck sparsification, frequency-domain compression, dynamic quantization, and differential coding, the present invention proposes a frequency difference compressor with multi-level gradient compression. Starting from the original gradient, each stage optimizes different characteristics of the gradient to compress the gradient with the maximum compression rate and retain key information as much as possible. Finally, a highly compressed representation is generated for transmission between the client and the server, significantly reducing the communication overhead in personalized federated learning while maintaining model accuracy and convergence performance. The client decompresses the data through the reverse process to recover the approximate gradient for updating the model.
[0044] The present invention makes full use of the spatial redundancy, frequency-domain sparsity, and temporal consistency of the gradient, and has good adaptability and generality, especially suitable for resource-constrained environments such as low bandwidth and high latency. In addition, the client adopts efficient tensor operations and hardware acceleration optimization strategies to ensure low computational overhead in the compression and decompression processes, while adapting to the resource-constrained distributed environment.
[0045] Embodiment 2 As Figure 2 shown, Embodiment 2 of the present invention provides a gradient compression method, which can be implemented by a gradient compression device (hereinafter referred to as the gradient compression device), and in particular, is executed by one or more processors in the gradient compression device.
[0046] In this embodiment, the gradient compression device can be an electronic device equipped with a processor, and the processor has a computer program of this frequency difference compressor based on precision perception and the method, and the computer program can be executed, such as a computer, a smart phone, a smart tablet, a workstation, etc., which is not limited here.
[0047] A gradient compression method includes steps S1 to S5: S1, obtain the original gradient data sent by the server; S2, remove the non-critical gradient redundant components in the original gradient data through a dynamic threshold, and retain the key gradients to generate sparse gradient data.
[0048] Specifically, it includes the following steps: Obtain the gradient statistical characteristics of the original gradient data, including the absolute value mean, two-norm, and current training loss of the gradient elements; Calculate the dynamic threshold based on the information bottleneck principle to distinguish the key gradients and non-key gradients in the original gradient data, ensuring that more information is retained in the early training and more strict compression in the later stage; the dynamic threshold is calculated by comprehensively considering the gradient amplitude, current training loss, and distribution characteristics of the original gradient data; the expression of the dynamic threshold is: ; Among them, is the dynamic threshold; is the trade-off coefficient; represents the two-norm; represents the gradient element; , representing the average absolute value of the gradient elements; , representing the dimension of the data; is the number of samples; is the channel; is the height; is the width; is the current loss; Let be the exponential coefficient to prevent logarithmic divergence. Other exponential values can also be set, which are not limited here.
[0049] Compare the gradient elements of the original gradient data with the dynamic threshold, retain the gradient elements with absolute values higher than the dynamic threshold and generate a position mask to obtain the sparse gradient, which is convenient for the client to restore the gradient structure during decompression; the expression of the sparse gradient is: ; Among them, represents the sparse gradient; 1(·) is the indicator function.
[0050] S3. Convert the sparse gradient data into a frequency-domain representation through discrete cosine transform to obtain frequency-domain coefficients, and only retain the first N high-energy frequency-domain coefficients according to the preset compression rate to generate a metadata packet containing the first N high-energy frequency-domain coefficients and their positions and quantities.
[0051] Specifically, it includes the following steps: Flatten the gradient tensor of the sparse gradient data into a one-dimensional vector; Perform an orthogonal discrete cosine transform on the flattened one-dimensional vector to generate frequency-domain coefficients; among them, the frequency-domain coefficients are concentrated in the low-frequency region; According to the preset compression rate, select the first N high-energy coefficients in the frequency-domain coefficients and set the remaining coefficients to zero to obtain the compressed coefficients; Use the positions and quantities of the first N high-energy coefficients as metadata, and transmit them together with the compressed coefficients to the client, so that the client can perform an inverse discrete cosine transform to restore the approximate gradient and reshape it into the original tensor shape.
[0052] S4. Adaptively map the high-precision floating-point gradient data of the metadata packet to a low-bit integer representation to generate dynamically quantized gradient data, so as to optimize the balance between communication data volume and information accuracy.
[0053] Specifically, it includes the following steps: Calculate the statistics of the gradient according to the gradient distribution characteristics of the meta data packet; wherein, the statistics include the minimum value, the maximum value and the standard deviation of the gradient; Based on the statistics, dynamically determine the quantization bit width, which is used to select a high bit width to retain the precision in the early stage of training and reduce the bit width in the later stage of training to improve the compression ratio. The expression is: ; where b is the quantization bit width; 、 、represent the maximum value, the minimum value and the variance of the gradient respectively; let be the exponential coefficient to prevent division by zero, and other exponential values can also be set, which will not be limited here; represents rounding down; Calculate the scaling factor according to the statistics and the quantization bit width to map the gradient to the integer range and round it to generate the quantized gradient in integer form. The expression is: ; ; where s is the scaling factor; represents the quantized gradient; represents the rounding operation; represents the gradient element; Transmit the quantized gradient, the scaling factor and the zero point information as dynamic quantized gradient data, so as to receive the dynamic quantized gradient data through the client and restore it to a high-precision floating-point gradient through linear transformation. The expression: ; where, is the restored high-precision floating-point gradient; If it is signed gradient quantization, the zero point information needs to be transmitted, that is, the data corresponding to the integer representing 0 in the quantized gradient, so that the client can decode correctly; if it is unsigned quantization and 0 is the minimum value, the zero point information is the minimum value of the gradient.
[0054] S5. According to the temporal correlation of the gradients between consecutive training rounds, only retain the gradient change part, that is, the differential gradient, in the dynamic quantized gradient data and transmit it to the client to complete data compression and transmission.
[0055] Specifically, it includes the following steps: Obtain and initialize the gradient data of the previous round of training stored on the server side; Calculate the difference between the gradient of the current training round and the gradient of the previous training round to generate a differential gradient. The expression is: ; where, Denote the differential gradient of the t-th round; Denote the gradient of the t-th round of training, i.e., the gradient of the current training round; Is the gradient of the previous training round. If the current training round is the first round, i.e., t = 1, then Is 0, that is, the complete gradient is transmitted in the first round, and the differential gradient is transmitted in subsequent rounds; Transmit the differential gradient to the client to combine with the gradient data of the previous round stored by the client, restore the complete gradient data of the current training round, and update the storage to prepare for the differential coding calculation of the next round.
[0056] Embodiment III The third embodiment of the present invention also provides a gradient compression device, which includes a memory and a processor. A computer program is stored in the memory, and the computer program can be executed by the processor to implement a gradient compression method as described above.
[0057] Embodiment IV The fourth embodiment of the present invention also provides a computer-readable storage medium. Computer-readable instructions are stored on the computer-readable storage medium. When the computer-readable instructions are executed by the processor of the device where the computer-readable storage medium is located, a gradient compression method as described above is implemented.
[0058] In several embodiments provided by the embodiments of the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are merely illustrative. For example, the flowcharts in the drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0059] In addition, the functional modules in each embodiment of the present invention can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part.
[0060] When the above-mentioned functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs. It should be noted that in this article, the terms "include", "comprise", or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such a process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article, or device including the said element.
[0061] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise.
[0062] It should be understood that the term "and / or" used herein is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0063] Depending on the context, the word "if" as used herein can be interpreted as "when", "while", "in response to determining", or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined", "in response to determining", "when detecting (stated condition or event)", or "in response to detecting (stated condition or event)".
[0064] The "first / second" mentioned in the embodiments is only used to distinguish similar objects and does not represent a specific order for the objects. It can be understood that the "first / second" can be interchanged in a specific order or sequence when permitted. It should be understood that the objects distinguished by the "first / second" can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.
[0065] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various modifications and changes. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A frequency difference compressor based on precision perception is used for data compression and transmission between a server and a client to reduce the communication overhead in personalized federated learning. It is characterized in that It includes: An information bottleneck sparsification unit, which is used to remove non-critical gradient redundancy components in the original gradient data through a dynamic threshold according to the obtained original gradient data sent by the server, and retain critical gradients to generate sparse gradient data; A frequency domain compression unit, which is used to convert the sparse gradient data into a frequency domain representation through discrete cosine transform to obtain frequency domain coefficients, and only retain the first N high-energy frequency domain coefficients according to a preset compression rate to generate a meta data packet containing the first N high-energy frequency domain coefficients and their positions and quantities; A dynamic quantization unit, which is used to adaptively map the high-precision floating-point gradient data of the meta data packet to a low-bit integer representation to generate dynamic quantization gradient data, so as to optimize the balance between communication data volume and information accuracy; A differential coding unit, which is used to only retain the gradient change part, that is, the differential gradient, in the dynamic quantization gradient data according to the temporal correlation of gradients between consecutive training rounds, and transmit it to the client to complete data compression and transmission.
2. The frequency difference compressor based on precision perception according to claim 1, characterized in that , The information bottleneck sparsification unit is specifically used for: Obtain the gradient statistical characteristics of the original gradient data, including the absolute value mean, second norm, and current training loss of gradient elements; Calculate a dynamic threshold based on the information bottleneck principle to distinguish critical gradients and non-critical gradients in the original gradient data; Compare the gradient elements of the original gradient data with the dynamic threshold, retain the gradient elements with absolute values higher than the dynamic threshold and generate a position mask to obtain sparse gradients.
3. The frequency difference compressor based on precision perception according to claim 2, wherein , The dynamic threshold is calculated by comprehensively considering the gradient amplitude, current training loss, and distribution characteristics of the original gradient data; the expression of the dynamic threshold is: ; Among them, is the dynamic threshold; is the trade-off coefficient; represents the two-norm; represents the gradient element; , represents the average absolute value of the gradient elements; , represents the dimension of the data; is the number of samples; is the channel; is the height; is the width; is the current loss; is the exponential coefficient, used to prevent logarithmic divergence.
4. The frequency difference compressor based on precision perception according to claim 3, wherein , The expression of the sparse gradient is: ; Among them, represents the sparse gradient; 1(·) is the indicator function.
5. A frequency difference compressor based on precision perception according to claim 1, characterized in that , The frequency domain compression unit is specifically used for: Flatten the gradient tensor of the sparse gradient data into a one-dimensional vector; Perform an orthogonal discrete cosine transform on the flattened one-dimensional vector to generate frequency domain coefficients; among them, the frequency domain coefficients are concentrated in the low-frequency region; According to the preset compression rate, select the first N high-energy coefficients in the frequency domain coefficients, and set the remaining coefficients to zero to obtain compressed coefficients; Transmit the positions and quantities of the first N high-energy coefficients, and the compressed coefficients as a meta data packet to the client, so that the client performs an inverse discrete cosine transform to restore an approximate gradient and reshape it into the original tensor shape.
6. The frequency difference compressor based on precision perception according to claim 1, wherein , The dynamic quantization unit is specifically used for: Calculate the statistics of the gradient according to the gradient distribution characteristics of the meta data packet; among them, the statistics include the minimum value, maximum value, and standard deviation of the gradient; Based on the statistics, dynamically determine the quantization bit width, which is used to select a high bit width to retain accuracy in the early stage of training and reduce the bit width in the later stage of training to improve the compression rate. The expression is: ; where b is the quantization bit width; , , represent the maximum value, minimum value, and variance of the gradient respectively; is the exponential coefficient used to prevent division by zero; represents floor function; Calculate a scaling factor according to the statistics and the quantization bit width to map the gradient to the integer range, and round it to generate a quantized gradient in integer form. The expression is: ; ; where s is a scaling factor; represents the quantization gradient; represents a rounding operation; represents a gradient element; Transmit the quantized gradient, the scaling factor, and the zero point information as dynamic quantization gradient data, so that the client receives the dynamic quantization gradient data and performs a linear transformation to restore it to a high-precision floating-point gradient. The expression: ; Among them, is the restored high-precision floating-point gradient; For signed gradient quantization, the zero-point information needs to be transmitted, that is, the data corresponding to the integer representing 0 in the quantized gradient, so that the client can decode correctly; for unsigned quantization and 0 being the minimum value, the zero-point information is the minimum value of the gradient.
7. A frequency difference compressor based on precision perception according to claim 1, characterized in that , the differential encoding unit is specifically configured to: Obtain and initialize the gradient data of the previous round of training stored on the server side; Calculate the difference between the gradient of the current training round and the gradient of the previous training round to generate a differential gradient, expressed as: ; Among them, represents the differential gradient of the t-th round; represents the gradient of the t-th round of training, that is, the gradient of the current training round; is the gradient of the previous training round; if the current training round is the first round, that is, t = 1, then is 0, that is, the complete gradient is transmitted in the first round, and the differential gradient is transmitted in subsequent rounds; Transmit the differential gradient to the client to combine with the gradient data of the previous round stored on the client side, restore the complete gradient data of the current training round, and update the storage to prepare for the next round of differential encoding calculation.
8. A gradient compression method, applied to a precision-aware frequency difference compressor as described in any one of claims 1-7, characterized in that, including: Obtain the original gradient data sent by the server; Remove the non-critical gradient redundancy components in the original gradient data through a dynamic threshold and retain the critical gradients to generate sparse gradient data; Convert the sparse gradient data into a frequency-domain representation through a discrete cosine transform to obtain frequency-domain coefficients, and only retain the first N high-energy frequency-domain coefficients according to a preset compression ratio to generate a meta-data packet containing the first N high-energy frequency-domain coefficients and their positions and quantities; Adaptive bit-width map the high-precision floating-point gradient data of the meta-data packet into a low-bit integer representation to generate dynamic quantized gradient data to optimize the balance between communication data volume and information accuracy; According to the temporal correlation of the gradients between consecutive training rounds, only retain the changing part of the gradients in the dynamic quantized gradient data, that is, the differential gradient, and transmit it to the client to complete data compression and transmission.
9. A gradient compression device, characterized in that, including a processor and a memory, where the memory stores a computer program that can be executed by the processor to implement a gradient compression method as claimed in claim 8.
10. A computer-readable storage medium, characterized in that, Computer-readable instructions are stored on the computer-readable storage medium, and when the computer-readable instructions are executed by the processor of the device where the computer-readable storage medium is located, a gradient compression method as claimed in claim 8 is implemented.
Citation Information
Patent Citations
Bidirectional adaptive gradient compression method and system for federated learning
CN116470920A
Model parameter adjustment method and device based on dynamic compression, equipment and medium
CN116484946A
Federal learning-based adaptive absolute gradient compressor, training method and system
CN118261238A
Method and apparatus for improved video decompression by predetermination of IDCT results based on image characteristics
US5872866A
Coding method, decoding method, coding apparatus, decoding apparatus, and electronic device
WO2024045672A1
Cited By
Gradient compressor based on information consistency driving and gradient compression method and equipment
CN120911525A
Gradient compressor and gradient compression method and device based on information consistency driving
CN120911525B
Dynamic gradient compression learning method of federal learning system
CN120952110A
Self-adaptive sensing communication compression method, device and system facing segmentation learning
CN121151480A
Adaptive perceptual communication compression method, device and system for segmentation learning
CN121151480B