Gradient compressor based on information consistency driving and gradient compression method and equipment
By using an information consistency-driven gradient compressor, which leverages mutual information constraint modeling, multi-scale sparse representation, and Bayesian adaptive quantization, the problem of high gradient communication overhead in distributed deep learning is solved, achieving efficient and robust gradient compression and ensuring model performance.
Patent Information
- Application Number
- CN202511453033.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-13
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-10-13
AI Technical Summary
In existing distributed deep learning, gradient communication overhead is too high, especially in low-bandwidth or high-latency network environments. Existing gradient compression methods struggle to balance high compression ratios with model performance and lack systematic multi-level integration.
A gradient compressor driven by information consistency is adopted, including a mutual information constraint modeling module, a multi-scale sparse representation module, and a Bayesian adaptive quantization module. The compression operator is optimized by minimizing the Lagrangian objective function, and the sparsity parameters and quantization bit width are dynamically adjusted by combining wavelet transform and Bayesian methods to ensure the integrity of gradient information and compression efficiency.
It significantly reduces the communication overhead of distributed deep learning while preserving gradient information to the maximum extent, ensuring training performance, adapting to the needs of different training stages, and achieving efficient communication and model convergence.
Smart Images

Figure CN120911525A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of distributed learning and data compression, and particularly relates to a gradient compressor and a gradient compression method and device based on information consistency driving. BACKGROUND
[0002] Distributed deep learning significantly improves the efficiency and privacy protection capability of model training by distributing computing tasks to multiple client devices, and is widely used in federated learning, edge computing and other scenarios. However, in the process of distributed training, high-dimensional feature gradients need to be frequently transmitted between the client and the server, and the communication overhead becomes the main bottleneck restricting the system performance, especially in low-bandwidth or high-latency network environments. For example, when training a deep neural network, the gradient tensor usually contains millions or even tens of millions of floating-point parameters, resulting in a large amount of data in a single communication, increasing the delay and energy consumption. In addition, client devices such as mobile devices or Internet of Things nodes often have limited computing and storage resources, further exacerbating the need for efficient communication.
[0003] In order to solve this problem, existing technologies propose a series of gradient compression methods to reduce data transmission. These methods include gradient sparsification, quantization, low-rank decomposition and differential coding. Gradient sparsification methods such as Top-k sparsification retain the top k% elements by sorting the gradient absolute value, or random sparsification discards elements with a fixed probability to reduce dimensionality; gradient quantization techniques such as SignSGD only transmit gradient signs and cooperate with scaling factors to approximate the original value, and QSGD uses a fixed bit width to discretize the gradient into a finite value set; low-rank decomposition methods such as PowerSGD approximate the gradient tensor as the product of low-dimensional matrices through singular value decomposition, thereby reducing the communication volume; differential coding methods such as EF-SGD utilize the temporal correlation of gradients, only transmitting the difference between the current round and the previous round, and combining an error feedback mechanism to compensate for compression loss.
[0004] Although these techniques alleviate the communication bottleneck to some extent, there are still significant limitations in balancing high compression rate and model performance. For example, Top-k or random sparsification methods lack adaptability, which may discard key small-amplitude gradients or introduce excessive noise, resulting in decreased convergence speed or impaired performance; fixed-bit-width quantization cannot adapt to changes in gradient distribution, and low-bit quantization causes information loss, especially when high precision is required during early training, while high-bit quantization is insufficient in compression rate; low-rank decomposition has high computational overhead and is not suitable for resource-constrained devices, and the low-rank approximation may not capture the complex structure of the gradient, affecting convergence. In addition, existing methods lack systematic multi-level integration, and the trade-off between compression rate and information preservation is not ideal, making it difficult to meet the needs of diverse distributed learning scenarios.
[0005] Therefore, the present application is proposed. SUMMARY
[0006] The present application aims to provide a gradient compressor and a gradient compression method and device based on information consistency driving to solve the problem of excessive communication overhead in distributed deep learning in the prior art, significantly reduce the communication data volume between the client and the server through multi-level optimization, and maximize the preservation of gradient information to ensure that the distributed training performance is comparable to centralized training.
[0007] To solve the above technical problems, the present application realizes the following technical scheme: A gradient compressor based on information consistency driving is used for gradient compression between a plurality of clients and a server in a federated learning system, comprising a mutual information constraint modeling module, a multi-scale sparse representation module, and a Bayesian adaptive quantization module. The mutual information constraint modeling module is used to receive the original gradient generated by the model of the client in the training process, and the mutual information constraint is used as the core target of gradient compression, and the optimal compression operator is obtained by minimizing the Lagrange objective function. The multi-scale sparse representation module is used to perform multi-scale sparse processing on the original gradient using wavelet transform to reduce redundant data and obtain a sparse gradient representation, and the mutual information constraint and the optimal compression operator are combined for evaluation to dynamically adjust the sparse parameter. The Bayesian adaptive quantization module is used to estimate the probability distribution of the sparse gradient representation by Bayesian method, dynamically adjust the quantization bit width according to the probability distribution of the gradient, generate compressed gradient, and transmit it back to the server end after decompression and aggregation, and then issue it to the client for client model update.
[0008] Preferably, in the mutual information constraint modeling module, the mutual information is the difference between the entropy and the conditional entropy, and the mutual information constraint is used as the core target of compression, and the expression is: ; ; Wherein, represents the original gradient; represents the reconstructed gradient, i.e. the approximate gradient after decompression; represents the compressed representation, which is the target of transmission between the client and the server; represents the constraint; represents the compression operator, i.e. the mapping function from the original gradient to the compressed representation ; represents the expectation operator, i.e. taking the average value on the data distribution; represents the original gradient L2 norm square between the reconstructed gradient , for measuring the reconstruction error; denotes L2 norm; denotes mutual information, i.e. the information about the original gradient preserved in the compressed representation; denotes information preservation threshold; denotes entropy, i.e. the overall uncertainty of the original gradient; denotes conditional entropy, i.e. the uncertainty of the gradient still existing given the compressed representation.
[0009] Preferably, the constrained optimization problem is transformed into unconstrained form by the Lagrangian objective function, and a balance term is introduced to flexibly adjust the importance of reconstruction error and information preservation, whose expression is: ; wherein, denotes Lagrangian objective function; denotes Lagrangian multiplier, for balancing the error term and the information term; denotes expected reconstruction error term; denotes information difference value.
[0010] Preferably, the expression of the optimal compression operator is: ; wherein, denotes optimal compression operator; denotes minimization operation; denotes Lagrangian objective function; denotes compression operator.
[0011] Preferably, when the original gradient is processed for multi-scale sparsification by wavelet transform, the original gradient is decomposed into low-frequency and high-frequency parts, and sparsification is performed in multi-scale space, whose expression is: ; wherein, denotes original gradient; denotes low-frequency coefficient, i.e. global trend; denotes scale function, for representing the low-frequency part; denotes high-frequency coefficient, i.e. local detail; denotes wavelet basis function, for representing the high-frequency part; , denote scale, position index respectively; At the same time, an energy preservation constraint is introduced to ensure that the low-frequency part and part of the high-frequency components can preserve most of the original signal energy, whose expression is: ; wherein, denotes the reserved high frequency set, denotes the energy keeping ratio; denotes the L2 norm; In the sparsification process, the high frequency components are selectively reserved, and the expression is: ; wherein, denotes the sparsified high frequency coefficient; so as to obtain the sparsified gradient representation : ; wherein, denotes the reserved low frequency coefficient set; denotes the sparsified high frequency set.
[0012] Preferably, the evaluation is combined with the optimal compression operator to dynamically adjust the sparsification parameter, specifically: calculate the mutual information of the original gradient and the sparsified gradient representation; combine the mutual information constraint and the Lagrange objective function to solve the optimal compression operator; when the mutual information and the optimal compression operator do not meet the set target, adjust the threshold of wavelet transform or the decomposition layer number.
[0013] Preferably, the probability distribution of the gradient is obtained through probability modeling, and the expression is: ; wherein, denotes the probability distribution function of the input gradient value ; denotes the number of components of the Gaussian mixture model; denotes the weight of the th Gaussian component, satisfying ; denotes the mean value of the th component, i.e. the center position; denotes the variance of the th component, i.e. the distribution expansion degree; denotes the Gaussian density function; denotes the probability density of the gradient value mean value , variance .
[0014] Preferably, when dynamically adjusting the quantization bit width according to the gradient distribution, the optimal bit width is selected by maximum a posteriori inference to find the number of bits that can maximize the posterior probability, so as to ensure that the error in the quantization process is minimized, and the expression is: ; Wherein, represents the optimal bit width, i.e., the number of bits finally selected; represents finding the bit width value that maximizes the posterior probability; b represents the bit width; represents the gradient value in the sparse gradient representation sample set; represents the gradient value in the sparse gradient representation sample set; represents the posterior probability of the bit width b given the gradient sample set . Next, based on the standard deviation of the gradient and the bit width size, a scaling factor is calculated to map the floating-point number to the integer space, and the expression is: ; Wherein, represents the scaling factor for adjusting the mapping ratio of the floating-point number to the integer domain; represents the standard deviation of the gradient distribution; Finally, in the quantization process, each gradient value is mapped to an integer, and the error is reduced by rounding, and the expression is: ; Wherein, represents the gradient value quantized integer value, i.e., compressed gradient; represents the gradient mean value for centering processing; represents the rounding operation.
[0015] The application also provides a gradient compression method based on information consistency driving, comprising: receiving the original gradient generated by the model of the client in the training process; constructing mutual information constraints based on the original gradient and the reconstructed gradient, and obtaining the optimal compression operator by minimizing the Lagrange objective function based on the mutual information constraints; based on the original gradient, using wavelet transform to perform multi-scale sparse processing on the original gradient to reduce redundant data, to obtain a sparse gradient representation, and combining the mutual information constraints and the optimal compression operator to evaluate to dynamically adjust the sparse parameter; estimate the probability distribution of the sparse gradient representation by Bayesian method, dynamically adjust the quantization bit width according to the probability distribution of the gradient, generate compressed gradient, and transmit back to the server side after decompression and aggregation, and then issue to the client for client model update.
[0016] The application further provides a gradient compression device based on information consistency driving, comprising a processor and a memory, wherein the memory stores a computer program which can be executed by the processor to implement the gradient compression method based on information consistency driving.
[0017] The application further provides a computer readable storage medium which stores computer readable instructions, and the computer readable instructions are executed by a processor of a device where the computer readable storage medium is located to implement the gradient compression method based on information consistency driving.
[0018] In summary, compared with the prior art, the application has the following beneficial effects: The application comprises an inter-information constraint modeling module, a multi-scale sparse representation module and a Bayesian adaptive quantization module, and significantly reduces the communication overhead in distributed deep learning through multi-level optimization while maximizing the gradient information.
[0019] Specifically, the inter-information constraint modeling module defines an information bottleneck optimization target to ensure that the compression result is not only numerically approximate but also maintains strong correlation with the original gradient in the information level, avoiding the risk of information loss in traditional compression methods. The multi-scale sparse representation module uses wavelet decomposition method to sparsify at different scales, effectively capturing the global trend and local details of the gradient, and significantly reducing the communication volume. The Bayesian adaptive quantization module automatically selects the optimal bit width according to the dynamic characteristics of the gradient distribution, using a higher bit width in the early training stage to ensure that information is not lost, and automatically reducing the bit width in the later training stage to improve the compression rate.
[0020] The gradient compressor based on information consistency driving has high efficiency, universality and robustness, and can be widely applied to federated learning, edge computing and other scenarios. In particular, the application significantly reduces the communication overhead through multi-level optimization while ensuring the integrity of the gradient information, solving the problem that the compression rate and model performance are difficult to balance in the prior art. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the application, and therefore should not be regarded as a limitation to the scope, and for those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0022] Figure 1 A framework schematic diagram of the gradient compressor based on information consistency driving provided for the first embodiment.
[0023] Figure 2 This is a flowchart of a gradient compression method based on information consistency driven by Example 2.
[0024] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0026] Example 1 like Figure 1 As shown, an information consistency-driven gradient compressor is used for gradient compression between several clients and servers in a federated learning system. It is characterized by including: a mutual information constraint modeling module, a multi-scale sparse representation module, and a Bayesian adaptive quantization module.
[0027] like Figure 1 As shown, in federated learning, each client (client1, ... clientn) uses its local private data to perform the training process on the client model (forward propagation to calculate the loss, back propagation to solve the parameter gradient), generating the original gradient (i.e. the gradient of the model parameters, reflecting the direction and magnitude of the parameter update), which is then compressed by the gradient compressor and uploaded to the server.
[0028] The gradient compressor framework consists of three modules: mutual information constraint modeling, multi-scale sparse representation, and Bayesian adaptive quantization. First, the mutual information constraint modeling module defines the overall optimization objective of the compression process from an information theory perspective, ensuring that the compression result retains key information. Then, the multi-scale sparse representation module reduces redundant data in the frequency and spatial domains, making transmission more efficient. Next, the Bayesian adaptive quantization module dynamically models the gradient distribution and adaptively selects the optimal bit width, achieving a balance between compression ratio and accuracy. These three modules simultaneously optimize communication efficiency and training accuracy.
[0029] In the compression task, simply minimizing numerical errors is not enough to guarantee the effectiveness of model training. Even if two gradients are very close in value, if the amount of information is insufficient, the update direction of the model may be biased. Therefore, we introduce mutual information constraints to maintain information as the core goal of compression. The mutual information constraint modeling module is used to receive the original gradient generated by the client's model during the training process, and take the mutual information constraint as the core goal of gradient compression, and obtain the optimal compression operator by minimizing the Lagrangian objective function.
[0030] Specifically, in the mutual information constraint modeling module, the mutual information is the difference between entropy and conditional entropy, and the mutual information constraint is taken as the core goal of compression, which is used to measure the information correlation between the original gradient and the compressed representation. The higher the mutual information, the less the information loss in the compression process. Its expression is: ; ; Where, represents the original gradient; represents the reconstructed gradient, i.e. the approximate gradient after decompression; represents the compressed representation, which is the target of transmission between the client and the server; represents the constraint; represents the compression operator, i.e. the mapping function from the original gradient to the compressed representation ; represents the expectation operator, i.e. taking the average value on the data distribution; represents the L2 norm square of the original gradient and the reconstructed gradient , which is used to measure the reconstruction error; represents the L2 norm; represents the mutual information, i.e. the information about the original gradient preserved in the compressed representation; represents the information retention threshold; represents the entropy, i.e. the overall uncertainty of the original gradient; represents the conditional entropy, i.e. the uncertainty of the gradient still existing after the given compressed representation.
[0031] Mutual information reveals the ability of the compressed representation to reduce uncertainty. In other words, if the compression result can greatly reduce the uncertainty of the original gradient, it means that the information is better preserved. This definition not only facilitates theoretical analysis, but also provides a measurement standard for subsequent quantization strategies.
[0032] In this way, the compressed representation not only guarantees numerical approximation, but also maintains strong correlation with the original gradient at the information level. This module provides a theoretical basis for subsequent sparsification and quantization, making the entire compression process have an information-theoretic interpretation.
[0033] Since the constrained optimization problem is difficult to solve directly, the constrained optimization problem is converted into an unconstrained form by the Lagrange objective function, and a balance term is introduced to flexibly adjust the importance of reconstruction error and information preservation, and its expression is: ; wherein, represents the Lagrange objective function; represents the Lagrange multiplier, which is used to balance the error term and the information term; represents the expected reconstruction error term; represents the information difference.
[0034] This transformation not only simplifies the optimization process, but also adapts to the needs of different training stages, achieving dynamic adjustment.
[0035] Finally, the optimal compression operator is obtained by minimizing the Lagrange objective function, and its expression is: ; wherein, represents the optimal compression operator; represents the minimization operation; represents the Lagrange objective function; represents the compression operator.
[0036] This compression operator can balance between numerical error and information preservation, ensuring that the compression process is both efficient and reliable.
[0037] The mutual information constraint provides a theoretical framework for the compression task, and the multi-scale sparse representation is the first step to realize compression. Gradient signals often contain both global trends and local details, and if only compressed at a single scale, key information may be lost. Therefore, a multi-scale sparse representation module is constructed.
[0038] The multi-scale sparse representation module is used to perform multi-scale sparse processing on the original gradient using wavelet transform based on the original gradient to reduce redundant data and obtain a sparse gradient representation, and the sparse gradient representation is evaluated in combination with the mutual information constraint and the optimal compression operator to dynamically adjust the sparsification parameters.
[0039] When performing multi-scale sparse processing on the original gradient using wavelet transform, the original gradient is decomposed into low-frequency and high-frequency parts, and sparse processing is performed in the multi-scale space, and its expression is: ; wherein, denotes the original gradient; denotes the low-frequency coefficients, i.e., the global trend; denotes the scale function, used to represent the low-frequency part; denotes the high-frequency coefficients, i.e., the local details; denotes the wavelet basis function, used to represent the high-frequency part; , denote the scale and position index, respectively.
[0040] In this way, a large amount of redundant data in the original gradient is reduced, the proportion of non-zero elements is reduced, and the most important components can be selectively retained, so that the information integrity is maintained while reducing the amount of data.
[0041] At the same time, in order to ensure that too much energy is not lost after compression, an energy preservation constraint is introduced to ensure that the low-frequency part and part of the high-frequency components can retain most of the original signal energy, thereby avoiding the decline in training performance, and its expression is: ; wherein, denotes the set of retained high-frequency components, denotes the energy preservation ratio; denotes the L2 norm; In the sparse process, the high-frequency components are selectively retained, and the expression is: ; wherein, denotes the sparse high-frequency coefficients; so as to obtain the sparse gradient representation : ; wherein, denotes the set of retained low-frequency coefficients; denotes the set of sparse high-frequency components.
[0042] The sparse gradient representation is evaluated in combination with the optimal compression operator to dynamically adjust the sparse parameter, specifically: calculate the mutual information of the original gradient and the sparse gradient representation; combine the mutual information constraint and the Lagrange objective function to solve the optimal compression operator; when the mutual information and the optimal compression operator do not meet the set target, adjust the threshold or the decomposition level of the wavelet transform.
[0043] For example, if the mutual information is lower than a preset threshold (indicating excessive information loss), the threshold of the wavelet transform is lowered to retain more gradient components; if the optimal compression operator does not meet expectations (indicating insufficient redundancy reduction), the threshold is raised to further sparsify the gradient. Through iterative adjustments, the sparse gradient representation satisfies both the compression ratio requirement and ensures information consistency.
[0044] After completing the multi-scale sparse representation, the dimensionality and redundancy of the gradient have been significantly reduced, but further reduction of the communication data volume is still needed. Traditional quantization methods typically use a fixed bit width, such as 8-bit or 16-bit representation, which is not necessarily optimal at different training stages. In the early stages of training, the gradient amplitude is large and the distribution is complex, and a higher bit width helps to preserve information; however, in the later stages of training, the gradient gradually converges, and an excessively high bit width will lead to resource waste. Therefore, we propose a Bayesian adaptive quantization module that dynamically selects the optimal bit width and automatically adjusts the quantization strategy.
[0045] The Bayesian adaptive quantization module is used to estimate the probability distribution (such as Gaussian or Laplace distribution) of the sparse gradient representation based on the Gaussian mixture model (GMM) using Bayesian methods (such as Markov chain Monte Carlo method). The quantization bit width is dynamically adjusted according to the probability distribution of the gradient to generate a compressed gradient, which is transmitted back to the server for decompression and aggregation before being sent to the client for model updates.
[0046] The probability distribution of the gradient is obtained through probability modeling, and its expression is: ; in, Represents the gradient value of the input. The probability distribution function; This represents the number of components in a Gaussian mixture model; Indicates the first The weights of the Gaussian components satisfy the following conditions: ; Indicates the first The mean of all components, i.e., the center position; Indicates the first The variance of each component, i.e., the extent of the distribution spread; Represents the Gaussian density function; Represents gradient value mean ,variance The probability density is given below.
[0047] When dynamically adjusting the quantization bit width based on the gradient distribution, the optimal bit width is selected by inferring the maximum a posteriori probability to find the number of bits that maximizes the a posteriori probability, ensuring that the error is minimized during quantization. The expression is: ; where, represents the optimal bit-width, i.e., the final selected number of bits; represents finding the bit-width value that maximizes the posterior probability; b represents the bit-width; represents the -th gradient value in the set of sparsified gradient representation samples; represents the set of gradient samples given the bit-width b.
[0048] Next, a scaling factor is calculated based on the standard deviation of the gradients and the bit-width size to map the floating-point numbers to the integer space, expressed as: ; where, represents the scaling factor used to adjust the mapping ratio of floating-point numbers to the integer domain; represents the standard deviation of the gradient distribution.
[0049] The scaling factor depends on the standard deviation of the gradients and the bit-width size, thus ensuring that the quantized values will not overflow.
[0050] Finally, each gradient value is mapped to an integer during the quantization process, reducing errors through rounding, expressed as: ; where, represents the gradient value quantized integer value, i.e., the compressed gradient; represents the gradient mean value, used for centering processing; represents the rounding operation.
[0051] During the quantization process, each gradient value is mapped to an integer, and rounding is used to reduce errors. The quantized result will be transmitted to the server, along with the scaling factor and mean value information for decoding.
[0052] The core idea of the Bayesian adaptive quantization module is to obtain gradient distribution information through probability modeling, and then infer the best quantization parameter according to the posterior probability, thus achieving a dynamic balance between compression rate and accuracy.
[0053] The compressed gradient is transmitted to the server, and the server performs the following operations: Decompression: According to the rules during compression (such as wavelet inverse transform, quantization decoding), the compressed gradient is restored to the form of "approximate original gradient".
[0054] Aggregation: Using aggregation algorithms such as Federated Averaging (FedAvg) algorithm, the decompressed gradients of each client are aggregated into global gradients (reflecting the comprehensive update requirements of all client data on the model).
[0055] The server distributes the aggregated global gradients to each client, which then updates its local model parameters based on the global gradients, completing one federated learning iteration. Subsequently, the client can continue the process of "generating original gradients → compressing → uploading → aggregating → updating" based on the updated model until the model converges.
[0056] In summary, compared with existing distributed training compression methods, this invention has significant advantages: First, this scheme models using mutual information constraints, considering not only numerical errors during optimization but also ensuring the preservation of information between the compressed representation and the original gradient. This design avoids the risk of numerical approximation but information loss in traditional compression methods, enabling the model to maintain stable convergence during training.
[0057] Secondly, this invention effectively captures the global trend and local details of gradients through multi-scale sparse representation. By using wavelet transform decomposition to perform sparsification at different scales, it can reduce redundant information while retaining key structural features, thereby significantly reducing communication overhead and ensuring high training accuracy even in low-bandwidth environments.
[0058] Furthermore, this invention employs a Bayesian adaptive quantization strategy, automatically selecting the optimal bit width based on the dynamic characteristics of the gradient distribution. In the early stages of training when the distribution is complex, the method tends to use a higher bit width to ensure no information loss; while in the later stages of training when the distribution stabilizes, the method automatically reduces the bit width to improve the compression ratio. Compared to traditional methods with fixed bit widths, this scheme can achieve an adaptive balance between communication efficiency and accuracy at different training stages.
[0059] Example 2 like Figure 2 As shown, the second embodiment of the present invention also provides a gradient compression method based on information consistency driving, including: S1 receives the original gradients generated by the client's model during training; S2, construct mutual information constraints based on the original gradient and the reconstructed gradient, and obtain the optimal compression operator by minimizing the Lagrange objective function based on the mutual information constraints; S3. Wavelet transform is used to perform multi-scale sparsification on the original gradient to reduce redundant data and obtain a sparse gradient representation. The sparse gradient representation is evaluated in combination with the mutual information constraint and the optimal compression operator to dynamically adjust the sparsification parameters. S4. Estimate the probability distribution of the sparse gradient representation using the Bayesian method, dynamically adjust the quantization bit width according to the probability distribution of the gradient, generate a compressed gradient, transmit it back to the server, decompress and aggregate it, and then send it to the client for client model updates.
[0060] Embodiment three The third embodiment of the present application also provides a gradient compression device based on information consistency driving, comprising a memory and a processor, the memory stores a computer program, and the computer program can be executed by the processor to implement the gradient compression method based on information consistency driving as described above.
[0061] Embodiment four The fourth embodiment of the present application also provides a computer readable storage medium, which stores computer readable instructions, and the computer readable instructions are executed by the processor of the device where the computer readable storage medium is located to implement the gradient compression method based on information consistency driving as described above.
[0062] In several embodiments provided by the embodiments of the present application, it should be understood that the disclosed apparatus and method can also be implemented by other means. The apparatus and method embodiments described above are only illustrative, for example, the flowchart in the accompanying drawings shows the possible implementation architecture, function and operation of the apparatus, method and computer program product according to the embodiments of the present application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment or a part of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in different order from that shown in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and sometimes they can be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for executing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0063] In addition, each functional module in the various embodiments of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0064] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or parts of the present application that essentially contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes. It should be noted that in this document, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the processes, methods, articles or devices that include a series of elements not only include those elements, but also include other elements not explicitly listed or inherent to such processes, methods, articles or devices. Without more limitations, the element defined by the statement "includes a" does not exclude the presence of other identical elements in the process, method, article or device that includes the element.
[0065] The terms used in the embodiments of the present application are merely for the purpose of describing particular embodiments and are not intended to limit the present application. The singular forms "a", "an" and "the" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0066] It should be understood that the term "and / or" used herein is merely a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the three cases of A alone, A and B together, and B alone. In addition, the character " / " in this document generally represents an "or" relationship between the front and rear associated objects.
[0067] Depending on the context, the word "if" as used herein can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if it is determined" or "if (a stated condition or event) is detected" can be interpreted as "when it is determined" or "in response to determining" or "when (a stated condition or event) is detected" or "in response to detecting (a stated condition or event)".
[0068] The "first / second" mentioned in the embodiments are only to distinguish similar objects, and do not represent a specific order for the objects. Understandably, the "first / second" can be interchanged in a specific order or sequence as appropriate. It should be understood that the objects distinguished by "first / second" can be interchanged as appropriate, so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.
[0069] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Various modifications and changes can be made to the present application by those skilled in the art. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A gradient compressor based on information consistency driven, used for gradient compression between a plurality of clients and a server in a federated learning system, characterized in that, The method comprises the following steps: The mutual information constraint modeling module is used to receive the original gradient generated by the model of the client in the training process, and the mutual information constraint is taken as the core target of gradient compression, and the optimal compression operator is obtained by minimizing the Lagrange objective function; The multi-scale sparse representation module is used to perform multi-scale sparse processing on the original gradient by using wavelet transform to reduce redundant data, and obtain a sparse gradient representation, and the mutual information constraint and the optimal compression operator are combined for evaluation to dynamically adjust the sparse parameter; The Bayesian adaptive quantization module is used to estimate the probability distribution of the sparse gradient representation by using the Bayesian method, dynamically adjust the quantization bit width according to the probability distribution of the gradient, generate a compressed gradient, and transmit the compressed gradient back to the server side, which is decompressed, aggregated and then sent to the client side for model updating. In the mutual information constraint modeling module, the mutual information is the difference between the entropy and the conditional entropy, and the mutual information constraint is taken as the core target of compression, and the expression is:
2. The gradient compressor based on information consistency driving according to claim 1, wherein The Lagrange objective function is used to convert the constrained optimization problem into an unconstrained form, and a balance term is introduced to flexibly adjust the importance of the reconstruction error and the information retention, and the expression is: ; ; wherein, denotes the original gradient; denotes the reconstructed gradient, i.e. the decompressed approximated gradient; denotes the compressed representation, which is the target of the transmission between the client and the server; denotes constrained to; denotes the compression operator, i.e. the mapping function from the original gradient to the compressed representation ; denotes the expectation operator, i.e. the averaging over the data distribution; denotes the original gradient and the reconstructed gradient , used to measure the reconstruction error; denotes the L2 norm; represents mutual information, i.e. information about the original gradient that is preserved in the compressed representation; represents an information retention threshold; represents entropy, i.e. overall uncertainty of the original gradient; represents conditional entropy, i.e. uncertainty that still exists about the gradient given the compressed representation.
3. The gradient compressor based on information consistency driving according to claim 2, characterized in that The expression of the optimal compression operator is: ; wherein, represents a Lagrangian objective function; represents a Lagrangian multiplier for balancing the error term and the information term; represents an expected reconstruction error term; represents an information difference.
4. The gradient compressor based on information consistency driving according to claim 3, characterized in that When the original gradient is processed by using the wavelet transform for multi-scale sparse processing, the original gradient is decomposed into a low-frequency part and a high-frequency part, and the sparse processing is performed in the multi-scale space, and the expression is: ; wherein, denotes an optimal compression operator; denotes a minimization operation; denotes a Lagrangian objective function; denotes a compression operator.
5. The gradient compressor based on information consistency driving according to claim 1, characterized in that At the same time, the energy retention constraint is introduced to ensure that the low-frequency part and part of the high-frequency components can retain most of the original signal energy, and the expression is: ; wherein, denotes the original gradient; denotes the low frequency coefficients, i.e. the global trend; denotes the scale function, used to represent the low frequency part; denotes the high frequency coefficients, i.e. the local details; denotes the wavelet basis function, used to represent the high frequency part; , denote scale, position index, respectively; In the sparse process, the high-frequency components are selectively retained, and the expression is: ; wherein, represents a reserved high frequency set, represents an energy holding ratio; represents an L2 norm; The mutual information between the original gradient and the sparse gradient representation is calculated. ; wherein, denotes the sparse high frequency coefficients; Thus, the sparse gradient representation is obtained : ; wherein, represents a set of reserved low frequency coefficients; represents a set of high frequencies after sparsification.
6. The gradient compressor based on information consistency driving according to claim 4, characterized in that The optimal compression operator is solved by combining the mutual information constraint and the Lagrange objective function. When the solved mutual information and the optimal compression operator do not meet the set target, the threshold value or the decomposition layer number of the wavelet transform is adjusted. The probability distribution of the gradient is obtained by probability modeling, and the expression is: When the quantization bit width is dynamically adjusted according to the gradient distribution, the optimal bit width is selected by maximum a posteriori probability inference to find the bit number that can maximize the posterior probability, so as to ensure that the error in the quantization process is minimized, and the expression is:
7. The gradient compressor based on information consistency driving according to claim 1, characterized in that Then, the scaling factor is calculated based on the standard deviation of the gradient and the bit width size to map the floating-point number to the integer space, and the expression is: ; wherein, denotes the input gradient value a probability distribution function; denotes the number of components of the Gaussian mixture model; denotes the weight of the th Gaussian component, satisfying ; denotes the mean, i.e. the center position, of the th component; denotes the variance, i.e. the spread of the distribution, of the th component; denotes the Gaussian density function; denotes the probability density under the gradient value mean , variance .
8. The gradient compressor based on information consistency driving according to claim 7, characterized in that Finally, each gradient value is mapped to an integer in the quantization process, and the error is reduced by rounding, and the expression is: ; in, This indicates the optimal bit width, i.e., the final number of bits selected; This indicates the search for the bit width value that maximizes the posterior probability; b represents the bit width. The sparsified gradient represents the gradient of the th sample in the sample set. One gradient value; Represents a given set of gradient samples Below, the posterior probability of bit width b; The method comprises the following steps: ; wherein, denotes a scaling factor, used to adjust the mapping scale of floating point numbers to the integer domain; denotes the standard deviation of the gradient distribution; Receiving the original gradient generated by the model of the client in the training process; ; wherein, denotes the gradient value quantized integer value, i.e. compressed gradient; denotes the gradient mean value, for centering; denotes the rounding operation.
9. A gradient compression method based on information consistency driving for federated learning, characterized in that Based on the original gradient and the reconstructed gradient, a mutual information constraint is constructed, and based on the mutual information constraint, an optimal compression operator is obtained by minimizing a Lagrange objective function; Based on the original gradient, the original gradient is processed by wavelet transform for multi-scale sparsification to reduce redundant data, to obtain a sparsified gradient representation, and the mutual information constraint is combined with the optimal compression operator for evaluation to dynamically adjust the sparsification parameter; The probability distribution of the sparsified gradient representation is estimated by a Bayesian method, the quantization bit width is dynamically adjusted according to the probability distribution of the gradient, the compressed gradient is generated, transmitted back to the server end, decompressed and aggregated, and then issued to the client end for client model updating.
10. A gradient compression apparatus driven based on information consistency, characterized in that, The method comprises a processor and a memory, wherein the memory stores a computer program which can be executed by the processor to implement the gradient compression method based on information consistency driving according to claim 9.
Citation Information
Patent Citations
Efficient federated learning method based on adaptive quantization and dynamic gradient coding
CN120146225A
Frequency difference compressor based on precision perception, gradient compression method, equipment and medium
CN120258058A
Distributed artificial intelligence system using transmission of compressed gradients and model parameter, and learning apparatus and method therefor
US20230196205A1