Gradient low-rank compression modeling method and system based on information entropy

By constructing a gradient eigenvalue distribution model and introducing information entropy, the compression rank is dynamically adjusted, which solves the problem of insufficient adaptability of existing low-rank compression methods in dynamic environments, achieves efficient gradient low-rank compression, and improves the stability of model training and communication efficiency.

CN120670700APending Publication Date: 2025-09-19SHANGHAI JIAOTONG UNIV +1

Patent Information

Application Number
CN202510778205.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing low-rank compression methods lack unified theoretical support and are difficult to adapt to the dynamic changes of gradients, resulting in loss of key information or excessive redundancy, affecting the accuracy and efficiency of model training, especially in complex model structures or diverse training environments, resulting in insufficient generalization and stability.

Method used

A gradient low-rank compression modeling method based on information entropy is constructed. By building a gradient eigenvalue distribution model, a functional relationship between compression error and compression rank is established. Marcenko-Pastur distribution and standard deviation are introduced as intermediaries. The compression rank is dynamically adjusted to adapt to the changes in gradient information entropy, thereby realizing a refined compression strategy.

Benefits of technology

It improves the adaptability and robustness of the compression strategy, enhances the stability and accuracy of model training, and maximizes the benefits of communication compression. It is versatile and scalable, and suitable for different hardware environments and training tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670700A_ABST
    Figure CN120670700A_ABST
Patent Text Reader

Abstract

The invention provides a theoretical modeling method and system for describing the relationship between the gradient entropy and the compression rank, and provides a theoretical basis for gradient compression strategy design and adaptive communication optimization. According to the method, modeling is carried out on gradient matrix eigenvalue distribution, internal relation between eigenvalues and compression errors is revealed by applying a random matrix theory and characteristic spectrum analysis, and a function relation between a compression rank and a reconstruction error is constructed. On the basis, a constraint condition that a compression error absolute value is kept constant is introduced, the standard deviation of the gradient matrix is incorporated into error estimation, a corresponding relation between the standard deviation and a compression rank is derived, and the effect of the standard deviation serving as a compression error intermediate variable is clarified from the statistical angle. According to the method, the minimum compressible rank under specific precision can be estimated, a unified analysis framework is provided for analyzing compressibility differences of different training stages or network layer gradients, and the method has good theoretical universality and engineering practicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of distributed model training in deep learning, and specifically to a gradient low-rank compression modeling method and system based on information entropy. Background Art

[0002] With the widespread application of large-scale language models (LLMs) in various fields, the number of model parameters and training scale has shown exponential growth. To support model training at these parameter levels, it is often necessary to build distributed training clusters containing hundreds to thousands of GPUs. In this process, the gradient synchronization mechanism under the data parallel training architecture has gradually become a performance bottleneck. Each round of training iteration requires the full gradient exchange between all computing nodes. This full gradient synchronization mechanism not only consumes a large amount of network bandwidth resources, but also significantly reduces system throughput and exacerbates the imbalance of resource load between different computing nodes. This problem is particularly prominent in environments with weak communication bandwidth or heterogeneous hardware clusters.

[0003] Gradient compression technology, as the core solution to alleviate the above problems, effectively reduces the amount of communication data by compressing the representation of the gradient tensor, thereby improving training efficiency and reducing energy consumption. By reducing the amount of gradient data that needs to be transmitted in each iteration, gradient compression technology can not only reduce the communication burden, but also effectively reduce energy consumption and accelerate the training process. The current mainstream gradient compression methods are mainly divided into three categories: gradient sparsification, gradient quantization, and matrix low-rank decomposition: gradient sparsification reduces the amount of data by selectively transmitting gradient elements with larger amplitudes and accumulating the rest in a local cache; gradient quantization converts high-precision floating-point gradients into low-bit width data (such as 8-bit or 1-bit) to reduce bandwidth occupancy; matrix low-rank decomposition technology, based on the low-rank characteristics of the gradient matrix, decomposes it into a combination of matrices with smaller dimensions, reducing the total amount of parameter transmission. Among them, low-rank compression methods have become a research hotspot due to their advantages in both theory and practice, including high compression rate and low reconstruction error.

[0004] However, existing low-rank compression methods suffer from significant drawbacks: most employ fixed rank settings or dynamically adjust rank values ​​based on heuristic methods, but lack a unified theoretical framework. These artificially pre-set compression parameters struggle to adapt to the dynamic nature of gradients, leading to loss of critical information or excessively redundant compression, ultimately compromising model training accuracy and compression efficiency. The generalizability and stability of these methods in complex model structures or diverse training environments remain to be verified.

[0005] Although recent research has attempted to introduce entropy theory to assess the information complexity of gradient tensors, positing that high entropy indicates a rich gradient containing valuable information and greater compression difficulty, while low entropy indicates a high proportion of redundant information and greater compression potential, this research remains in its early stages, relying primarily on empirical or heuristic approaches and lacking a systematic theoretical model or reusable framework.

[0006] A search of patent documents revealed an invention patent with publication number CN117910521A, which discloses a gradient compression method, apparatus, device, distributed cluster, and storage medium. This invention belongs to the field of distributed computing and is used to adjust the degree of gradient compression based on two indicators: the model performance optimization rate and the current single-step training duration. This solves the problem of being unable to balance model performance and communication overhead when performing gradient compression on low-speed networks. The invention uses a single training step as the granularity. After obtaining gradient data at any training step after the warm-up phase, if the model performance optimization rate does not meet the standard, the gradient compression degree is reduced to improve model performance. If the model performance optimization rate meets the standard and the current single-step training duration exceeds the standard, the gradient compression degree can be amplified to reduce communication overhead. The invention can dynamically adjust the degree of gradient data compression based on the influence of network conditions, thereby minimizing communication overhead while taking into account both model performance and network conditions. This patent primarily relies on training indicators to dynamically adjust the degree of compression, without involving information entropy, random matrix theory, and eigenvalue distribution modeling. The theoretical derivation of compression rank and error is not in-depth, and its generality is insufficient.

[0007] To sum up, in response to the above-mentioned problems of the existing technology, studying a gradient low-rank compression modeling method and system based on information entropy has become a key task that needs to be solved urgently. For the first time, a functional relationship model between gradient information entropy and compression rank is proposed theoretically. Summary of the Invention

[0008] In view of the defects in the prior art, the purpose of the present invention is to provide a gradient low-rank compression modeling method and system based on information entropy.

[0009] According to the present invention, a gradient low-rank compression modeling method based on information entropy is provided, which includes the following sub-steps:

[0010] Step S1, constructing a distribution model of gradient eigenvalues;

[0011] Step S2, constructing a functional relationship between compression error and compression rank;

[0012] Step S3, constructing a cumulative distribution function that conforms to the Marcenko-Pastur distribution;

[0013] Step S4: Based on the functional relationships between the cumulative distribution function, the compression error, and the compression rank, and the distribution model, establish the mathematical relationship between the compression error and the compression rank; Step S5: Based on the mathematical relationship between the compression error and the compression rank, establish the mathematical relationship between the standard deviation and the compression rank;

[0014] Step S6: Construct the mathematical relationship between the gradient entropy and the standard deviation;

[0015] Step S7: Based on the mathematical relationship between the standard deviation and the compression rank and the mathematical relationship between the gradient entropy and the standard deviation, construct the mapping expression between the gradient entropy and the compression rank.

[0016] Preferably, in Step S1, let G be the probability distribution on, the mean of G is 0, and the variance is 1, denote the set of real numbers, let Y be a p×n-dimensional random matrix, the elements of Y are independently and identically distributed according to G, define the matrix S = n -1 YY T , T is the matrix transpose symbol. When p, n → ∞ and satisfy p / n → γ ∈ (0, 1), the empirical spectral distribution F S of the matrix S converges to the Marcenko-Pastur distribution in the sense of weak convergence, and the probability density function is:

[0017]

[0018] where: γ represents p / n,

[0019] x represents the eigenvalue variable of the spectral distribution, and y represents the subscript of the function symbol, not a variable.

[0020] Preferably, in Step S2, let the matrix have a singular value decomposition, that is, A = U∑V T , let r = rank(A) be the compression rank of A, take k < r, and define the truncated matrix

[0021] where:

[0022] m represents the number of rows of the matrix A, n represents the number of columns of the matrix A, U represents the left singular vector matrix, V represents the right singular vector matrix. r represents the rank of the matrix A, the number of non-zero singular values, rank() is synonymous with r, representing the matrix rank, k represents the rank retained after truncation, controlling the compression degree.

[0023] σ i is the singular value of A, u i and v i are the corresponding left and right singular vectors respectively;

[0024] For all compressed matrices B of rank k, the minimum error under the 2-norm is given by A k gives:

[0025] min rank(B)=k ‖AB‖2=‖AA k ‖2=σ k+1 Formula (2)

[0026] Similarly, for the Frobenius norm, the following conclusion holds:

[0027]

[0028] Among them, ‖‖2 represents the spectral norm of the matrix, which is equal to its maximum singular value and is used to measure the maximum error of the matrix in the sense of 2-norm;

[0029] ‖‖ F It represents the Frobenius norm (F is the first letter of Frobenius), which is the square root of the sum of the squares of all matrix elements and is used to measure the overall approximation error.

[0030] Preferably, in step S3, is a random matrix, and the standard deviation of each element in A is 1, then AA T The cumulative distribution function corresponding to the eigenvalue λ is expressed as:

[0031]

[0032] Preferably, in step S4, the random matrix The elements are independent and identically distributed, with a mean of 0 and a variance of 1, denoted by A r Estimate the compression error by finding the optimal low-rank approximation of A with a compression rank of r

[0033] Preferably, the compression error is estimated The following sub-steps are included:

[0034] Step S4.1, determine the eigenvalue range [a, b] according to step S1, and uniformly sample the eigenvalue candidate set {λ0} within the eigenvalue range;

[0035] Step S4.2, based on the cumulative distribution function F(λ0) of the matrix eigenvalues ​​in step S3, calculate the cumulative probability {p0} corresponding to each λ0 and construct a mapping relationship, i.e., a pair {(λ0, p0)};

[0036] Step S4.3: uniformly sample m random probability values ​​{p} from the interval (0, 1), and use the mapping relationship obtained in step S4.2 to obtain the corresponding eigenvalue samples {λ} by interpolation;

[0037] Step S4.4, sum the smallest mr values ​​in the eigenvalue sample {λ} as the compression error Estimated value of:

[0038]

[0039] where λ i AA T The i-th eigenvalue of , in descending order,

[0040] Under the premise of knowing the compression rank r and matrix dimension (m,n), the compression error ∈=‖AA is estimated by sampling r ‖ F , the process is formally expressed as:

[0041] ∈=g(r;m,n) Formula (15)

[0042] Among them, g() means that under the premise of known matrix dimension m×n and compression rank r, based on – The expected value or unbiased approximation of the compression error ∈ obtained by Monte Carlo sampling estimation of the Pastur distribution provides an error estimation path in the absence of the true matrix.

[0043] Preferably, in step S5, a fixed compression error constraint is imposed during the LLM training process, denoted as ∈ ini , and is set when the gradient compression mechanism is first enabled,

[0044] Stochastic gradient matrix The standard deviation of is σ0, and the compression error ∈=‖AA in the whole training process r ‖ F The fixed compression error constraint is always satisfied. When the standard deviation of the gradient matrix changes from σ0 to σ1 during training, the compression rank is adjusted from the original r0 to r1:

[0045]

[0046] Preferably, in step S6, for a normal distribution with a mean of μ and a standard deviation of σ, the information entropy is expressed as:

[0047]

[0048] Assume that the random variable X obeys a standard normal distribution with a mean of 0 and a variance of 1. According to the definition of information entropy, substitute the probability density function f(x) into Equation 22 to obtain:

[0049]

[0050] Here, e represents a natural constant, namely the Euler number; the first integral term corresponds to the probability density function of the normal distribution, and its integral result is 1. For the second integral term, let the variable substitution x = yσ + μ, then dx = σdy, and Equation 23 becomes:

[0051]

[0052] Preferably, in step S7, during the training process, if the entropy value of the gradient matrix a changes from H0 to H1, under the premise of a fixed compression error, its compression rank is adjusted from r0 to r1, satisfying the following relationship:

[0053]

[0054] The present invention also provides a gradient low-rank compression modeling system based on information entropy, comprising:

[0055] Module M1, constructs the distribution model of gradient eigenvalues;

[0056] Module M2, constructs the functional relationship between compression error and compression rank;

[0057] Module M3, constructs the cumulative distribution function that conforms to the Marcenko-Pastur distribution;

[0058] Module M4, establishing a mathematical relationship between the compression error and the compression rank based on the cumulative distribution function, the functional relationship between the compression error and the compression rank, and the distribution model;

[0059] Module M5, establishing a mathematical relationship between the standard deviation and the compression rank based on the mathematical relationship between the compression error and the compression rank;

[0060] Module M6, constructs the mathematical relationship between gradient entropy and standard deviation;

[0061] Module M7 constructs a mapping expression between gradient entropy and compression rank based on the mathematical relationship between standard deviation and compression rank and the mathematical relationship between gradient entropy and standard deviation.

[0062] Compared with the prior art, the present invention has the following beneficial effects:

[0063] 1. This invention represents a significant innovation in the field of low-rank gradient compression. Current mainstream methods rely on fixed rank values ​​or dynamically adjust compression strength during training through heuristic methods. However, due to a lack of systematic modeling of the internal statistical structure of the gradient tensor, these methods struggle to accurately adapt to gradient distribution characteristics at different stages and levels, leading to insufficient or redundant compression and impacting training stability and model accuracy.

[0064] 2. Starting from the essential statistical properties of the gradient tensor, this paper constructs the first quantitative relationship model between compression rank and eigenvalue distribution, compression error, standard deviation, and information entropy. This model is a complete and solvable theoretical modeling framework, shifting the compression rank setting from being driven by experience to being driven by data and theory. By introducing gradient eigenvalue spectrum analysis and standard deviation as error control mediators, the minimum compressible rank of the tensor can be estimated within the target error range, enabling the design of refined compression strategies.

[0065] 3. The present invention introduces information entropy as a measure of global information complexity, making the compression strategy highly sensitive to changes in gradient compressibility. The entropy-guided compression scheduling mechanism enhances the adaptability and robustness of the compression strategy, providing a unified theoretical benchmark for compression regulation, which helps to maximize communication compression benefits while maintaining model accuracy.

[0066] 4. The model of the present invention is universal and extensible, and can be embedded in existing low-rank compression frameworks (such as SVD-based compression methods), supporting cross-model and cross-level migration applications. Its theoretical rank selection mechanism improves the feasibility of compression strategies in different hardware environments and training tasks, laying a theoretical foundation for building an efficient and adaptive large-scale distributed training system. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:

[0068] Figure 1 This is a flowchart of a gradient low-rank compression modeling method based on information entropy in an embodiment of the present invention. DETAILED DESCRIPTION

[0069] The present invention is described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art further understand the present invention, but do not limit the present invention in any way. It should be noted that, for those skilled in the art, several variations and improvements can be made without departing from the scope of the present invention. These are all within the scope of protection of the present invention. The present invention provides a theoretical modeling method and system for describing the relationship between gradient entropy and compression rank, providing a theoretical basis for gradient compression strategy design and adaptive communication optimization. This method aims to systematically characterize the functional relationship between compression rank and multiple statistical properties of the gradient tensor, providing theoretical support for rank value selection during gradient compression. The method first models the distribution of the eigenvalues ​​of the gradient matrix, applies random matrix theory and eigenspectral analysis, reveals the intrinsic connection between eigenvalues ​​and compression error, and constructs a functional relationship between compression rank and reconstruction error. On this basis, a constraint is introduced to maintain a constant absolute value of the compression error, and the standard deviation of the gradient matrix is ​​incorporated into the error estimate. The corresponding relationship between the standard deviation and compression rank is derived, and the role of the standard deviation as a mediating variable for compression error is statistically clarified.

[0070] Furthermore, in view of the changing characteristics of the information entropy of the gradient tensor during training, the present invention establishes a function mapping between gradient entropy and standard deviation, realizes the explicit representation of the compression rank as an information entropy function, reflects the theoretical lower bound of the gradient compressibility, and enables the rank value selection to dynamically adapt to the information complexity of the gradient itself.

[0071] Through the above modeling process, the present invention can not only estimate the minimum compressible rank under a specific accuracy, but also provide a unified analysis framework for analyzing the differences in gradient compressibility at different training stages or network layers, with good theoretical versatility and engineering practicality.

[0072] Example 1:

[0073] Figure 1 Flowchart of a gradient low-rank compression modeling method based on information entropy in an embodiment of the present invention. Figure 1 As shown, this embodiment provides a gradient low-rank compression modeling method based on information entropy, including the following sub-steps:

[0074] Step S1, constructing a distribution model of gradient eigenvalues.

[0075] Specifically, in step S1, let G be A probability distribution on G with mean 0 and variance 1, Represents a set of real numbers, let Y be a p×n dimensional random matrix, the elements of Y are independent and identically distributed in G, and define the matrix S=n -1 YY T , T is the matrix transpose symbol, when p,n→∞ and satisfy p / n→γ∈(0,1), the empirical spectrum distribution F of the matrix SS Approaches the Marchenko - Pastur distribution in a weakly convergent manner, with the probability density function given by:

[0076]

[0077] where: γ represents p / n,

[0078] x represents the eigenvalue variable of the spectral distribution, and y represents the subscript of the function symbol, not a variable.

[0079] Based on the distribution model of gradient eigenvalues, a probability estimate of the eigenvalue truncation error for any compression rank k is carried out, providing a theoretical basis for subsequent modeling.

[0080] Step S2, construct the functional relationship between the compression error and the compression rank.

[0081] Specifically, in step S2, let the matrix Have a singular value decomposition (SVD), that is, A = U∑V T , Let r = rank(A) be the compression rank of A (in the compression method of low - rank decomposition, the compression rank is equivalent to the compression ratio), take k < r, and define the truncated matrix

[0082] where:

[0083] m represents the number of rows of matrix A, n represents the number of columns of matrix A, U represents the left singular vector matrix, V represents the right singular vector matrix. r represents the rank of matrix A, the number of non - zero singular values, rank() is synonymous with r, representing the matrix rank, k represents the rank retained after truncation, controlling the degree of compression.

[0084] σ i Is the singular value of A, u i And v i Are the corresponding left and right singular vectors respectively.

[0085] For all matrices B with compression rank k, the minimum error in the 2 - norm is given by A k As:

[0086] min rank(B)=k ‖A - B‖2 = ‖A - A k ‖2 = σ k+1 Equation (2)

[0087] Similarly, for the Frobenius norm, the conclusion still holds:

[0088]

[0089] where,

[0090] ‖‖2 represents the spectral norm of the matrix, which is equal to its maximum singular value and is used to measure the maximum error of the matrix in the sense of 2-norm;

[0091] ‖‖ F It represents the Frobenius norm (F is the first letter of Frobenius), which is the square root of the sum of the squares of all matrix elements and is used to measure the overall approximation error.

[0092] This estimation formula is directly used to inversely find the lower bound of the compression rank k under specific error constraints.

[0093] Step S3: constructing a cumulative distribution function that conforms to the Marcenko-Pastur distribution.

[0094] Specifically, in step S3, step S1 describes the asymptotic distribution characteristics of the eigenvalues ​​of the large-scale random covariance matrix. is a random matrix, and the standard deviation of each element in A is 1, then AA T The cumulative distribution function (CDF) corresponding to the eigenvalue λ is expressed as:

[0095]

[0096] Proof: Consider the indefinite integral:

[0097]

[0098] Let y = λ - a and c = ba, then dy = dλ, and Equation 5 can be rewritten as:

[0099]

[0100] Using trigonometric substitution, let y = csin 2 t, then dy = 2c sin t cos t dt = c sin 2t dt. Substituting this into Equation 6, the integral becomes:

[0101]

[0102] Further simplifying Formula 7, let c(d-1)=2a, and we get:

[0103]

[0104] Splitting and arranging the numerator of formula 8, the integral becomes:

[0105]

[0106] Where t represents the angle variable introduced in the trigonometric substitution.

[0107] d represents an intermediate parameter introduced by the identity transformation c(d-1)=2a, which is used to simplify the integral expression.

[0108] The first integral is calculated by the universal substitution m = tant, since cos 2t = (1-m 2 ) / (1+m 2 ), dt=dm / (1+m 2 ), the integral is:

[0109]

[0110] According to the integral form of the inverse tangent function, we continue to simplify formula 10 and obtain the first integral:

[0111]

[0112] Substituting Equation 10 into Equation 9 and combining the results of the other two integrals, we obtain:

[0113]

[0114] Finally, substituting y = λ - a, c = ba, and c(d - 1) = 2a into Equation 12 yields the final integral result:

[0115]

[0116] This expression represents the cumulative distribution function of the Marcenko-Pastur distribution, which describes the behavior of the eigenvalue distribution of large random matrices under certain conditions.

[0117] Step S4: establishing a mathematical relationship between the compression error and the compression rank based on the cumulative distribution function, the functional relationship between the compression error and the compression rank, and the distribution model.

[0118] Specifically, in step S4, let the random matrix The elements are independent and identically distributed, with a mean of 0 and a variance of 1, denoted by A r is the optimal low-rank approximation of A with a compression rank of r, then its squared compression error is Estimate as follows:

[0119] Step S4.1, determine the upper and lower bounds a and b of the eigenvalue according to step S1, and uniformly sample the eigenvalue candidate set {λ0} within this range;

[0120] Step S4.2, based on the cumulative distribution function F(λ0) of the matrix eigenvalues ​​in step S3, calculate the cumulative probability {p0} corresponding to each λ0 and construct a mapping relationship, i.e., a pair {(λ0, p0)};

[0121] Step S4.3: uniformly sample m random probability values ​​{p} from the interval (0,1), and use the mapping relationship obtained in step S4.2 to interpolate and obtain the corresponding eigenvalue samples {λ}; S4.1: sample the eigenvalue candidate set {λ0}—> S4.2: calculate the cumulative probability {p0} corresponding to each λ0, and construct the mapping relationship—> S4.3: uniformly sample from {λ0} to obtain the eigenvalue samples {λ}. Therefore, the eigenvalue samples are sampled from the eigenvalue candidate set. And they are not replaced by λ i .

[0122] Step S4.4, sum the smallest mr values ​​in the eigenvalue sample {λ} as the compression error Estimated value of

[0123] Proof: From step S2, let A r is the optimal low-rank approximation of A with a compression rank of r, then its squared reconstruction error measured by the Frobenius norm satisfies the following relationship:

[0124]

[0125] where λ i AA T The i-th eigenvalue of is arranged in descending order. Therefore, the error is equivalent to the matrix AA T The sum of the smallest mr eigenvalues ​​in . According to step S1, the matrix AA T The eigenvalues ​​of obey a fixed distribution, and the cumulative distribution function of this distribution is shown in Formula 13. However, since it is difficult to directly solve the analytical form of the sum of the eigenvalues ​​under this distribution, the Monte Carlo sampling method is used to study the statistical characteristics. The sample {λ} obtained by the above steps S1 to S3 is the matrix AA T Therefore, the sum of the smallest mr eigenvalues ​​in the sample is taken as the compression error An unbiased estimate of .

[0126] This step provides an implementation path for the compression error estimation based on random matrix theory. Specifically, under the premise of knowing the compression rank r and matrix dimension (m, n), the compression error ∈ = ‖AA is estimated by sampling. r ‖ F , the process is formally expressed as:

[0127] ∈=g(r;m,n) Formula (15)

[0128] Among them, g() means that under the premise of known matrix dimension m×n and compression rank r, based on – The expected value or unbiased approximation of the compression error ∈ obtained by Monte Carlo sampling estimation of the Pastur distribution provides an error estimation path in the absence of the true matrix. ∈=‖AA r ‖ F is a strict mathematical definition; ∈ = g(r; m, n) is an approximate estimation formula based on sampling and random matrix theory; both point to the same error amount (but the former is a theoretical definition and the latter is a numerical estimate.

[0129] Step S5: Based on the mathematical relationship between the compression error and the compression rank, a mathematical relationship between the standard deviation and the compression rank is established.

[0130] Specifically, in step S5, in order to facilitate the dynamic adjustment of the compression rank, a fixed compression error constraint is imposed during the LLM training process, denoted as ∈ ini This error value is set when the gradient compression mechanism is first enabled.

[0131] Assume the random gradient matrix The standard deviation of is σ0, and the compression error ∈=‖AA in the whole training process r ‖ F The fixed compression error constraint is always satisfied. When the standard deviation of the gradient matrix changes from σ0 to σ1 during training, its compression rank is adjusted from the original r0 to:

[0132]

[0133] Proof process: For a random gradient matrix A with a standard deviation of σ0, if its compression rank is set to r0, then its compression error is ∈. In the subsequent training process, the gradient matrix evolves to A′, and its standard deviation decreases to σ1 (i.e., σ0<σ1). At this time, the ratio of the matrix Frobenius norm satisfies:

[0134]

[0135] When the compression rank is fixed to r, the compression error and standard deviation satisfy the following relationship:

[0136] ∈0·σ1=∈1·σ0 Formula (18)

[0137] That is, the compression error is proportional to the standard deviation, so:

[0138] ∈∝g(r;m,n)·σ0 Equation (19)

[0139] Among them, ∝ indicates that there is a proportional relationship between two quantities, which is the proportional symbol

[0140] Therefore, under the condition that the absolute value of the compression error remains unchanged (i.e., ‖AA r ‖ F=||A′-A′ r || F ), the compression rank and standard deviation required for matrix decomposition satisfy:

[0141] g(r0)σ0=g(r1)σ1 Equation (20)

[0142] So the update formula of compression rank is:

[0143]

[0144] This relationship shows that, under the constraint of a fixed compression error, if the standard deviation of the gradient distribution decreases, the required compression rank also decreases, thereby improving communication efficiency. If the standard deviation increases, the compression rank should be increased to maintain the upper bound on the error. This functional relationship provides a theoretical basis and computational foundation for the dynamic regulation of compression strategies during training.

[0145] Step S6: constructing a mathematical relationship between gradient entropy and standard deviation.

[0146] In this embodiment, gradient entropy, entropy value, and entropy all refer to the entropy obtained by gradient calculation, which refers to the specific calculation results. Information entropy is a theoretical definition.

[0147] Specifically, in step S6, for a normal distribution with a mean of μ and a standard deviation of σ, the information entropy is expressed as:

[0148]

[0149] Assume that the random variable X obeys a standard normal distribution with a mean of 0 and a variance of 1. According to the definition of information entropy, substitute the probability density function f(x) into Equation 22 to obtain:

[0150]

[0151] Here, e represents a natural constant, namely the Euler number. The first integral term corresponds to the probability density function of the normal distribution, and its integral result is 1. For the second integral term, let the variable substitution x = yσ + μ, then dx = σdy, and Equation 23 becomes:

[0152]

[0153] The gradient distribution gradually converges during training and is approximately considered to be normally distributed. Under this assumption, step S6 shows that there is a linear logarithmic relationship between entropy and standard deviation, which means that a decrease in gradient variance will lead to a decrease in entropy, providing a theoretical basis for tracking training dynamics based on entropy.

[0154] Step S7: constructing a mapping expression between the gradient entropy and the compression rank based on the mathematical relationship between the standard deviation and the compression rank and the mathematical relationship between the gradient entropy and the standard deviation.

[0155] Specifically, in step S7, during the training process, if the entropy value of the gradient matrix A changes from H0 to H1, under the premise of a fixed compression error, its compression rank is adjusted from r0 to r1, satisfying the following relationship:

[0156]

[0157] Derivation process: According to step S6, the entropy of a random variable that obeys a normal distribution is a logarithmic function of its standard deviation σ. Therefore, when the standard deviation of the gradient matrix changes from σ0 to σ1, the change in its entropy is:

[0158]

[0159] Transforming formula 26 yields:

[0160]

[0161] Since the compression rank depends on the standard deviation, according to step S5, we get:

[0162]

[0163] Through the seven steps described above, the present invention constructs a compression rank modeling framework that uses information entropy as a driving signal and combines compression error and standard deviation estimation. This method offers advantages such as computational closed-form expression, strong adaptability, and a solid theoretical foundation. It is suitable for communication compression scheduling in distributed training of various deep neural networks, significantly reducing communication overhead while ensuring model accuracy.

[0164] Example 2:

[0165] The present invention also provides a gradient low-rank compression modeling system based on information entropy. The gradient low-rank compression modeling system based on information entropy can be implemented by executing the process steps of the gradient low-rank compression modeling method based on information entropy, that is, those skilled in the art can understand the gradient low-rank compression modeling method based on information entropy as a preferred implementation of the gradient low-rank compression modeling system based on information entropy.

[0166] Specifically, the information entropy-based gradient low-rank compression modeling system includes:

[0167] Module M1, constructs the distribution model of gradient eigenvalues;

[0168] Module M2, constructs the functional relationship between compression error and compression rank;

[0169] Module M3, constructs the cumulative distribution function that conforms to the Marcenko-Pastur distribution;

[0170] Module M4, establishing a mathematical relationship between the compression error and the compression rank based on the cumulative distribution function, the functional relationship between the compression error and the compression rank, and the distribution model;

[0171] Module M5, establishing a mathematical relationship between the standard deviation and the compression rank based on the mathematical relationship between the compression error and the compression rank;

[0172] Module M6, constructs the mathematical relationship between gradient entropy and standard deviation;

[0173] Module M7 constructs a mapping expression between gradient entropy and compression rank based on the mathematical relationship between standard deviation and compression rank and the mathematical relationship between gradient entropy and standard deviation.

[0174] Specifically, in module M4, let the random matrix The elements are independent and identically distributed, with a mean of 0 and a variance of 1, denoted by A r is the optimal low-rank approximation of A with a compression rank of r, then its squared compression error is Estimated by the following modules:

[0175] Module M4.1, determines the upper and lower bounds a and b of the eigenvalue according to module M1, and uniformly samples the eigenvalue candidate set {λ0} within this range;

[0176] Module M4.2, based on the cumulative distribution function F(λ0) of the matrix eigenvalues ​​in module M3, calculates the cumulative probability {p0} corresponding to each λ0 and constructs a mapping relationship, i.e., the {(λ0,p0)} pair;

[0177] Module M4.3 uniformly samples m random probability values ​​{p} from the interval (0, 1), and uses the mapping relationship obtained in module M4.2 to obtain the corresponding eigenvalue samples {λ} through interpolation;

[0178] Module M4.4, sums the smallest mr values ​​in the eigenvalue sample {λ} as the compression error estimated value.

[0179] Those skilled in the art will appreciate that, in addition to implementing the system and its various devices, modules, and units provided by the present invention in purely computer-readable program code, it is entirely possible to implement the same functions of the system and its various devices, modules, and units provided by the present invention in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; the devices, modules, and units for implementing various functions can also be considered as both software modules implementing the method and structures within the hardware component.

[0180] The above describes specific embodiments of the present invention. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art may make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. The embodiments of this application and the features in the embodiments may be combined with each other in any manner unless there is a conflict.

Claims

1. A gradient low-rank compression modeling method based on information entropy, characterized in that: The following sub-steps are included: Step S1, constructing a distribution model of gradient eigenvalues; Step S2, constructing a functional relationship between compression error and compression rank; Step S3, constructing a cumulative distribution function that conforms to the Marcenko-Pastur distribution; Step S4, establishing a mathematical relationship between the compression error and the compression rank based on the cumulative distribution function, the functional relationship between the compression error and the compression rank, and the distribution model; Step S5, establishing a mathematical relationship between the standard deviation and the compression rank based on the mathematical relationship between the compression error and the compression rank; Step S6, constructing a mathematical relationship between gradient entropy and standard deviation; Step S7: constructing a mapping expression between the gradient entropy and the compression rank based on the mathematical relationship between the standard deviation and the compression rank and the mathematical relationship between the gradient entropy and the standard deviation.

2. The gradient low-rank compression modeling method based on information entropy according to claim 1 is characterized in that: In step S1, let G be The probability distribution on G has a mean of 0 and a variance of 1. Represents a set of real numbers, let Y be a p×n dimensional random matrix, the elements of Y are independent and identically distributed in G, and define the matrix S=n -1 YY T , T is the matrix transpose symbol, when p,n→∞ and satisfy p / n→γ∈(0,1), the empirical spectrum distribution F of the matrix S S It approaches the Marcenko-Pastur distribution in a weakly convergent manner, and the probability density function is: Where: γ represents p / n, x represents the eigenvalue variable of the spectral distribution, and y represents the subscript of the function symbol, not a variable.

3. The gradient low-rank compression modeling method based on information entropy according to claim 2 is characterized in that: In the said step S2, let the matrix have a singular value decomposition, that is, A = U∑V T , let r = rank(A) be the compression rank of A, take k < r, and define the truncated matrix in: m represents the number of rows of matrix A, n represents the number of columns of matrix A, U represents the left singular vector matrix, V represents the right singular vector matrix, r represents the rank of matrix A, the number of non-zero singular values, rank() is synonymous with r, representing the matrix rank, and k represents the rank retained by truncation, which controls the degree of compression; σ i is a singular value of A, u i and v i are the corresponding left and right singular vectors respectively; For all compressed matrices B of rank k, the minimum error under the 2-norm is given by A k gives: minutes rank(B)=k ‖AB‖2=‖AA k ‖2=σ k+1 expression(2) Similarly, for the Frobenius norm, the following conclusion holds: Among them, ‖‖2 represents the spectral norm of the matrix, which is equal to its maximum singular value and is used to measure the maximum error of the matrix in the sense of 2-norm; ‖‖ F represents the Frobenius norm, which is the square root of the sum of the squares of all matrix elements and is used to measure the overall approximation error.

4. The gradient low-rank compression modeling method based on information entropy according to claim 3 is characterized in that: In the step S3, is a random matrix, and the standard deviation of each element in A is 1, then AA T The cumulative distribution function corresponding to the eigenvalue λ is expressed as:

5. The gradient low-rank compression modeling method based on information entropy according to claim 4 is characterized in that: In step S4, the random matrix The elements are independent and identically distributed, with a mean of 0 and a variance of 1, denoted by A r Estimate the compression error by finding the optimal low-rank approximation of A with a compression rank of r 6. The gradient low-rank compression modeling method based on information entropy according to claim 5 is characterized in that: estimating the compression error The following sub-steps are included: Step S4.1, determining the eigenvalue range [a, b] according to step S1, and uniformly sampling the eigenvalue candidate set {λ0} within the eigenvalue range; Step S4.2, based on the cumulative distribution function F(λ0) of the matrix eigenvalues ​​in step S3, calculate the cumulative probability {p0} corresponding to each λ0, and construct a mapping relationship, i.e., a pair {(λ0, p0)}; Step S4.3: uniformly sample m random probability values ​​{p} from the interval (0, 1), and use the mapping relationship obtained in step S4.2 to obtain the corresponding eigenvalue samples {λ} by interpolation; Step S4.4, sum the smallest mr values ​​in the eigenvalue sample {λ} as the compression error Estimated value of: where λ i AA T The i-th eigenvalue of , in descending order, Under the premise of knowing the compression rank r and matrix dimension (m,n), the compression error ∈=‖AA is estimated by sampling r ‖ F , the process is formally expressed as: ∈=g(r;m,n) Formula (15) Among them, g() means that under the premise of known matrix dimension m×n and compression rank r, based on – The expected value or unbiased approximation of the compression error ∈ obtained by Monte Carlo sampling estimation of the Pastur distribution provides an error estimation path in the absence of the true matrix.

7. The gradient low-rank compression modeling method based on information entropy according to claim 6 is characterized in that: In step S5, a fixed compression error constraint is imposed during the LLM training process, denoted as ∈ ini , and is set when the gradient compression mechanism is first enabled, Stochastic gradient matrix The standard deviation of is σ0, and the compression error ∈=‖AA in the whole training process r ‖ F The fixed compression error constraint is always satisfied. When the standard deviation of the gradient matrix changes from σ0 to σ1 during training, the compression rank is adjusted from the original r0 to r1:

8. The gradient low-rank compression modeling method based on information entropy according to claim 7 is characterized in that: In step S6, for a normal distribution with a mean of μ and a standard deviation of σ, the information entropy is expressed as: Assume that the random variable X obeys a standard normal distribution with a mean of 0 and a variance of 1. According to the definition of information entropy, substitute the probability density function f(x) into Equation 22 to obtain: Here, e represents a natural constant, namely the Euler number. The first integral term corresponds to the probability density function of the normal distribution, and its integral result is 1. For the second integral term, let the variable replacement x = yσ + μ, then dx = σdy, and Equation 23 becomes:

9. The gradient low-rank compression modeling method based on information entropy according to claim 8 is characterized in that: In step S7, during the training process, if the entropy value of the gradient matrix A changes from H0 to H1, under the premise of a fixed compression error, its compression rank is adjusted from r0 to r1, satisfying the following relationship:

10. A gradient low-rank compression modeling system based on information entropy, characterized in that: include: Module M1, constructs the distribution model of gradient eigenvalues; Module M2, constructs the functional relationship between compression error and compression rank; Module M3, constructs the cumulative distribution function that conforms to the Marcenko-Pastur distribution; Module M4, establishing a mathematical relationship between the compression error and the compression rank based on the cumulative distribution function, the functional relationship between the compression error and the compression rank, and the distribution model; Module M5, establishing a mathematical relationship between the standard deviation and the compression rank based on the mathematical relationship between the compression error and the compression rank; Module M6, constructs the mathematical relationship between gradient entropy and standard deviation; Module M7 constructs a mapping expression between gradient entropy and compression rank based on the mathematical relationship between the standard deviation and the compression rank and the mathematical relationship between the gradient entropy and the standard deviation.

Citation Information

Patent Citations

  • Gradient compression method, device and equipment, distributed cluster and storage medium

    CN117910521A

Cited By

  • Gradient compression method and gradient compressor based on adaptive neural coding

    CN121303211A

  • A gradient compression method and gradient compressor based on adaptive neural coding

    CN121303211B