Large-scale language model compression method for fine-grained size control in variable memory environment
By calculating the activation value-aware weight matrix in a large language model and decomposing, dynamically loading or unloading residual data blocks, the memory and performance problems during deployment of large language models in variable memory environments are solved, and efficient model deployment and excellent performance are achieved.
Patent Information
- Application Number
- CN202510043329.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-16
AI Technical Summary
Existing large language model compression methods are difficult to achieve dynamic optimization in variable memory environments, resulting in problems such as memory footprint and performance damage when deployed in local devices.
By calculating the activation value-aware weight matrix and decomposing it into a symbol matrix and an absolute value matrix, iteratively decomposes to obtain multiple residual data blocks. Depending on the current available memory capacity and the importance of residual data blocks, residual data blocks are loaded or unloaded dynamically to build highly adaptable compression models.
Dynamic optimization of large language models in variable memory environments is achieved, ensuring that the model can run efficiently when deployed in local devices and maintains performance close to the original model under extreme compression ratios without additional training.
Smart Images

Figure CN120012842A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of model compression, and in particular relates to a large language model compression method with fine-grained size control in a variable memory environment. Background Art
[0002] Large language models (LLMs) have shown excellent performance in various benchmarks and are gradually being used in daily life, such as general language assistants, search engines, and code assistants. As the scale increases, the deployment of large language models has gradually shifted from performance improvement to usability improvement, such as memory usage becoming a bottleneck. Loading large models requires significant memory resources, which poses a challenge to deployment on low-resource devices.
[0003] The compression methods for large language models mainly include quantization, pruning, knowledge distillation, and weight decomposition. Among them, traditional compression methods such as quantization, pruning, and knowledge distillation require a preset compression ratio, have limited adaptability, and have complex compression steps. Although the previous weight decomposition method can reduce storage space requirements, it will significantly damage performance at high compression ratios and is difficult to adapt to dynamic memory changes.
[0004] In order to make large language models benefit more people, deploying models on local devices has gradually become a strong demand. Deploying large language models on local devices requires taking into account both memory and performance, and existing methods are difficult to achieve dynamic optimization in a variable memory environment. Existing methods are difficult to adapt to the limited and rapidly changing available memory capacity in local devices due to the need to preset compression ratios, and switching compression models with different compression ratios will bring great memory loading delays and storage burdens. Therefore, it is currently necessary to propose a model compression method that can adapt to variable memory environments to make efficient local model deployment possible. Summary of the invention
[0005] The present invention is made to solve the above problems, and aims to provide a large language model compression method that can realize dynamic optimization in a variable memory environment. The present invention adopts the following technical solutions:
[0006] The present invention provides a large language model compression method with fineness size control in a variable memory environment. The method has such a technical feature that it comprises the following steps: step S1, calculating the activation value corresponding to each weight matrix of the large language model to be compressed through a predetermined calibration data set, and obtaining an activation value-aware weight matrix based on the activation value; step S2, decomposing each of the activation value-aware weight matrices into its symbol matrix and absolute value matrix, and iteratively decomposing the absolute value matrix, and obtaining a plurality of residual data blocks based on the symbol matrix and the iterative decomposition result; step S3, evaluating and sorting the plurality of residual data blocks in terms of importance; step S4, dynamically loading or unloading the residual data block according to the change of the current available memory capacity and the importance of the residual data block, and obtaining a compression model adapted to the current available memory capacity.
[0007] The large-scale language model compression method with fine-grained size control in a variable memory environment provided by the present invention may also have such a technical feature, wherein, in step S1, a scaling factor is calculated based on the activation value, and the corresponding rows in each weight matrix are scaled using the scaling factor to obtain the activation value-aware weight matrix.
[0008] The large language model compression method with fine-grained size control in a variable memory environment provided by the present invention may also have such a technical feature, wherein, in step S1, an activation value matrix is obtained based on the activation value, and the L2 norm of the activation value matrix is counted as the scaling factor:
[0009] s=[||x1||2,||x2||2,…,|x n ||2]
[0010] The corresponding rows in each weight matrix are scaled by the scaling factor to obtain the activation value-aware weight matrix:
[0011] XW=Xdiag(1 / s)diag(s)W=Xdiag(1 / s)W scaled
[0012] In the formula, x n is the value of each column of the activation value matrix, X is the activation value matrix, W is the weight matrix, W scaled is the activation value-aware weight matrix.
[0013] The large-scale language model compression method with fineness size control under a variable memory environment provided by the present invention may also have such a technical feature, wherein step S2 includes the following sub-steps: step S2-1, decomposing the activation value-aware weight matrix into the product of the symbol matrix and the absolute value matrix; step S2-2, performing singular value decomposition on the absolute value matrix to be decomposed, and retaining the symbol matrix and the first k singular vectors and singular values obtained by singular value decomposition, to obtain the residual data block composed of the symbol matrix and the first k singular vectors; step S2-3, judging whether a predetermined number of iterations has been reached, and obtaining a plurality of decomposed residual data blocks when the judgment is yes; step S2-4, when the judgment is no in step S2-3, calculating the approximate residual based on the absolute value matrix to be decomposed and the residual data block, to obtain a residual absolute value matrix; step S2-5, using the residual absolute value matrix as a new absolute value matrix to be decomposed, and returning to step S2-2 to continue decomposing the residual absolute value matrix.
[0014] The large language model compression method with fine-grained size control in a variable memory environment provided by the present invention may also have such a technical feature, wherein in step S2-1, the decomposition of the activation value-aware weight matrix is expressed as:
[0015] W scaled =W sign ☉|W|
[0016] In step S2-2, the singular value decomposition is expressed as:
[0017]
[0018] The residual data block obtained after n-times iterative decomposition is expressed as:
[0019]
[0020] Where W sign is a symbol matrix, |W| is an absolute value matrix, A' is a left singular matrix, B' is a right singular matrix, is the residual data block, is the symbol matrix of the residual data block.
[0021] The large-scale language model compression method with fine-grained size control in a variable memory environment provided by the present invention may also have such a technical feature, wherein the symbol matrix only contains 1 and -1, which are packaged as a data type supported by the CPU for storage, and each parameter occupies 1 bit of memory capacity.
[0022] The large-scale language model compression method with fine grain size control in a variable memory environment provided by the present invention may also have such a technical feature, wherein, in step S3, a predetermined calibration data set is used to calculate the perplexity of different compression models constructed when loading different residual data blocks, so as to evaluate the importance of different residual data blocks, and globally sort the multiple residual data blocks according to the importance.
[0023] The large-scale language model compression method with fine grain size control in a variable memory environment provided by the present invention may also have such a technical feature, wherein, in step S4, based on the current available memory capacity and the importance of the residual data blocks, the top m residual data blocks in order of importance are selected to form the compression model.
[0024] The large-scale language model compression method with fine grain size control in a variable memory environment provided by the present invention may also have such a technical feature, wherein, in step S4, when the current available memory capacity increases, a number of residual data blocks starting from the m+1th in importance ranking are loaded into the compression model accordingly according to the increase amount to form a new compression model; when the current available memory capacity decreases, a number of residual data blocks with the lowest importance ranking are unloaded from the compression model accordingly according to the decrease amount to form a new compression model.
[0025] Functions and Effects of the Invention
[0026] According to the large language model compression method with fineness size control in a variable memory environment provided by the present invention, the method includes the steps of calculating an activation value-aware weight matrix, iteratively decomposing the absolute value matrix of the activation value-aware weight matrix, evaluating and sorting the importance of multiple residual data blocks decomposed iteratively, and dynamically loading residual data blocks according to the change of available content capacity and importance sorting to form different compression models. Through such a method, the compression model can be well adapted to the variable memory environment, so that the model can be deployed in the local device, and because the residual data blocks with higher importance are loaded preferentially, the compression model still has excellent performance close to the original large language model, even in the case of extreme compression ratios, so the compression model does not need to be trained again, making the deployment of the model in the local device more convenient and efficient. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 It is a flow chart of a large language model compression method with fine-grained size control in a variable memory environment according to an embodiment of the present invention;
[0028] Figure 2 is a flow chart of step S2 in an embodiment of the present invention;
[0029] Figure 3 is a schematic diagram of a residual data block in an embodiment of the present invention;
[0030] Figure 4 is a schematic diagram of dynamically loading residual data blocks in an embodiment of the present invention;
[0031] Figure 5 It is a schematic diagram of dynamically unloading residual data blocks in an embodiment of the present invention. DETAILED DESCRIPTION
[0032] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the large-scale language model compression method with fine grain size control under a variable memory environment of the present invention is specifically described below in combination with embodiments and drawings.
[0033] Figure 1 It is a flow chart of a large language model compression method with fine-grained size control in a variable memory environment in this embodiment.
[0034] like Figure 1 As shown, the method comprises the following steps:
[0035] Step S1, calculating activation values corresponding to each weight matrix of the large language model to be compressed through a predetermined calibration data set, and obtaining an activation value-aware weight matrix based on the activation values.
[0036] Step S2, decomposing the weight matrix of each activation value perception into its sign matrix and absolute value matrix, and iteratively decomposing the absolute value matrix, and obtaining multiple residual data blocks based on the sign matrix and the iterative decomposition result.
[0037] Step S3, evaluating and sorting the importance of multiple residual data blocks.
[0038] Step S4, dynamically loading or unloading the residual data block according to the current available memory capacity of the local device and the importance of the residual data block, to obtain a compression model that is compatible with the current available memory capacity.
[0039] The above steps will be described in detail below.
[0040] Step S1, calculating activation values corresponding to each weight matrix of the large language model to be compressed through a predetermined calibration data set, and obtaining an activation value-aware weight matrix based on the activation values.
[0041] Among them, the L2 norm of each column of the statistical activation value matrix is used as the scaling factor:
[0042] s=[||x1||2,||x2||2,…,||x n ||2]
[0043] In the formula, x n are the values of each column of the activation value matrix.
[0044] Then, the corresponding rows in each weight matrix are scaled by the scaling factor s to obtain the activation value-aware weight matrix:
[0045] XW=Xdiag(1 / s)diag(s)W=Xdiag(1 / s)W scaled
[0046] In the formula, X is the activation value matrix, W is the weight matrix, and W scaled is the activation-value-aware weight matrix. Using the activation-value-aware weight matrix, the rows in the weight matrix W that have a greater impact on the model output are restored as much as possible in the subsequent decomposition.
[0047] Step S2, decomposing the weight matrix of each activation value perception into its sign matrix and absolute value matrix, and iteratively decomposing the absolute value matrix, and obtaining multiple residual data blocks based on the sign matrix and the iterative decomposition result.
[0048] Figure 2 is a flow chart of step S2 in this embodiment.
[0049] like Figure 2 As shown, step S2 specifically includes the following sub-steps:
[0050] Step S2-1, decomposing the activation value-aware weight matrix into the product of its sign matrix and absolute value matrix.
[0051] W scaled =W sign ⊙|W|
[0052] Where W sign is a symbol matrix, and |W| is an absolute value matrix.
[0053] Step S2-2, performing singular value decomposition on the absolute value matrix to be decomposed, and retaining the symbol matrix and the first k singular vectors and singular values obtained by the singular value decomposition, to obtain a residual data block consisting of the symbol matrix and the first k singular vectors.
[0054]
[0055] Where A' is the left singular matrix and B' is the right singular matrix.
[0056] Among them, since the symbol matrix W sign It only contains 1 and -1, so it can be packaged as a data type supported by the CPU for storage in the local device, and each parameter occupies 1 bit of memory.
[0057] Step S2-3, judging whether the predetermined number of iterations n has been reached, and entering the end state when the judgment is yes, to obtain a plurality of decomposed residual data blocks.
[0058] Step S2-4, when the judgment in step S2-3 is no, the approximate residual is calculated based on the absolute value matrix to be decomposed (before this round of iterative decomposition) and the residual data block to obtain a residual absolute value matrix.
[0059] Step S2-5, taking the obtained residual absolute value matrix as a new absolute value matrix to be decomposed, and returning to step S2-2, that is, repeating the above steps to continue to decompose the residual absolute value matrix.
[0060] After n iterations of decomposition, the reconstructed matrix is:
[0061]
[0062] In the formula, is the residual data block, is the symbol matrix of the residual data block.
[0063] Figure 3 is a schematic diagram of a residual data block in this embodiment.
[0064] Through step S2, the basic unit decomposed is a "residual block", that is, Its composition is as follows Figure 3 shown.
[0065] Figure 4 is a schematic diagram of dynamically loading residual data blocks in this embodiment, Figure 5 It is a schematic diagram of dynamically unloading residual data blocks in this embodiment, which shows the situation of loading / unloading residual data blocks for the i-th layer in the model. Exemplarily, multiple residual data blocks are stored in a hard disk (such as an SSD or HDD), the model is loaded into the memory, and the residual data blocks are dynamically loaded / unloaded into the memory.
[0066] like Figure 4 and Figure 5 As shown, by loading multiple residual data blocks, a compression model of a large language model can be obtained, and by loading different residual data blocks, compression models of different sizes can be constructed, thereby achieving the adaptability of the compression model in the change of available content. Specifically, for each weight matrix of a large language model, one or more residual data blocks can be loaded. In this embodiment, in order to reduce the space required for searching residual data blocks, the difference in the number of residual data blocks loaded in different weight matrices is limited to no more than 1.
[0067] Since a residual data block contains only one symbol matrix and k singular vectors, and the number of singular vectors is much smaller than the dimension of the original weight matrix, the memory usage of each residual data block is very small, making the memory usage of the compression model only at the MB level.
[0068] Step S3, evaluating and sorting the importance of multiple residual data blocks.
[0069] A calibration dataset containing 32 samples is used to calculate the perplexity of different compression models constructed by loading different residual data blocks, so as to evaluate the importance of different residual data blocks. Then, multiple residual data blocks are globally sorted according to their importance.
[0070] Step S4, dynamically loading or unloading the residual data block according to the current available memory capacity of the local device and the importance of the residual data block, to obtain a compression model that is compatible with the current available memory capacity.
[0071] For example, the current compression model is first formed by the first m residual data blocks in terms of importance. Then, when the current available memory capacity increases, several residual data blocks starting from the m+1th in importance are loaded into the current compression model accordingly according to the increase, forming a new compression model to make full use of the available memory and improve the model accuracy. When the current available memory capacity decreases, several residual data blocks ranked last in importance are unloaded from the current compression model according to the decrease, forming a new compression model to adapt to the change in memory capacity and maintain the model accuracy as much as possible.
[0072] In the above process, since the residual data blocks with higher importance are loaded first, the compressed model still has excellent performance close to the original large language model, even in the case of extreme compression ratio. And since the performance is very close to the original large language model in this way, the compressed model does not need to be trained again.
[0073] Functions and Effects of the Embodiments
[0074] According to the large language model compression method with fine-grained size control in a variable memory environment provided by this embodiment, the method includes the steps of calculating an activation value-aware weight matrix, iteratively decomposing the absolute value matrix of the activation value-aware weight matrix, evaluating and sorting the importance of multiple residual data blocks decomposed iteratively, and dynamically loading residual data blocks to form different compression models according to the change of available content capacity and importance sorting. Through such a method, the compression model can be well adapted to the variable memory environment, so that the model can be deployed in the local device, and because the residual data blocks with higher importance are loaded preferentially, the compression model still has excellent performance close to the original large language model, even in the case of extreme compression ratios, so the compression model does not need to be trained again, making the deployment of the model in the local device more convenient and efficient.
[0075] In the embodiment, the symbol matrix and the first k singular vectors obtained by singular value decomposition are used as residual data blocks. Since the k value is much smaller than the dimension of the original weight matrix, the memory occupancy of the residual data block is very small, so that the memory occupancy of the compression model is only at the MB level, and a high compression ratio can be achieved.
[0076] In the embodiment, a heuristic residual data block importance ranking method is adopted to evaluate the importance of different residual data blocks by calculating the perplexity of the compression model composed of loading different residual data blocks. This can not only effectively evaluate the importance, so that the model can preferentially load the residual data blocks that have a greater impact on performance, but also has good interpretability.
[0077] The above embodiments are only used to illustrate the specific implementation of the present invention, and the present invention is not limited to the description scope of the above embodiments. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and descriptions are only to illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, and these changes and improvements fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. A large language model compression method with fine-grained size control in a variable memory environment, characterized in that: The following steps are involved: Step S1, calculating activation values corresponding to each weight matrix of the large language model to be compressed through a predetermined calibration data set, and obtaining an activation value-aware weight matrix based on the activation values; Step S2, decomposing each activation value-aware weight matrix into its sign matrix and absolute value matrix, and iteratively decomposing the absolute value matrix, and obtaining a plurality of residual data blocks based on the sign matrix and the iterative decomposition result; Step S3, evaluating and sorting the importance of the plurality of residual data blocks; Step S4, dynamically loading or unloading the residual data blocks according to the change of the available memory capacity and the importance ranking of the residual data blocks, to obtain a compression model adapted to the available memory capacity.
2. The large language model compression method with fine-grained size control in a variable memory environment according to claim 1, characterized in that: in, In step S1, a scaling factor is calculated based on the activation value, and the corresponding row in each weight matrix is scaled by the scaling factor to obtain the activation value-aware weight matrix.
3. The large language model compression method with fineness control in a variable memory environment according to claim 2, characterized in that: in, In step S1, an activation value matrix is obtained based on the activation value, and the L2 norm of the activation value matrix is counted as the scaling factor: s=[||x1||2,||x2||2,…,||x n ||2] The corresponding rows in each weight matrix are scaled by the scaling factor to obtain the activation value-aware weight matrix: XW=Xdiag(1 / s)diag(s)W=Xdiag(1 / s)W scaled In the formula, x n is the value of each column of the activation value matrix, X is the activation value matrix, W is the weight matrix, W scaled is the activation value-aware weight matrix.
4. The large language model compression method with fine-grained size control in a variable memory environment according to claim 2, Features: Wherein, step S2 includes the following sub-steps: Step S2-1, decomposing the activation value-aware weight matrix into the product of the sign matrix and the absolute value matrix; Step S2-2, performing singular value decomposition on the absolute value matrix to be decomposed, and retaining the symbol matrix and the first k singular vectors and singular values obtained by the singular value decomposition, to obtain the residual data block consisting of the symbol matrix and the first k singular vectors; Step S2-3, determining whether a predetermined number of iterations has been reached, and obtaining the decomposed plurality of residual data blocks when the number of iterations has been reached. Step S2-4, when the judgment in step S2-3 is no, an approximate residual is calculated based on the absolute value matrix to be decomposed and the residual data block to obtain a residual absolute value matrix; Step S2-5: taking the residual absolute value matrix as a new absolute value matrix to be decomposed, and returning to step S2-2 to continue decomposing the residual absolute value matrix.
5. The large language model compression method with fineness control in a variable memory environment according to claim 4, characterized in that: in, In step S2-1, the decomposition of the activation value-aware weight matrix is expressed as: IN scaled =In sign ⊙|W| In step S2-2, the singular value decomposition is expressed as: The residual data block obtained after n-times iterative decomposition is expressed as: Where W sign is a symbol matrix, |W| is an absolute value matrix, A' is a left singular matrix, B' is a right singular matrix, is the residual data block, is the symbol matrix of the residual data block.
6. The large language model compression method with fineness size control in a variable memory environment according to claim 5, characterized in that: in, The symbol matrix only contains 1 and -1, which are packed into a data type supported by the CPU for storage, and each parameter occupies 1 bit of memory capacity.
7. The large language model compression method with fineness control in a variable memory environment according to claim 1, characterized in that: in, In step S3, the perplexity of different compression models constructed by loading different residual data blocks is calculated using a predetermined calibration data set, so as to evaluate the importance of different residual data blocks. The plurality of residual data blocks are globally sorted according to the importance.
8. The large language model compression method with fineness control in a variable memory environment according to claim 7, characterized in that: in, In step S4, according to the available memory capacity and the importance of the residual data blocks, the first m residual data blocks in order of importance are selected to form the compression model.
9. The large language model compression method with fineness size control in a variable memory environment according to claim 8, characterized in that: in, In step S4, when the available memory capacity increases, a plurality of residual data blocks starting from the m+1th in importance order are loaded into the compression model accordingly according to the increase amount to form a new compression model. When the available memory capacity decreases, a number of residual data blocks ranked last in importance are unloaded from the compression model according to the amount of decrease to form a new compression model.