A large model quantization and compression method and system based on the polar coordinate system
By using an extreme coordinate system to separate and optimize vector direction and magnitude, the method addresses the inefficiencies of traditional LLM quantization, achieving improved compression and deployment efficiency.
Patent Information
- Application Number
- CN202510510093.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-23
AI Technical Summary
The existing large-language model quantization methods cannot guarantee the low-cost hardware deployment and inference efficiency of large models at the same time. When traditional methods quantize under the Cartesian coordinate system, the coupling of the direction of the vector and the modular length leads to unstable quantization sensitivity, affecting the compression effect of the model.
The quantitative compression method based on the polar coordinate system is adopted, and the original weight parameter distribution of the large language model is processed into a standard Gaussian distribution, and the direction codebook and the modular length codebook are respectively constructed, and the direction and modular length are independently optimized. The enhanced Hadamar rotation and Llyod-MAX algorithm are used for quantization to form a vector codebook that conforms to the distribution characteristics and quantization sensitivity characteristics.
It improves model compression performance, reduces memory usage, reduces hardware deployment costs, and improves inference efficiency, breaking through the problem of mismatch between quantization sensitivity and quantization methods caused by the coupling of direction and modular length in traditional methods.
Smart Images

Figure CN120046664B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a large model quantization and compression method and system based on a polar coordinate system. Background Technique
[0002] The statements in this part only provide background technical information related to the present invention, and do not necessarily constitute prior art.
[0003] Large Language Model (LLM) is a natural language processing model based on deep learning. Its huge number of parameters brings great challenges to efficient inference and practical deployment. Post Training Quantization (PTQ) methods (for example, using Vector Quantization (VQ)) can transform the parameters of the large language model from high-precision floating-point values to low-precision integer values to effectively compress the model size. Traditional VQ methods directly cluster and optimize the original data distribution through Mean Square Error (MSE) in the Cartesian coordinate system, and use the cluster centers as new vector representations. It does not analyze the characteristics of the original data distribution, and the effect of the quantization model largely depends on the iterative optimization effect of the clustering algorithm; moreover, in the weights of the large language model, the direction of the vector is more sensitive to the quantization process than the magnitude of the vector.
[0004] In summary, the existing large language model quantization methods cannot simultaneously ensure low-cost hardware deployment and inference efficiency of large models. Summary of the Invention
[0005] In order to solve the technical problems existing in the above background technique, the present invention provides a large model quantization and compression method and system based on a polar coordinate system, which can simultaneously ensure low-cost hardware deployment and inference efficiency of large models.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] The first aspect of the present invention provides a large model quantization and compression method based on a polar coordinate system.
[0008] A large model quantization and compression method based on a polar coordinate system, comprising:
[0009] Retrieve the original weight parameter distribution of the pre-stored large language model from the first storage unit, and process it into a standard Gaussian distribution to obtain the corresponding weight vectors that conform to the standard Gaussian distribution;
[0010] Convert each of the weight vectors that conform to the standard Gaussian distribution into corresponding polar coordinate representations;
[0011] Construct the direction codebook and magnitude codebook for each of the polar coordinate representations respectively, obtain a vector codebook that conforms to the distribution characteristics and quantization sensitivity characteristics, and transmit it to the second storage unit for storage;
[0012] Among them, the storage format of the vector codebook in the second storage unit is direction index, magnitude index, and element value; the direction index and magnitude index are concatenated to form a final index, and the final index is integrated in units of the original weight parameters of the large language model to form a quantized model weight matrix; the quantized model weight matrix is used to provide to the computing unit to run the large language model on the computing unit.
[0013] As an implementation, the process of processing the original weight parameter distribution of the large language model into a standard Gaussian distribution includes:
[0014] Calculate the covariance matrix of the original weight parameters of the large language model;
[0015] Perform a Karhunen-Loeve transform on the covariance matrix to obtain an enhanced matrix;
[0016] Multiply the enhanced matrix by the Hadamard matrix for fusion to obtain an enhanced rotation matrix;
[0017] Among them, the parameter distributions in the enhanced rotation matrix are the standard Gaussian distribution.
[0018] As an implementation, the weight vector that conforms to the standard Gaussian distribution contains two variables, namely: direction and magnitude.
[0019] As an implementation, the process of constructing the direction codebook for each of the polar coordinate representations includes:
[0020] Select the vector directions of the E8 lattice at a given magnitude as the base class;
[0021] Select a specified number of subsets from the base class using the greedy algorithm according to the quantization bits of the direction, that is, obtain the direction codebook.
[0022] As an implementation, the process of constructing the magnitude codebook for each of the polar coordinate representations includes:
[0023] According to the root distribution of the chi-square distribution that the weight vector conforming to the standard Gaussian distribution conforms to, obtain the distribution of the magnitude, and then calculate the probability density function of the magnitude;
[0024] According to the probability density function of the magnitude, use the Llyod-MAX algorithm to quantize the magnitude to form the magnitude codebook.
[0025] As an implementation manner, the process of quantifying the modulus length using the Lloyd-MAX algorithm is as follows:
[0026] Select the initial quantization lattice points and quantization boundaries;
[0027] According to the initial quantization lattice points and quantization boundaries, calculate the centroid and boundaries in sequence;
[0028] Judge whether the Lloyd-MAX algorithm converges. If the change in quantization distortion between two adjacent iterations is less than a preset threshold or the preset number of iterations is reached, stop the iteration; otherwise, continue the iterative calculation of calculating the centroid and boundaries in sequence until the Lloyd-MAX algorithm converges, and finally obtain the quantized codebook lattice points.
[0029] The second aspect of the present invention provides a large model quantization compression system based on the polar coordinate system.
[0030] A large model quantization compression system based on the polar coordinate system includes:
[0031] A weight parameter distribution standardization module, which is used to retrieve the original weight parameter distribution of the large language model pre-stored in the first storage unit and process it into a standard Gaussian distribution to obtain a corresponding weight vector conforming to the standard Gaussian distribution;
[0032] A weight parameter polar coordinate conversion module, which is used to convert each of the weight vectors conforming to the standard Gaussian distribution into a corresponding polar coordinate representation;
[0033] A vector codebook construction and storage module, which is used to construct a direction codebook and a modulus length codebook for each of the polar coordinate representations respectively, obtain a vector codebook conforming to the distribution characteristics and quantization sensitivity characteristics, and transmit it to the second storage unit for storage;
[0034] Among them, the storage format of the vector codebook in the second storage unit is direction index, modulus length index, and element value; the direction index and the modulus length index are spliced to form a final index, and the final index is integrated in units of the original weight parameters of the large language model to form a quantized model weight matrix; the quantized model weight matrix is used to provide to the computing unit to run the large language model on the computing unit.
[0035] The third aspect of the present invention provides a computer-readable storage medium.
[0036] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the above-mentioned large model quantization compression method based on the polar coordinate system.
[0037] The fourth aspect of the present invention provides a computer program product.
[0038] A computer program product includes a computer program / instruction, which, when executed by a processor, implements the steps in the large model quantization compression method based on the polar coordinate system as described above.
[0039] A fifth aspect of the present invention provides an electronic device.
[0040] An electronic device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the large model quantization compression method based on the polar coordinate system as described above are implemented.
[0041] Compared with the prior art, the present invention has the following beneficial effects:
[0042] (1) The present invention is applicable to low-bit quantization compression of a large language model. The distribution of the original weight parameters of the large language model is reshaped into a Gaussian distribution, and the vector direction and modulus are separated by a polar coordinate system. An independently optimized direction codebook and modulus codebook are constructed, thereby obtaining a vector codebook that conforms to the distribution characteristics and quantization sensitivity characteristics. The vector codebook is only stored in the format of direction index, modulus index and element value storage, thereby avoiding the quantization error instability problem caused by clustering effect differences in traditional methods. While improving the performance of the compression model, it also ensures low-cost hardware deployment of the large language model.
[0043] (2) The polar coordinate quantization framework proposed in the present invention supports independent bit allocation for direction and modulus length. It alleviates its high quantization sensitivity by increasing the proportion of directional bits and reduces redundant storage of modulus length, thus reducing the memory usage of the codebook exponentially. This breaks through the problem of mismatch between quantization sensitivity and quantization method caused by the coupling of direction and modulus length in traditional methods.
[0044] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0046] Figure 1 It is a comparison diagram of the quantization of the direction and modulus of the weight parameter vector in the LLaMA-2-70B model using the existing vector quantization method;
[0047] Figure 2 is a flow chart of a large model quantization compression method based on a polar coordinate system according to an embodiment of the present invention;
[0048] Figure 3It is a process diagram for processing the original weight parameter distribution of the large language model into a standard Gaussian distribution in an embodiment of the present invention;
[0049] Figure 4(a) is the effect diagram of traditional Hadamard rotation;
[0050] Figure 4(b) is the enhanced Hadamard rotation effect diagram proposed by the present invention;
[0051] Figure 5 It is a process diagram for constructing a direction codebook for polar coordinate representation in an embodiment of the present invention;
[0052] Figure 6 It is a process diagram for constructing a magnitude codebook for polar coordinate representation in an embodiment of the present invention;
[0053] Figure 7 It is a schematic diagram of vector quantization direction error and magnitude error in an embodiment of the present invention;
[0054] Figure 8 It is a schematic diagram of polar coordinate quantization lattice points in an embodiment of the present invention;
[0055] Figure 9 It is a schematic diagram of the large model quantization and compression system structure based on the polar coordinate system in an embodiment of the present invention. Detailed implementation manners
[0056] The present invention will be further described below in conjunction with the drawings and embodiments.
[0057] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further descriptions of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0058] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "include" and / or "comprise" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0059] Term explanation:
[0060] Large Language Model (LLM) is a natural language processing model based on deep learning. Through training with a vast amount of text data, it forms an ultra-large-scale neural network with billions to trillions of parameters, capable of understanding semantics, generating coherent text, and performing various language tasks. For example, storing the parameters of the LLaMA-2-70B model in FP16 format requires 140GB of memory, which exceeds the capacity of high-performance GPUs and makes multi-GPU deployment difficult. Such a large scale has a significant impact on memory capacity and hard disk storage, and requires a large amount of bandwidth during inference. Among them, the LLaMA-2-70B model is a well-known large language model with 70 billion parameters.
[0061] Figure 1 Compares the effects of quantifying the direction and magnitude of the weight parameter vectors in the LLaMA-2-70B model respectively using existing vector quantization methods. From Figure 1 it can be seen that under the experimental settings of various quantization bit numbers (number of cluster centers), the accuracy loss of direction quantization is greater than that of magnitude quantization. However, the existing measurement index MSE (Mean Squared Error) is more sensitive to changes in magnitude, which is contrary to the actual situation. And the codebooks of the existing methods couple the direction and magnitude of the vector, and cannot adjust the weights of the direction and magnitude in the VQ method, ultimately resulting in the inability to simultaneously ensure low-cost hardware deployment and inference efficiency of large models.
[0062] Figure 2 Gives the process of the large model quantization and compression method based on the polar coordinate system in the embodiments of the present invention. Combining Figure 2 , the large model quantization and compression method based on the polar coordinate system in the embodiments of the present invention specifically includes the following steps S201~S203.
[0063] S201: Retrieve the original weight parameter distribution of the pre-stored large language model from the first storage unit, and process it into a standard Gaussian distribution to obtain the corresponding weight vector that conforms to the standard Gaussian distribution.
[0064] First, deep learning frameworks default to using Gaussian distribution for network parameter initialization. Second, normalization layers in deep neural networks, such as Batch Norm and Layer Norm, further stabilize the mean and variance, making the weights and activation values close to Gaussian distribution. Finally, L2 regularization is a commonly used technique in the training of most large language models, and its prior distribution is essentially Gaussian distribution, which strengthens the effectiveness of modeling weight parameters with Gaussian distribution.
[0065] The standard Hadamard rotation cannot align the variances between channels, which also limits its ability to transform the original distribution into a Gaussian distribution. This embodiment proposes an enhanced Hadamard rotation method to process the original weight parameter distribution of the large language model into a standard Gaussian distribution, as Figure 3 shown. The process of processing the original weight parameter distribution of the large language model into a standard Gaussian distribution includes:
[0066] S2011: Calculate the covariance matrix of the original weight parameters of the large language model;
[0067] Given a model weight parameter matrix (satisfying the following assumption: the columns of the matrix are zero-mean) , with its dimension being (n, m), the Hadamard transform matrix H has dimension (m, m), and the covariance matrix of the transformation matrix can be expressed as formula (1):
[0068] (1);
[0069] where K represents the matrix of eigenvectors obtained from eigenvalue decomposition, represents the diagonal eigenvalue matrix, and the superscript T represents the transpose operation of the matrix. Then, the -th diagonal element can be expressed as formula (2):
[0070] (2);
[0071] where: is the -th diagonal value of . For a given value, formula (2) represents the variance of the
[0072] S2012: Perform the Karhunen-Loeve transform on the covariance matrix to obtain an enhanced matrix;
[0073] Since the vector varies with l, the channel variances cannot be numerically proven to be similar. In addition, considering that H is a fixed matrix while K and are input-dependent, the Hadamard transform cannot adjust for channel differences in all scenarios. Therefore, this property of the prior art inevitably leads to differences in variances between channels, resulting in suboptimal Gaussian distribution transformation ability.
[0074] The enhanced Hadamard rotation method based on the Karhunen-Loeve transform (KLT) is used to align the variances between the channels of the model weight parameters and enhance the Gaussianization effect of the rotation.
[0075] Specifically, perform eigenvalue decomposition on the covariance matrix to apply KLT, as shown in Equation (3):
[0076] (3);
[0077] S2013: Multiply the enhanced matrix by the Hadamard matrix for fusion to obtain an enhanced rotation matrix;
[0078] Among them, the distribution of each parameter in the enhanced rotation matrix is a standard Gaussian distribution.
[0079] Calculate the KLT enhanced rotation matrix HK, as shown in Equation (4):
[0080] (4);
[0081] Equation (1) is transformed into Equation (5):
[0082] (5);
[0083] Among them, represents the identity matrix.
[0084] Correspondingly, Equation (2) is transformed into Equation (6):
[0085] (6);
[0086] In this way, the variances of each channel become the same, which can make the distribution after the rotation transformation closer to the Gaussian distribution. Figure 4(a) is the effect diagram of the traditional Hadamard rotation, and Figure 4(b) is the effect diagram of the enhanced Hadamard rotation proposed by the present invention. By comparison, it can be found that the method proposed in the embodiment of the present invention can further reduce outliers and make the quantization errors of more data smaller, reducing the overall quantization error.
[0087] Rotation can suppress the outliers of the original weight parameters to a certain extent and increase the representation ability of the codebook. At the same time, it is easy to solve the vector direction and modulus length that conform to the Gaussian distribution, which is helpful for the construction of the direction and modulus length codebooks. In this embodiment, an enhanced matrix is introduced before the Hadamard rotation, which can align the variances between different channels of the original weight parameters, reduce outliers to a certain extent, and is more conducive to the Hadamard rotation to process the original weight parameters into a standard Gaussian distribution. When the distribution characteristics of the quantization variable are known, the quantization lattice points that conform to the original distribution can be theoretically deduced.
[0088] S202: Convert each of the weight vectors conforming to the standard Gaussian distribution into corresponding polar coordinate representations.
[0089] The direction of the weight vector of the large language model is more sensitive to quantization than the magnitude (because the direction has a higher dimension, higher spatial freedom, and higher compression loss in the case of low bits). As Figure 1 shown, the embodiments of the present invention designed a set of control verification experiments, clustering the direction and magnitude of the model weight parameter vectors at different bit widths respectively, and then evaluating the mean of the zero-shot accuracy of the quantization model on multiple data sets. The results show that under the same bit width setting, the accuracy of magnitude quantization is higher than that of direction quantization, and as the quantization bit decreases, the accuracy of direction quantization drops more significantly, while the accuracy of magnitude quantization still remains close to floating point.
[0090] However, the existing (mean square error)-based clustering method is more sensitive to magnitude changes. The first reason is that is essentially a measure of the relative distance between spaces. It only covers the consideration of direction because in VQ, it is artificially stipulated that the starting points of all vectors are at the origin. From the formula form, and the magnitude ( ) have a similar calculation form, as shown in formulas (7) and (8):
[0091] (7);
[0092] (8);
[0093] where and represent the original vector to be quantized and the clustering center respectively, represents the dimension of the vector. On the other hand, can also be expressed by formula (9):
[0094] (9);
[0095] where, and represent and 's magnitudes respectively, represents and 's included angle, represents the cosine function. If is expressed in the form of ( represents the change amount of the magnitude), then formula (9) can be further evolved into:
[0096] (10);
[0097] As can be seen from formula (10), the influence of the error of the modulus length on MSE is of the square order. In the case of clustering, is usually small. Under the Taylor expansion and is approximated. Therefore, the influence of the direction on MSE is of the linear order. In addition, the embodiments of the present invention also conducted the following verification experiments: statistically all the direction errors of vector quantization as shown in formula (11) and the modulus length errors as shown in formula (12):
[0098] (11);
[0099] (12);
[0100] Where: represents the direction error of vector quantization; represents the modulus length error; represents the sine function, represents the absolute value operation.
[0101] The errors in the direction and modulus length are reflected by a visualization method in Figure 7 . It is found that in almost 100% of the cases, the direction error is three times the modulus length error, which shows that the traditional clustering method is more inclined to make the modulus length approximate.
[0102] To solve the above problems, this embodiment proposes a vector quantization method using polar coordinate representation, which converts the vector representation in the original Cartesian coordinate system into a combination of direction and modulus length, and quantizes the respective sets of direction and modulus length separately. This method can allocate more clustering centers for the direction, so as to adapt to the great quantization sensitivity of the direction. For any -dimensional vector , its polar coordinate representation is the modulus length and the direction :
[0103] (13);
[0104] (14).
[0105] When constructing the vector codebook, the direction and modulus length are quantized separately; when quantizing and encoding, the optimal direction and the optimal modulus length are selected from the direction codebook and the modulus length codebook respectively, and combined into the optimal vector.
[0106] In this embodiment, the direction and magnitude of the vector codebook are separated and quantized independently, reducing the mutual interference between them. After Gaussianization, the direction and magnitude of the weight parameters of the large language model are independent random variables. Therefore, quantization in polar coordinates conforms to its distribution characteristics. In addition, the importance ratio can be adjusted by adjusting the number of bits allocated to the direction and magnitude in the index to adapt to the different sensitivities of the direction and magnitude to quantization.
[0107] S203: Construct the direction codebook and magnitude codebook for each of the polar coordinate representations respectively, obtain a vector codebook that conforms to the distribution characteristics and quantization sensitivity characteristics, and transmit it to the second storage unit for storage;
[0108] Among them, the storage format of the vector codebook in the second storage unit is direction index, magnitude index, and element value; the direction index and magnitude index are concatenated to form the final index, and the final index is integrated with the original weight parameters of the large language model as a unit to form a quantized model weight matrix; the quantized model weight matrix is used to provide to the calculation unit to run the large language model on the calculation unit.
[0109] Among them, the directions of the Gaussian distribution are uniform on each magnitude. Therefore, it is necessary to construct a direction codebook with uniform spatial directions. The definition of direction uniformity is: maximizing the minimum adjacent angle within this set of directions. The present invention proposes to use a greedy algorithm based on the E8 lattice directions to construct a spatially uniform direction codebook. The E8 lattice is the only positive definite, even, unimodular lattice in 8-dimensional space. Its definition requires that all coordinates are integers or half-integers and the sum is even, as shown in formula (15):
[0110] (15);
[0111] Among them: represents the eight-dimensional integer set, is used to represent an arbitrary vector; represents taking the union; represents taking the intersection.
[0112] It has self-duality, that is, its structure completely coincides with its dual lattice. This property stems from the unimodular property (determinant is ±1). In the sphere packing problem, the E8 lattice achieves the maximum density packing in 8-dimensional space. At a given magnitude, the directions from the origin to the E8 lattice points are spatially uniform. However, its quantity may not necessarily meet the requirements of the number of bits.
[0113] As Figure 5 shown, the process of constructing the direction codebook for each of the polar coordinate representations includes:
[0114] S20311: Select the vector directions of the E8 lattice at a given magnitude as the base class;
[0115] S20312: Select a specified number of subsets from the base class using the greedy algorithm according to the quantization bits of the direction, that is, obtain the direction codebook.
[0116] Among them, the process of selecting a specified number of subsets from the base class using the greedy algorithm is as follows:
[0117] Step a1: Initialize a selected set A and add any one direction.
[0118] Step a2: Take the remaining directions as the candidate set B.
[0119] Step a3: Select a direction from set B and add it to set A, and the sum of the angles between this direction and all directions in set A is the largest.
[0120] Step a4: Repeat step a3 until the number of directions in set A meets the requirements.
[0121] This embodiment improves the uniformity and selection flexibility of the direction. The E8 lattice is the optimal spherical packing scheme in eight-dimensional space. The existing scheme uses the vectors within the limited modulus length of the E8 lattice as the codebook. However, this approach ignores the phase problem between the vector directions at different modulus lengths and cannot represent as many directions as possible to the greatest extent. In contrast, taking the vector directions of the E8 lattice at a given modulus length increases the difference between the centers of the direction codebook and improves the modeling ability in terms of direction. In addition, the number of directions of E8 at each type of modulus length is fixed. Using the greedy algorithm helps to improve the flexibility of direction bit selection and is conducive to adaptively coordinating the important relationship between direction and modulus length.
[0122] Considering that the modulus length of the model weight parameter vector after Gaussianization conforms to the root distribution of the chi-square distribution, the present invention proposes a Lloyd-MAX algorithm based on distribution characteristics. As Figure 6 shown, the process of constructing the modulus length codebook for each of the polar coordinate representations includes:
[0123] S20321: According to the fact that the weight vector conforming to the standard Gaussian distribution conforms to the root distribution of the chi-square distribution, obtain the distribution of the modulus length, and then calculate the probability density function of the modulus length.
[0124] The definition of the chi-square distribution of dimension is the sum of the squares of standard Gaussian random variables, so its root distribution is the distribution of the modulus length. Suppose there is a modulus length variable such that obeys the chi-square distribution with degrees of freedom . First, the probability density function (Probability Density Function, PDF) of the chi-square distribution is:
[0125] (16);
[0126] where \(y > 0\), is the gamma function and \(e\) is the natural number. Since and is non - negative, the PDF of can be derived by variable transformation. The transformation function is:
[0127] (17);
[0128] Its inverse function is:
[0129] (18);
[0130] According to the variable transformation formula, we have:
[0131] (19);
[0132] where: is the derivative of.
[0133] where is the PDF of. Substituting formula (18) and its derivative, we get:
[0134] (20);
[0135] Substitute with , we obtain:
[0136] (21);
[0137] Solving for gives the expression:
[0138] (22);
[0139] Substituting into formula (16) gives:
[0140] (23);
[0141] Formula (23) characterizes the magnitude distribution of the Gaussianized model weight parameter vector.
[0142] S20322: Quantize the magnitude using the Llyod - MAX algorithm according to the probability density function of the magnitude to form a magnitude codebook.
[0143] Among them, the process of quantifying the modulus length using the Lloyd-MAX algorithm is as follows:
[0144] Step b1: Select the initial quantization lattice points and quantization boundaries; select the initial quantization lattice points and quantization boundaries .
[0145] Step b2: According to the initial quantization lattice points and quantization boundaries, calculate the centroid and boundaries in sequence;
[0146] The calculation formula for the centroid is: ;
[0147] The calculation formula for the boundary is: ;
[0148] Step b3: Determine whether the Lloyd-MAX algorithm converges. If the change in quantization distortion between two adjacent iterations is less than the preset threshold or the preset number of iterations is reached, stop the iteration; otherwise, return to step b2 to continue the iteration until the Lloyd-MAX algorithm converges, and finally obtain the quantized codebook lattice points.
[0149] Such as Figure 8 the codebook lattice points shown taking the two-dimensional space as an example. Figure 8 The blue line in it represents the direction of spatial uniformity, and the orange line represents the modulus length that conforms to the root distribution of the chi-square distribution.
[0150] It should also be noted here that the first storage unit and the second storage unit can be implemented using other memories such as a disk or computer storage space (such as memory).
[0151] In this embodiment, only the direction index, modulus length index, and element values of the model weight vector are stored in the second storage unit. This reduces the space occupied by model storage, and can effectively reduce the model storage requirement to 12.5% of the original model. For example, a 7B model only requires 1.7 GB of disk space to complete the model storage.
[0152] During the deployment process of the quantized model, the following steps may be included:
[0153] Load the quantized model from the second storage unit into the GPU (Graphics Processing Unit) video memory;
[0154] Construct the direction and modulus length codebooks and load them into the L1 cache (primary cache) of the computing unit;
[0155] According to the model inference process, load the indexes from the computing unit video memory into the L1 cache (primary cache) in sequence;
[0156] Divide the index into a direction index and a magnitude index according to the bits of the direction and magnitude, and the left and right positions.
[0157] Decode the direction and magnitude from the direction and magnitude codebooks respectively according to the direction and magnitude indexes.
[0158] Send the direction and magnitude to the computing unit for operation with the activation value.
[0159] In this embodiment, during the process of quantized model inference, the direction and magnitude codebooks are loaded into the L1 cache of the GPU only once. During each inference process, only the direction and magnitude indexes of the vector are loaded from the video memory to the L1 cache for table lookup, and then the table lookup result is sent to the computing unit for operation with the activation.
[0160] This embodiment does not need to move the codebooks between the L1 cache and the GPU video memory multiple times, and transmits information through indexes, reducing the memory access time of the system and improving the inference efficiency of the large model (for example, about 8 times of inference acceleration can be achieved on the LLaMA-2-70B model).
[0161] Figure 9 The structural schematic diagram of the large model quantization and compression system based on the polar coordinate system according to the embodiment of the present invention is given. Combining Figure 9 , a large model quantization and compression system based on the polar coordinate system according to an embodiment of the present invention at least includes:
[0162] A weight parameter distribution standardization module 901, which is used to retrieve the original weight parameter distribution of the pre-stored large language model from the first storage unit and process it into a standard Gaussian distribution to obtain a corresponding weight vector that conforms to the standard Gaussian distribution.
[0163] A weight parameter polar coordinate conversion module 902, which is used to convert each of the weight vectors that conform to the standard Gaussian distribution into a corresponding polar coordinate representation.
[0164] A vector codebook construction and storage module 903, which is used to construct a direction codebook and a magnitude codebook for each of the polar coordinate representations respectively, obtain a vector codebook that conforms to the distribution characteristics and quantization sensitivity characteristics, and transmit it to the second storage unit for storage.
[0165] Wherein, the storage format of the vector codebook in the second storage unit is a direction index, a magnitude index, and an element value; the direction index and the magnitude index are spliced to form a final index, and the final index is integrated in units of the original weight parameters of the large language model to form a quantized model weight matrix; the quantized model weight matrix is used to provide to the computing unit to run the large language model on the computing unit.
[0166] It should be noted here that each module in the large model quantization and compression system based on the polar coordinate system in the embodiments of the present invention corresponds one by one to each step in the above-mentioned large model quantization and compression method based on the polar coordinate system, and the specific implementation process is the same, so it will not be elaborated here.
[0167] For the large model quantization and compression method and system based on the polar coordinate system in this embodiment, first, the enhanced Hadamard rotation method is used to process the original weight parameter distribution into an approximately standard Gaussian distribution, so that the morphological representation of the distribution is solvable; secondly, the rotated weight parameter vector is represented in polar coordinates, and it is proposed to quantize the direction and the modulus length respectively to adapt to the quantization sensitivity difference between the direction and the modulus length; finally, based on considering the data distribution characteristics after rotation, construction schemes are designed for the direction codebook and the modulus length codebook respectively, so that the codebook distribution approximately conforms to the original distribution and reduces the quantization error; through the storage and loading methods of the quantization model of the present invention, the deployment and inference efficiency of the large model are improved.
[0168] In one or more embodiments, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the large model quantization and compression method based on the polar coordinate system as described above Figure 2 shown.
[0169] In one or more embodiments, a computer program product is provided, including a computer program / instructions, and when the computer program / instructions are executed by a processor, it implements the steps in the large model quantization and compression method based on the polar coordinate system as described above Figure 2 shown.
[0170] In one or more embodiments, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the large model quantization and compression method based on the polar coordinate system as described above Figure 2 shown.
[0171] The present invention is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products of the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of the flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.
[0172] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A large model quantization and compression method based on the polar coordinate system, characterized in that, Comprising: Retrieve the original weight parameter distribution of the pre-stored large language model from the first storage unit, and process it into a standard Gaussian distribution to obtain the corresponding weight vectors conforming to the standard Gaussian distribution; Convert each of the weight vectors conforming to the standard Gaussian distribution into corresponding polar coordinate representations; Construct a direction codebook and a magnitude codebook for each of the polar coordinate representations respectively, obtain a vector codebook conforming to the distribution characteristics and quantization sensitivity characteristics, and transmit it to the second storage unit for storage; Among them, the storage format of the vector codebook in the second storage unit is direction index, magnitude index, and element value; the direction index and the magnitude index are concatenated to form a final index, and the final index is integrated in units of the original weight parameters of the large language model to form a quantized model weight matrix; the quantized model weight matrix is used to be provided to the computing unit to run the large language model on the computing unit; Among them, the process of constructing the direction codebook for each of the polar coordinate representations includes: Select the vector directions of the E8 lattice at a given magnitude as the base class; Use the greedy algorithm to select a specified number of subsets from the base class according to the quantization bits of the direction to obtain the direction codebook; The process of constructing the magnitude codebook for each of the polar coordinate representations includes: According to the root distribution of the chi-square distribution that the weight vectors conforming to the standard Gaussian distribution conform to, obtain the distribution of the magnitude, and then calculate the probability density function of the magnitude; Quantize the magnitude according to the probability density function of the magnitude using the Llyod-MAX algorithm to form the magnitude codebook; The process of quantizing the magnitude using the Llyod-MAX algorithm is: Select the initial quantization grid points and quantization boundaries; Calculate the centroid and boundaries in turn according to the initial quantization grid points and quantization boundaries; Judge whether the Llyod-MAX algorithm converges. If the quantization distortion change between two adjacent iterations is less than the preset threshold or reaches the preset number of iterations, stop the iteration; otherwise, continue to calculate the centroid and boundaries iteratively until the Llyod-MAX algorithm converges, and finally obtain the quantized codebook grid points.
2. The large model quantization and compression method based on the polar coordinate system according to claim 1, wherein The process of processing the original weight parameter distribution of the large language model into a standard Gaussian distribution includes: Calculate the covariance matrix of the original weight parameters of the large language model; Perform the Karhunen-Loeve transform on the covariance matrix to obtain an enhanced matrix; Multiply the enhanced matrix by the Hadamard matrix for fusion to obtain an enhanced rotation matrix; Among them, the parameter distributions in the enhanced rotation matrix are the standard Gaussian distribution.
3. The large model quantization and compression method based on the polar coordinate system according to claim 1, characterized in that The weight vectors conforming to the standard Gaussian distribution contain two variables, namely: direction and magnitude.
4. A large model quantization and compression system based on the polar coordinate system, characterized in that, Comprising: A weight parameter distribution standardization module, which is used to retrieve the original weight parameter distribution of the pre-stored large language model from the first storage unit and process it into a standard Gaussian distribution to obtain the corresponding weight vectors conforming to the standard Gaussian distribution; A weight parameter polar coordinate conversion module, which is used to convert each of the weight vectors conforming to the standard Gaussian distribution into corresponding polar coordinate representations; A vector codebook construction and storage module, which is used to construct the direction codebook and the magnitude codebook of each of the polar coordinate representations respectively, obtain a vector codebook that conforms to the distribution characteristics and quantization sensitivity characteristics, and transmit it to the second storage unit for storage; Among them, the storage format of the vector codebook in the second storage unit is direction index, magnitude index, and element value; the direction index and the magnitude index are spliced to form a final index, and the final index is integrated in units of the original weight parameters of the large language model to form a quantized model weight matrix; the quantized model weight matrix is used to provide to the calculation unit to run the large language model on the calculation unit; Among them, the process of constructing the direction codebook of each of the polar coordinate representations includes: Select the vector directions of the E8 lattice at a given magnitude as the base class; Select a specified number of subsets from the base class using the greedy algorithm according to the quantization bits of the direction, that is, obtain the direction codebook; The process of constructing the magnitude codebook of each of the polar coordinate representations includes: According to the root distribution of the chi-square distribution of the weight vector conforming to the standard Gaussian distribution, obtain the distribution of the magnitude, and then calculate the probability density function of the magnitude; According to the probability density function of the magnitude, use the Llyod-MAX algorithm to quantize the magnitude to form the magnitude codebook; The process of quantizing the magnitude using the Llyod-MAX algorithm is: Select the initial quantization lattice points and quantization boundaries; According to the initial quantization lattice points and quantization boundaries, calculate the centroid and boundaries in turn; Judge whether the Llyod-MAX algorithm converges. If the quantization distortion change between two adjacent iterations is less than the preset threshold or reaches the preset number of iterations, stop the iteration; otherwise, continue to calculate the centroid and boundaries iteratively until the Llyod-MAX algorithm converges, and finally obtain the quantized codebook lattice points.
5. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the large model quantization and compression method based on the polar coordinate system described in any one of claims 1-3.
6. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, it implements the steps in the large model quantization and compression method based on the polar coordinate system described in any one of claims 1-3.
7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the large model quantization and compression method based on the polar coordinate system described in any one of claims 1-3.
Citation Information
Patent Citations
Large model quantization algorithm based on low-rank dictionary
CN118095371A
Upsampling of compressed financial time-series data using a jointly trained Vector Quantized Variational Autoencoder neural network
US12229679B1