Model quantification method, device, equipment, medium and computer program product
By iterating the weight matrix of the large language model, combining the weight block data and codebook data, the large language model is quantified, which solves the problem of difficult balance between precision and data volume compression in the existing technology, and realizes efficient model deployment and inference in resource-constrained environments.
Patent Information
- Application Number
- CN202510004236.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-13
AI Technical Summary
The existing large-language model quantization methods cannot maintain model accuracy while reducing the amount of model data, and cannot achieve a balance between precision and data volume compression.
Through an iterative quantization method based on the weight matrix, the weight block data and codebook data are determined, and the weight blocks are iteratively quantized to obtain the quantitative model. At the same time, the confusion comparison between the to-process model and the quantization model is compared to obtain the model quantization results.
On the basis of reducing the amount of model data, maintain or approach the accuracy of the original model, reduce the storage space and computing resource requirements, and improve the model's inference speed and efficiency, which is suitable for resource-constrained environments.
Smart Images

Figure CN119990221A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of model quantization technology, and in particular to a model quantization method, device, equipment, medium and computer program product. Background Art
[0002] A large language model generally refers to a deep learning model obtained by training a large number of data samples, which can be used to process natural language tasks. However, the large language model itself has a large amount of data, which makes the execution efficiency of the large language model on specific hardware low. In addition, the storage and computing costs of the large language model are high, making it impossible for the large language model to run in a resource-constrained environment. Quantizing the large language model can reduce the data volume of the model, improve the execution efficiency of the large language model, and reduce the model storage and computing costs. However, the existing large language model quantization method cannot achieve a balance between model accuracy and data volume compression. Summary of the invention
[0003] The present invention provides a model quantization method, device, equipment, medium and computer program product, which are used to solve the defect of existing large language model quantization technology that it is impossible to achieve a balance between model accuracy and data volume compression, and to achieve quality maintenance of model accuracy on the basis of reducing the model data volume.
[0004] The present invention provides a model quantization method, comprising the following steps.
[0005] Determine a weight matrix based on the acquired weight information of the model to be processed; Based on the weight matrix, determining weight block data and codebook data of the model to be processed; Based on the weight block data and the code book data, iteratively quantize the weight block to obtain a quantization model; The perplexity of the model to be processed and the quantized model is compared to obtain a model quantization result.
[0006] According to a model quantization method provided by the present invention, the weight block data includes the number of weight blocks; the code book data includes the group size, the number of columns processed by each code book and the number of code books; the weight block data and code book data of the model to be processed are determined based on the weight matrix, including: Based on the number of rows of the weight matrix W and number of columns , and preset block sizes , determine the number of weight blocks i ;in, ; In the case where the group size is In the case of , the number of columns processed by each code book is determined as , the number of codebooks is ; Wherein, the initialization code book Ci is a zero matrix of d rows and k columns, and the initialization code book is used to store the code book of each group; ; .
[0007] According to a model quantization method provided by the present invention, the weight block is iteratively quantized based on the weight block data and the code book data to obtain a quantization model, which includes: Initialize the quantized weight matrix Q and quantization error E to a zero matrix with r rows and c columns; Processing the input of the weight block of the current layer of the model to be processed to obtain a matrix decomposition expression Hinv; For the currently processed column Quantify and get the quantitative results ; Based on the formula Determine the columns The quantization error ,in, For Column The original matrix of Column-based The quantization error The weight block is iteratively quantized to obtain a quantized model.
[0008] According to a model quantization method provided by the present invention, the input of the weight block of the current layer of the model to be processed is processed to obtain a matrix decomposition expression Hinv including: Determine a Hessian matrix H based on an input X of a weight block of a current layer of the model to be processed; Perform Cholesky decomposition on the inverse matrix of the Hessian matrix to obtain a matrix decomposition expression Hinv.
[0009] According to a model quantization method provided by the present invention, the currently processed column Quantify and get the quantitative results include: In the case where the currently processed weight block is the first weight block in the group, the group index g is determined; where, ; Process the xth column to the x+m-1th column in the currently processed weight block and the xth column to the x+m-1th column in the scaling factor matrix to initialize the codebook Cg of the group; Performing operation processing on the corresponding columns in the weight matrix W and the scaling factor matrix, and performing element-by-element operation on the result of the operation processing and the initialized code book Cg to obtain an operation value; Mapping the operation value to a vector closest to the operation value in the code book Cg to obtain a column quantization result; The column quantization result is stored in the corresponding column range of the weight matrix Q to obtain the quantization result .
[0010] According to a model quantization method provided by the present invention, the perplexity comparison between the model to be processed and the quantized model to obtain the model quantization result includes: Based on the formula Determining the perplexity ; Perplexity Including the first perplexity obtained by the model to be processed to predict the target text, and the second perplexity obtained by the quantization model to predict the target text; wherein N is the total number of words in the target text; is the predicted probability of the ath word in the target text; is the context word of the ath word in the target text; Based on a comparison result of the first perplexity and the second perplexity, a model quantization result is determined.
[0011] The present invention also provides a model quantization device, comprising the following modules: A weight matrix determination module, used to determine a weight matrix based on the acquired weight information of the model to be processed; A module for determining data to be processed, used to determine weight block data and codebook data of the model to be processed based on the weight matrix; An iterative quantity module, used for iteratively quantizing the weight block based on the weight block data and the code book data to obtain a quantized model; The model quantization module is used to compare the perplexity of the model to be processed and the quantized model to obtain a model quantization result.
[0012] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the model quantization method described above is implemented.
[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the model quantization method described in any one of the above is implemented.
[0014] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the model quantization method described above is implemented.
[0015] The model quantization method, device, equipment, medium and computer program product provided by the present invention initialize the relevant data of each weight block in the weight matrix of the model to be processed, iteratively optimize each weight block in the weight matrix of the model to be processed, obtain the quantized model, and compare the perplexity of the model to be processed before quantization and the model after quantization to obtain the model quantization result. The model quantization method provided by the present application can reduce the storage space and required computing resources of a large language model, improve the reasoning speed and efficiency of the model, and thus realize efficient deployment and reasoning of the model in a resource-constrained environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0017] Figure 1 This is one of the flow charts of the model quantization method provided by the present invention.
[0018] Figure 2 This is the second flow chart of the model quantization method provided by the present invention.
[0019] Figure 3 It is a structural schematic diagram of the model quantization device provided by the present invention.
[0020] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0022] Combine the following Figure 1-Figure 4 The present invention describes the model quantization method, device, apparatus, medium and computer program product.
[0023] Figure 1 This is one of the flow charts of the model quantization method provided by the present invention, such as Figure 1 As shown, the method includes the following: Step 100: Determine a weight matrix based on the acquired weight information of the model to be processed; Specifically, the Large Language Model (LLM) can handle complex natural language tasks, but its high computational cost leads to a high inference delay. In order to improve the inference speed of the large language model, the large language model can be quantized, and the weight matrix and activation matrix of the large language model can be quantized to a lower precision than the original model, so as to use faster low-precision computing units to accelerate the inference of the large language model.
[0024] Taking the weight matrix as an example, the entire weight matrix can share the same quantization parameters (such as scaling factors and quantization zero points). Since there will be some outlier parameter data in the weight matrix (for example, large / small weight values), this method is difficult to quantize the entire weight matrix to 4 bits or lower bit widths while ensuring accuracy, and it is difficult to fully utilize the same computing unit of the processor to further accelerate the reasoning of large language models. The reason for the low accuracy is that the weight data range of the large language model is too large. When the large language model after parameter quantization is 4 bits, outlier parameter data will cause significant quantization errors.
[0025] In order to reduce the quantization error caused by outlier parameter data, different quantization bit widths can be used for outlier parameter data and normal data (for example, 16 bits for outlier parameter data and 8 bits for normal data). However, in the inference process of a large language model, this method will cause different threads to perform computing tasks with different load sizes and require calling different computing units, which is not convenient for deploying large language models on general-purpose processors (such as graphics processing units (GPUs)) and for efficient parallel computing.
[0026] Step 200: Determine the weight block data and codebook data of the model to be processed based on the weight matrix; The method proposed in this application for iteratively quantizing model weights mainly includes the following steps.
[0027] Step 1: Initialize the relevant data and codebook data of each weight block.
[0028] The weight block data includes the number of weight blocks into which the weight matrix is divided; the codebook data includes the group size of the codebook, the number of columns processed by each codebook, and the number of codebooks.
[0029] The code book is initialized. The initialized code book Ci is a zero matrix with d rows and k columns. The initialized code book is used to store the code book (also called code book) of each group.
[0030] Step 300: iteratively quantize the weight block based on the weight block data and the codebook data to obtain a quantization model; Specifically, the method for gradually quantizing model weights through iteration proposed in the present application also includes the following steps.
[0031] Step 2: Iterate and quantize each weight block i one by one. Step 2 mainly includes matrix decomposition, iterative processing of columns in the weight block, and quantization of columns in the weight block. After all columns are quantized, the quantized weight matrix Q is obtained, and the quantized model corresponding to the model to be processed is obtained.
[0032] Step 400: compare the perplexity of the model to be processed and the quantized model to obtain a model quantization result.
[0033] Step 3: For the evaluation of the quantitative model, this application proposes a method for quantitative evaluation based on the perplexity of the model before and after quantization. In information theory, perplexity is used to measure the quality of a probability distribution or probability model in predicting samples. Perplexity can also be used to compare two probability distributions or probability models, comparing the pros and cons of the two in predicting samples. The probability distribution model or probability model with low perplexity can better predict the sample. That is, the lower the perplexity, the better the prediction effect of the model.
[0034] This embodiment initializes the relevant data of each weight block in the weight matrix of the model to be processed, iteratively optimizes each weight block in the weight matrix of the model to be processed, obtains a quantized model, and compares the perplexity of the model to be processed before quantization and the model after quantization to obtain a model quantization result. The model quantization method provided in this application can reduce the storage space and required computing resources of a large language model, improve the reasoning speed and efficiency of the model, and thus achieve efficient deployment and reasoning of the model in a resource-constrained environment.
[0035] In one embodiment, the model quantization method provided in the embodiment of the present application may further include: Step 210: Based on the number of rows of the weight matrix W and number of columns , and preset block sizes , determine the number of weight blocks i ;in, ; Step 220: When the group size is In the case of , the number of columns processed by each code book is determined as , the number of codebooks is ; Wherein, the initialization code book Ci is a zero matrix of d rows and k columns, and the initialization code book is used to store the code book of each group; ; .
[0036] Specifically, the above step 1 also includes the following steps.
[0037] Step 1.1, calculate the number of weight blocks: For a weight W with r rows and c columns, assuming it is divided into weight blocks of size r×B, the number of each weight block i can be calculated , the number of weight blocks indicates how many blocks the weight W is divided into.
[0038] Step 1.2: Calculate the number of columns in each codebook group: Assuming the group size of the codebook is l, and the codebook is to be divided into groups (codebooks) with l columns, the number of columns in each codebook group can be calculated. , codebooks, the number of codebooks .
[0039] Step 1.3, initialize the quantized weight Q and quantization error E, the quantized weight Q and quantization error E are both zero matrices with r rows and c columns.
[0040] Step 1.4, initialize the codebook Ci of each group: Ci is a zero matrix with d rows and k columns. The initialized codebook Ci is used to store the codebook of each group.
[0041] This embodiment calculates weight block data and code book data and initializes them to lay a data foundation for the quantization of the model.
[0042] Figure 2 This is the second flow chart of the model quantization method provided by the present invention, such as Figure 2 As shown, the method may also include: Step 310, initializing the quantized weight matrix Q and the quantization error E to be a zero matrix with r rows and c columns; Step 320: Process the input of the weight block of the current layer of the model to be processed to obtain a matrix decomposition expression Hinv; Step 330: For the currently processed column Quantify and get the quantitative results ; Step 340: Based on the formula Determine the columns The quantization error ,in, For Column The original matrix of Step 350: Column-based The quantization error The weight block is iteratively quantized to obtain a quantized model.
[0043] Specifically, the content of the above step 2 also includes: Step 2.1: Perform matrix decomposition on the input X based on the weights of this layer of the model to obtain the matrix decomposition expression Hinv.
[0044] Step 2.2: Initialize the code book Ci.
[0045] Step 2.3: Iterate the columns in the current weight block to obtain the quantized results. ,List The quantization error and columns The original matrix, then based on the column The quantization error The weight blocks are iteratively quantized to obtain a quantized model.
[0046] This embodiment iterates the columns in the current weight block one by one and updates the quantized weight matrix to obtain a quantized model.
[0047] In one embodiment, the model quantization method provided in the embodiment of the present application may further include: Step 321: Determine the Hessian matrix H based on the input X of the weight block of the current layer of the model to be processed; Step 322: Perform Cholesky decomposition on the inverse matrix of the Hessian matrix to obtain a matrix decomposition expression Hinv.
[0048] Specifically, the content of the above step 2.1 specifically includes: Based on the input X of the weight of the current layer of the model to be processed, the Hessian matrix H is calculated, and then the inverse matrix of the Hessian matrix H is calculated. Perform Cholesky decomposition and get The Cholesky expression Hinv. Among them, Cholesky decomposition is the decomposition of a symmetric positive definite matrix into a product of a lower triangular matrix L and its transpose. Cholesky decomposition requires that all eigenvalues of the matrix must be greater than zero, so the diagonal elements of the decomposed lower triangle are also greater than zero. Cholesky decomposition is also called square root method.
[0049] In this embodiment, a matrix decomposition expression Hinv is obtained through Cholesky decomposition. The matrix decomposition expression Hinv can be used to calculate the column quantization error Ep.
[0050] In one embodiment, the model quantization method provided in the embodiment of the present application may further include: Step 331: When the weight block currently being processed is the first weight block in the group, determine the group index g; wherein, ; Step 332: Process the xth column to the x+m-1th column in the currently processed weight block and the xth column to the x+m-1th column in the scaling factor matrix to initialize the codebook Cg of the group; Step 333: perform operation processing on the corresponding columns of the weight matrix W and the scaling factor matrix, and perform element-by-element operation on the result of the operation processing and the initialized code book Cg to obtain an operation value; Step 334: Map the operation value to the vector closest to the operation value in the code book Cg to obtain a column quantization result; Step 335: Store the column quantization result in the corresponding column range of the weight matrix Q to obtain the quantization result .
[0051] The content of the above step 2.2 specifically includes: If the currently processed weight block is the first block in a group (i%m=0), then determine the group index , initialize the code book Cg of the group, specifically, divide the weight block from the xth column to the x+m-1th column by the (corresponding) value from the xth column to the ix+m-1th column of the scaling factor matrix to initialize the code book Cg.
[0052] The content of the above step 2.3 specifically includes: Step 2.3.1: Determine the column range p currently being processed.
[0053] Step 2.3.2: Use the vector quantization method to quantize the weight column p to obtain the quantized column result Qp.
[0054] Step 2.3.3, by formula Calculate the quantization error Ep, which means that the error Ep of column p is determined by the original weight of column p After quantification The difference is based on the Hinv obtained above and multiplied by it to weight the error.
[0055] Step 2.3.4: Update the weights of the current column range in the weight matrix, minus the weighted quantization error Ep.
[0056] Step 2.4: Update the weights of the columns after the current weight block, minus the weighted sum of the errors of the entire block.
[0057] Step 2.5: Finally, return the quantized weight matrix Q.
[0058] The specific contents of the above step 2.3.2 include: selecting a specific column from the weight matrix W, dividing it by the corresponding column in the scaling factor matrix, using the divided value and the above initialized code book Cg to perform element-by-element multiplication operation, mapping these values obtained by the element-by-element multiplication operation to the closest vector in the code book, and then storing the quantized results in the corresponding column range of the matrix Q.
[0059] This embodiment obtains a quantization model by iteratively processing each weight block.
[0060] In one embodiment, the model quantization method provided in the embodiment of the present application may further include: Step 410: Based on the formula Determining the perplexity ; Perplexity Including the first perplexity obtained by the model to be processed to predict the target text, and the second perplexity obtained by the quantization model to predict the target text; wherein N is the total number of words in the target text; is the predicted probability of the ath word in the target text; is the context word of the ath word in the target text; Step 420: Determine a model quantization result based on a comparison result of the first perplexity and the second perplexity.
[0061] Specifically, the specific content of the above step 3 is as follows: Comparing the perplexity of the quantized models , and compare it with the different parameter versions of the model before quantization to evaluate the model quantization effect. The perplexity is calculated based on the likelihood (prediction probability) of the model for the target text. The lower the perplexity, the more accurate the model's prediction. Specifically, it is the relationship between the probability distribution predicted by the model and the true distribution. The specific calculation is shown in Formula 1.
[0062] ; (1)
[0063] Where N is the total number of words in the target text; is the predicted probability of the ath word in the target text; is the context word of the ath word in the target text.
[0064] This embodiment measures the change in accuracy before and after the evaluation model is quantized by comparing the perplexity, provides a data basis for flexibly adjusting the process of model quantization, and also enhances the flexibility of model quantization.
[0065] The model quantization device provided by the present invention is described below. The model quantization device described below and the model quantization method described above can be referenced to each other.
[0066] Please refer to Figure 3 The present invention also provides a model quantization device, comprising: A weight matrix determination module 301 is used to determine a weight matrix based on the acquired weight information of the model to be processed; The to-be-processed data determination module 302 is used to determine the weight block data and codebook data of the to-be-processed model based on the weight matrix; An iteration module 303, configured to iteratively quantize the weight block based on the weight block data and the codebook data to obtain a quantization model; The model quantization module 304 is used to compare the perplexity of the model to be processed and the quantized model to obtain a model quantization result.
[0067] Optionally, the weight block data includes the number of weight blocks; the code book data includes the group size, the number of columns processed by each code book and the number of code books; the module for determining data to be processed includes: A weight block data determination unit for determining the weight block data based on the number of rows of the weight matrix W. and number of columns , and preset block sizes , determine the number of weight blocks i ;in, ; The codebook data determination unit is used to determine the group size of In the case of , the number of columns processed by each code book is determined as , the number of codebooks is ; Wherein, the initialization code book Ci is a zero matrix of d rows and k columns, and the initialization code book is used to store the code book of each group; ; .
[0068] Optionally, the iterative quantization module includes: A matrix initialization unit, used to initialize the quantized weight matrix Q and the quantization error E into a zero matrix with r rows and c columns; A matrix decomposition unit, used for processing the input of the weight block of the current layer of the model to be processed to obtain a matrix decomposition expression Hinv; Matrix column quantization unit, used to quantize the currently processed column Quantify and get the quantitative results ; The quantization error determination unit is used to determine the Determine the columns The quantization error ,in, For Column The original matrix of Iterative quantization unit for column-based The quantization error The weight block is iteratively quantized to obtain a quantized model.
[0069] Optionally, the matrix decomposition unit includes: A Hessian matrix determining unit, configured to determine a Hessian matrix H based on an input X of a weight block of a current layer of the model to be processed; The inverse matrix decomposition unit is used to perform Cholesky decomposition on the inverse matrix of the Hessian matrix to obtain a matrix decomposition expression Hinv.
[0070] Optionally, the matrix column quantization unit includes: A group index determination unit, used to determine the group index g when the weight block currently being processed is the first weight block in the group; wherein, ; A codebook initialization unit, used to process the xth column to the x+m-1th column in the currently processed weight block and the xth column to the x+m-1th column in the scaling factor matrix, and initialize the codebook Cg of the group; An operation value determination unit, configured to perform operation processing on the corresponding columns in the weight matrix W and the scaling factor matrix, and perform element-by-element operation on the result of the operation processing and the initialized code book Cg to obtain an operation value; A column quantization result determination unit, configured to map the operation value to a vector closest to the operation value in the code book Cg, to obtain a column quantization result; A column quantization result determination unit is used to store the column quantization result in the corresponding column range of the weight matrix Q to obtain a quantization result .
[0071] Optionally, the model quantization module includes: The perplexity determination unit is used to determine the perplexity based on the formula Determining the perplexity ; Perplexity Including the first perplexity obtained by the model to be processed to predict the target text, and the second perplexity obtained by the quantization model to predict the target text; wherein N is the total number of words in the target text; is the predicted probability of the ath word in the target text; is the context word of the ath word in the target text; A perplexity comparison unit is used to determine a model quantization result based on a comparison result of the first perplexity and the second perplexity.
[0072] Figure 4An example of a physical structure diagram of an electronic device is shown in FIG. Figure 4 As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430 and a communication bus 440, wherein the processor 410, the communication interface 420 and the memory 430 communicate with each other through the communication bus 440. The processor 410 may call the logic instructions in the memory 430 to execute the model quantization method, which includes: determining a weight matrix based on the acquired weight information of the model to be processed; determining the weight block data and code book data of the model to be processed based on the weight matrix; iteratively quantizing the weight block based on the weight block data and the code book data to obtain a quantized model; and comparing the perplexity of the model to be processed and the quantized model to obtain a model quantization result.
[0073] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0074] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the model quantization method provided by the above methods, which includes: determining a weight matrix based on the acquired weight information of the model to be processed; determining weight block data and code book data of the model to be processed based on the weight matrix; iteratively quantizing the weight block based on the weight block data and the code book data to obtain a quantized model; comparing the perplexity of the model to be processed and the quantized model to obtain a model quantization result.
[0075] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it is implemented to execute the model quantization method provided by the above-mentioned methods, the method comprising: determining a weight matrix based on the acquired weight information of the model to be processed; determining weight block data and code book data of the model to be processed based on the weight matrix; iteratively quantizing the weight block based on the weight block data and the code book data to obtain a quantized model; comparing the perplexity of the model to be processed and the quantized model to obtain a model quantization result.
[0076] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0077] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A model quantization method, characterized in that: include: Determine a weight matrix based on the acquired weight information of the model to be processed; Based on the weight matrix, determining weight block data and codebook data of the model to be processed; Based on the weight block data and the code book data, iteratively quantize the weight block to obtain a quantization model; The perplexity of the model to be processed and the quantized model is compared to obtain a model quantization result.
2. The model quantization method according to claim 1, characterized in that: The weight block data includes the number of weight blocks; the code book data includes the group size, the number of columns processed by each code book and the number of code books; the weight block data and code book data of the model to be processed are determined based on the weight matrix, including: Based on the number of rows of the weight matrix W and the number of columns , and preset block sizes , determine the number of weight blocks i ;in, ; In the case where the group size is In the case of , the number of columns processed by each code book is determined as , the number of codebooks is ; Wherein, the initialization code book Ci is a zero matrix of d rows and k columns, and the initialization code book is used to store the code book of each group; ; .
3. The model quantization method according to claim 2, characterized in that: The iterative quantization of the weight block based on the weight block data and the code book data to obtain a quantization model includes: Initialize the quantized weight matrix Q and quantization error E to a zero matrix with r rows and c columns; Processing the input of the weight block of the current layer of the model to be processed to obtain a matrix decomposition expression Hinv; For the currently processed column Quantify and get the quantitative results ; Based on the formula Determine the columns The quantization error ,in, For Column The original matrix of Column-based The quantization error The weight block is iteratively quantized to obtain a quantized model.
4. The model quantization method according to claim 3, characterized in that: The input of the weight block of the current layer of the model to be processed is processed to obtain a matrix decomposition expression Hinv including: Determine a Hessian matrix H based on an input X of a weight block of a current layer of the model to be processed; Perform Cholesky decomposition on the inverse matrix of the Hessian matrix to obtain a matrix decomposition expression Hinv.
5. The model quantization method according to claim 3, characterized in that: The currently processed column Quantify and get the quantitative results include: In the case where the currently processed weight block is the first weight block in the group, the group index g is determined; where, ; Process the xth column to the x+m-1th column in the currently processed weight block and the xth column to the x+m-1th column in the scaling factor matrix to initialize the codebook Cg of the group; Performing operation processing on the corresponding columns in the weight matrix W and the scaling factor matrix, and performing element-by-element operation on the result of the operation processing and the initialized code book Cg to obtain an operation value; Mapping the operation value to a vector closest to the operation value in the code book Cg to obtain a column quantization result; The column quantization result is stored in the corresponding column range of the weight matrix Q to obtain the quantization result .
6. The model quantization method according to claim 1, characterized in that: The comparing the perplexity of the model to be processed and the quantized model to obtain the model quantization result includes: Based on the formula Determining the perplexity ; Perplexity Including the first perplexity obtained by the model to be processed to predict the target text, and the second perplexity obtained by the quantization model to predict the target text; wherein N is the total number of words in the target text; is the predicted probability of the ath word in the target text; is the context word of the ath word in the target text; Based on a comparison result of the first perplexity and the second perplexity, a model quantization result is determined.
7. A model quantization device, characterized in that: include: A weight matrix determination module, used to determine a weight matrix based on the acquired weight information of the model to be processed; A module for determining data to be processed, used to determine weight block data and codebook data of the model to be processed based on the weight matrix; An iterative quantity module, used for iteratively quantizing the weight block based on the weight block data and the code book data to obtain a quantized model; The model quantization module is used to compare the perplexity of the model to be processed and the quantized model to obtain a model quantization result.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the model quantization method according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the model quantization method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the model quantization method according to any one of claims 1 to 6 is implemented.