Large model parameter mixing precision quantification method and device

Through the significance analysis of large model weight parameters and mixed precision quantization, the problem of inability to take into account both accuracy and efficiency in the prior art is solved, and efficient storage and calculation balance is achieved.

CN120430356APending Publication Date: 2025-08-05INST OF AUTOMATION CHINESE ACAD OF SCI +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510933198.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

In the prior art, the large-model quantization process cannot take into account both model accuracy and inference efficiency. Traditional quantization methods still require additional calculations when reducing storage requirements, or it is difficult to ensure overall accuracy when improving computing efficiency.

Method used

Through significance analysis, the weight parameters with higher significance are quantized 8-bit quantization, the parameters with lower significance are quantized 2-bit quantization, and the quantization type is recorded by the identification matrix to ensure the accuracy and storage efficiency of key parts.

Benefits of technology

On the basis of ensuring the accuracy of key parts of the model, it reduces storage requirements and improves inference efficiency, while supporting efficient access and matrix multiplication calculation of mixed storage of data of different quantized bit widths.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120430356A_ABST
    Figure CN120430356A_ABST
Patent Text Reader

Abstract

The invention discloses a large model parameter mixing precision quantification method and device. The method comprises the steps of obtaining significance of each weight parameter of a target model; performing 8-bit quantization on the weight parameters of which the saliency meets a high saliency condition to obtain a first quantization result; performing 2-bit quantization on the weight parameters of which the saliency does not meet the high saliency condition to obtain a second quantization result; based on the first quantization result and the second quantization result, obtaining a quantized weight matrix of the target model; the quantized weight matrix and the identification matrix are used for processing input data of the target model, and the identification matrix indicates the quantization type of each weight parameter in the quantized weight matrix.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of artificial intelligence technology, and more specifically, to a method and apparatus for mixed-precision quantization of large model parameters. Background Art

[0002] In recent years, the rapid development of deep learning and large-scale pre-trained models has achieved remarkable results in fields such as natural language processing and computer vision. However, the surge in the number of model weight parameters has also brought enormous storage and computational pressures, which has become a major bottleneck restricting the deployment and inference efficiency of large models in practical applications. To address this problem, quantization technology has emerged. By mapping floating-point weight parameters to a low-bit representation, it not only significantly reduces storage usage but also improves computational throughput by leveraging low-bit computing units.

[0003] Currently, the industry has proposed various quantization schemes, such as quantizing only the model's weight parameters and jointly quantizing the model's weight parameters and input data. The former has achieved some success in reducing model storage requirements, but it still requires additional dequantization operations during inference, failing to fundamentally reduce the amount of computation. The latter attempts to utilize low-bit computing units to directly perform matrix multiplication, thereby further improving computational efficiency. However, the lower quantization complexity of the model's input data makes it more difficult to ensure overall accuracy. Summary of the Invention

[0004] The embodiments of the present disclosure provide a mixed-precision quantization method and device for large model parameters, which can effectively solve the problem in the prior art of being unable to balance model accuracy and inference efficiency during the quantization of large models.

[0005] In a general aspect, a large model parameter mixed-precision quantization method is provided, comprising: obtaining the significance of each weight parameter of a target model; performing 8-bit quantization on the weight parameters whose significance meets a high significance condition to obtain a first quantization result; performing 2-bit quantization on the weight parameters whose significance does not meet the high significance condition to obtain a second quantization result; obtaining a quantized weight matrix of the target model based on the first quantization result and the second quantization result; and processing input data of the target model using the quantized weight matrix and an identification matrix, wherein the identification matrix indicates the quantization type of each weight parameter in the quantized weight matrix.

[0006] Optionally, the weight parameters whose significance meets the high significance condition are determined in the following manner: the weight parameters of each column in the weight matrix of the target model are grouped to obtain multiple weight parameter groups, wherein the number of weight parameters contained in each weight parameter group is equal to the number of processes that the processor can process in parallel; based on the significance of the weight parameters contained in each weight parameter group, the significance of each weight parameter group is determined; for the weight parameters in the weight parameter groups ranked in the first predetermined number of significance, it is determined that the significance of the weight parameters meets the high significance condition.

[0007] Optionally, the weight parameters whose significance meets the high significance condition are quantized to 8 bits to obtain a first quantization result, including: for each weight parameter group in the first predetermined number of weight parameter groups ranked in significance, performing the following processing: dividing the weight parameters in the current weight parameter group into a second predetermined number of small groups, wherein each small group contains 4 weight parameters; obtaining the average value of the weight parameters in each small group; performing 8-bit quantization on the average value of each small group respectively; determining the quantization result of the current weight parameter group based on the quantization result of each small group; in response to all weight parameter groups having completed the above processing, determining the quantization results of all weight parameter groups as the first quantization result.

[0008] Optionally, weight parameters whose significance does not meet the high significance condition are quantized by 2 bits to obtain a second quantization result, including: dividing the weight parameters in all weight parameter groups whose significance is not ranked in the first predetermined number into multiple blocks; for each block in the multiple blocks, performing the following processing: for each column in the current block, quantizing the weight parameters of the current column by 2 bits; determining the quantization error of the current column based on the weight parameters after quantization and the weight parameters before quantization of the current column; updating the unquantized weight parameters in the current block based on the quantization error of the current column; in response to the completion of quantization of all columns of the current block, updating the unquantized blocks in the multiple blocks based on the quantization error of the current block, wherein the quantization error of the current block is determined based on the quantization error of each column in the current block; in response to the completion of the above processing for multiple blocks, determining the quantization result of the weight parameter group as the second quantization result.

[0009] Optionally, based on the quantization results of each group, the quantization results of the current weight parameter group are determined, including: splicing the quantization results of each group into a second predetermined number of bytes for storage, and using the spliced quantization results as the quantization results of the current weight parameter group; wherein, before determining the quantization results of the weight parameter group as the second quantization results, the mixed precision quantization method also includes: for each weight parameter group, splicing the quantized weight parameters in the weight parameter group into a second predetermined number of bytes for storage, and using the spliced weight parameters as the quantization results of the weight parameter group.

[0010] Optionally, the input data of the target model is processed using the quantized weight matrix and identification matrix, including: grouping the data of each row of the input data, and quantizing each group of data to 8 bits, wherein the number of each group of data is equal to the number of processes that the processor can process in parallel; for each group of data, performing the following processing: taking out the weight parameters of the second predetermined number of bytes from the column corresponding to the current group of data in the quantized weight matrix; in response to the identification matrix indicating that the quantization type of the weight parameters of the second predetermined number of bytes is 2-bit quantization, converting the weight parameters of the second predetermined number of bytes into 8-bit weight parameters respectively; in response to the identification matrix indicating that the quantization type of the weight parameters of the second predetermined number of bytes is 8-bit quantization, copying the weight parameter of each byte in the second predetermined number of bytes 4 times to obtain an 8-bit weight parameter; and performing multiplication calculations with the current group of data using the 8-bit weight parameter.

[0011] In another general aspect, a large model parameter mixed precision quantization device is provided, comprising: a first acquisition unit, configured to acquire the significance of each weight parameter of a target model; a second acquisition unit, configured to perform 8-bit quantization on the weight parameters whose significance meets a high significance condition, to obtain a first quantization result; a third acquisition unit, configured to perform 2-bit quantization on the weight parameters whose significance does not meet the high significance condition, to obtain a second quantization result; a fourth acquisition unit, configured to obtain a quantized weight matrix of the target model based on the first quantization result and the second quantization result; and a processing unit, configured to process input data of the target model using the quantized weight matrix and an identification matrix, wherein the identification matrix indicates the quantization type of each weight parameter in the quantized weight matrix.

[0012] Optionally, the weight parameters whose significance meets the high significance condition are determined in the following manner: the weight parameters of each column in the weight matrix of the target model are grouped to obtain multiple weight parameter groups, wherein the number of weight parameters contained in each weight parameter group is equal to the number of processes that the processor can process in parallel; based on the significance of the weight parameters contained in each weight parameter group, the significance of each weight parameter group is determined; for the weight parameters in the weight parameter groups ranked in the first predetermined number of significance, it is determined that the significance of the weight parameters meets the high significance condition.

[0013] Optionally, the second acquisition unit is further configured to perform the following processing for each weight parameter group in the first predetermined number of weight parameter groups ranked in the top in significance: divide the weight parameters in the current weight parameter group into a second predetermined number of small groups, wherein each small group contains 4 weight parameters; obtain the average value of the weight parameters in each small group; perform 8-bit quantization on the average value of each small group; determine the quantization result of the current weight parameter group based on the quantization result of each small group; in response to all weight parameter groups having completed the above processing, determine the quantization results of all weight parameter groups as the first quantization results.

[0014] Optionally, the third acquisition unit is further configured to divide the weight parameters in all weight parameter groups that are not ranked in the first predetermined number in terms of significance into multiple blocks; for each block in the multiple blocks, perform the following processing: for each column in the current block, perform 2-bit quantization on the weight parameters of the current column; determine the quantization error of the current column based on the weight parameters after quantization and the weight parameters before quantization of the current column; update the unquantized weight parameters in the current block based on the quantization error of the current column; in response to the completion of quantization of all columns of the current block, update the unquantized blocks in the multiple blocks based on the quantization error of the current block, wherein the quantization error of the current block is determined based on the quantization error of each column in the current block; in response to the completion of the above processing for multiple blocks, determine the quantization result of the weight parameter group as the second quantization result.

[0015] Optionally, the second acquisition unit is further configured to splice the quantization results of each group into a second predetermined number of bytes for storage, and use the spliced quantization results as the quantization results of the current weight parameter group; wherein, the third acquisition unit is further configured to, for each weight parameter group, before determining the quantization result of the weight parameter group as the second quantization result, splice the quantized weight parameters in the weight parameter group into a second predetermined number of bytes for storage, and use the spliced weight parameters as the quantization result of the weight parameter group.

[0016] Optionally, the processing unit is further configured to group the data of each row of the input data and perform 8-bit quantization on each group of data, wherein the number of each group of data is equal to the number of processes that the processor can process in parallel; for each group of data, perform the following processing: take out the weight parameters of the second predetermined number of bytes from the column corresponding to the current group of data in the quantized weight matrix; in response to the identification matrix indicating that the quantization type of the weight parameters of the second predetermined number of bytes is 2-bit quantization, convert the weight parameters of the second predetermined number of bytes into 8-bit weight parameters respectively; in response to the identification matrix indicating that the quantization type of the weight parameters of the second predetermined number of bytes is 8-bit quantization, copy the weight parameter of each byte in the second predetermined number of bytes 4 times to obtain an 8-bit weight parameter; and perform multiplication calculation with the current group of data using the 8-bit weight parameter.

[0017] In another general aspect, a computer-readable storage medium storing instructions is provided, wherein, when the instructions are executed by at least one computing device, the at least one computing device is prompted to perform any large model parameter mixed precision quantization method as described above.

[0018] In another general aspect, a system is provided comprising at least one computing device and at least one storage device storing instructions, wherein the instructions, when executed by the at least one computing device, cause the at least one computing device to perform any of the large model parameter mixed precision quantization methods described above.

[0019] In another general aspect, a computer program product is provided, comprising computer instructions, which, when executed by a processor, implement any of the above large model parameter mixed precision quantization methods.

[0020] According to the large model parameter mixed precision quantization method and device of the embodiment of the present disclosure, through significance analysis, the weight parameters with higher significance are quantized to 8 bits to ensure the accuracy of the key parts of the model, and the weight parameters with lower significance are quantized to 2 bits to ensure that less storage is occupied. By combining the two, the reasoning efficiency of the model can be improved on the basis of ensuring the accuracy of the key parts of the model and occupying less storage; moreover, the present disclosure also records the quantization type of each weight parameter in the quantized weight matrix through an identification matrix, so that even if data of different quantization bit widths are mixed and stored in the quantized weight matrix, random access and matrix multiplication calculations can be performed efficiently. Therefore, through the present disclosure, the problem that the model accuracy and reasoning efficiency cannot be taken into account in the large model quantization process in the prior art can be effectively solved.

[0021] Additional aspects and / or advantages of the present general inventive concept will be set forth in part in the following description and in part will be apparent from the description, or may be learned through practice of the present general inventive concept. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The above and other objects and features of the embodiments of the present disclosure will become more apparent through the following description in conjunction with the accompanying drawings showing the embodiments, in which: Figure 1 is a flowchart illustrating a mixed-precision quantization method for large model parameters according to an embodiment of the present disclosure; Figure 2 is a schematic diagram illustrating the grouping of weight parameter groups according to an embodiment of the present disclosure; Figure 3 is a flowchart illustrating an efficient large-model mixed-precision quantization method according to an embodiment of the present disclosure; Figure 4 2 is a block diagram illustrating a large model parameter mixed precision quantization apparatus according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0023] The following detailed description is provided to help the reader gain a comprehensive understanding of the methods, devices and / or systems described herein. However, various changes, modifications and equivalents of the methods, devices and / or systems described herein will be clear after understanding the disclosure of the present application. For example, the order of operations described herein is merely an example and is not limited to those orders set forth herein, but can be changed as will be clear after understanding the disclosure of the present application, except for operations that must occur in a specific order. In addition, for greater clarity and conciseness, descriptions of features known in the art may be omitted.

[0024] The features described herein can be implemented in different forms and should not be construed as limited to the examples described herein. Rather, the examples described herein are provided to illustrate only some of the many possible ways to implement the methods, devices, and / or systems described herein, which will become clear after understanding the disclosure of this application.

[0025] As used herein, the term "and / or" includes any one of the associated listed items and any combination of any two or more.

[0026] Although terms such as "first," "second," and "third" may be used herein to describe various members, components, regions, layers, or portions, these members, components, regions, layers, or portions should not be limited by these terms. Instead, these terms are used solely to distinguish one member, component, region, layer, or portion from another member, component, region, layer, or portion. Thus, what is referred to as a first member, first component, first region, first layer, or first portion in the examples described herein may also be referred to as a second member, second component, second region, second layer, or second portion without departing from the teachings of the examples.

[0027] In the specification, when an element (such as a layer, region, or substrate) is described as being “on,” “connected to,” or “coupled to” another element, the element may be directly “on,” “connected to,” or “coupled to” the other element, or one or more other elements may be present therebetween. Conversely, when an element is described as being “directly on,” “directly connected to,” or “directly coupled to” another element, there may be no other elements present therebetween.

[0028] The terms used herein are only used to describe various examples and are not intended to limit the disclosure. Unless the context clearly indicates otherwise, the singular is intended to include the plural. The terms "comprise," "include," and "have" indicate the presence of the recited features, quantities, operations, components, elements, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, quantities, operations, components, elements, and / or combinations thereof.

[0029] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure pertains after understanding the present disclosure. Unless expressly defined as such herein, terms (such as those defined in commonly used dictionaries) should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and should not be interpreted in an idealized or overly formal manner.

[0030] Furthermore, in describing the examples, when it is deemed that a detailed description of well-known related structures or functions would cause an obscure interpretation of the present disclosure, such detailed description will be omitted.

[0031] Currently, the quantization process usually uses affine transformation to transform the original tensor Mapped to a lower bit width discrete numeric space as follows: (1) in, is the scaling factor, is zero point, is the rounding function, To limit the function value range to the low-bit representation space, For the low-bit spatial range, the inverse quantization process is: (2) Since quantization compresses continuous values into a finite number of discrete levels, the bit width used in the quantization process directly determines the number of quantization levels and the ultimate accuracy of the model. For example, increasing the bit width from 4 bits to 5 bits doubles the number of quantization levels compared to 4-bit quantization, significantly reducing the loss of accuracy.

[0032] Traditional quantization methods primarily include per-channel (or per-token) quantization and per-group quantization. While per-channel quantization is simple and efficient, it is prone to large errors when faced with uneven data distribution. Group quantization, on the other hand, calculates scaling factors and zero points on a smaller subset of parameters, better adapting to local data characteristics and thus reducing quantization errors. However, this also increases the complexity of hardware implementation and computational kernel design. Furthermore, to simplify computation, some work employs symmetric quantization (i.e., the zero point is fixed to 0). While this approach can reduce computational complexity, in practice, it often results in significant accuracy loss due to the asymmetric data distribution.

[0033] Currently, the industry has proposed various quantization schemes, such as quantizing only the model's weight parameters and jointly quantizing the model's weight parameters and input data. The former has achieved some success in reducing model storage requirements, but it still requires additional dequantization operations during inference, failing to fundamentally reduce the amount of computation. The latter attempts to utilize low-bit computing units to directly perform matrix multiplication, thereby further improving computational efficiency. However, the lower quantization complexity of the model's input data makes it more difficult to ensure overall accuracy.

[0034] In order to solve the above problems, the present disclosure provides a new and efficient path, that is, through significance analysis, the weight parameters with higher significance are quantized to 8 bits to ensure the accuracy of the key parts of the model, and the weight parameters with lower significance are quantized to 2 bits to ensure that less storage is occupied. By combining the two, the reasoning efficiency of the model can be improved while ensuring the accuracy of the key parts of the model and occupying less storage; moreover, the present disclosure also records the quantization type of each weight parameter in the quantized weight matrix through an identification matrix, so that even if data of different quantization bit widths are mixed and stored in the quantized weight matrix, random access and matrix multiplication calculations can be performed efficiently.

[0035] The large model parameter mixed precision quantization method and device disclosed in the present invention are described in detail below with reference to the accompanying drawings.

[0036] This paper proposes a mixed precision quantization method for large model parameters. Figure 1 1 is a flow chart illustrating a large model parameter mixed precision quantization method according to an embodiment of the present disclosure. Figure 1 , the large model parameter mixed precision quantization method comprises the following steps: In step S101 , the significance of each weight parameter of the target model is obtained.

[0037] Specifically, the weight parameters of the model can be divided into significant weights and non-significant weights, among which the significant weights contribute the most to the model accuracy. Therefore, by calculating the significance of each weight parameter, the key weight parameters of the model can be accurately identified.

[0038] As an example, the significance of each weight parameter can be obtained as follows: (3) in, represents the weight parameter of the i-th row and j-th column in the weight matrix, is the Hessian matrix obtained based on the model input data. The Hessian matrix is obtained as follows: (4) in, represents the input data of the model, express The transposed matrix of represents the identity matrix, Represents the hyperparameters of the model.

[0039] It should be noted that the target model disclosed in the present invention can be used for image recognition, image generation, image enhancement, image retrieval, image detection, etc., and can also be used for speech recognition, speech enhancement, etc. In addition, the target model can also be applied to visual question answering (VQA), text generation, human-computer dialogue, etc., which is not limited by the present disclosure.

[0040] As an example, when the target model is used in scenarios of image recognition, image generation, image enhancement, image retrieval, and image detection, the input data of the target model is image data. In the image recognition scenario, the target model can identify defect information, target objects, etc. in the image based on the input image data; in the image generation scenario, the target model generates the required target image based on the input image data; in the image enhancement scenario, the target model enhances the quality of the image data based on the input image data to obtain quality-enhanced image data; in the image detection scenario, the target model detects information such as the position and category of the target object in the image data based on the input image data.

[0041] As an example, when the target model is used in speech recognition and speech enhancement scenarios, the input data of the target model is speech data. In the speech recognition scenario, the target model can identify defects in the speech based on the input speech data (such as helping students improve their oral English skills); in the speech enhancement scenario, the target model can enhance the speech quality based on the input speech data and obtain speech data with enhanced quality.

[0042] For example, when the target model is applied to a visual question answering (VQA) task, the input data of the target model can be image data and corresponding text questions. The target model can automatically generate accurate answers to the corresponding text questions based on the input image data and related text questions. As an example, when the target model is applied to a text generation task, the input data of the target model is text data, and the target model can generate coherent and content-rich text information based on the given text prompts; As an example, when the target model is applied to a human-computer dialogue task, the input data of the target model is text data or voice data. The target model can achieve a smooth and natural interactive dialogue with the user based on the text or voice input by the user, thereby improving the user experience.

[0043] In step S102, the weight parameters whose significance meets the high significance condition are quantized to 8 bits to obtain a first quantization result.

[0044] As an example, the high significance condition may be that the significance of the weight parameter is greater than a first preset value, or that the weight parameters are grouped and the significance of the group to which the weight parameters belong is greater than a second preset value, which is not limited in this disclosure.

[0045] The following is an introduction using the example where the high significance condition is that the significance of the group where the weight parameter belongs is greater than the second preset value.

[0046] According to an embodiment of the present disclosure, the weight parameters whose significance meets the high significance condition can be determined in the following manner: the weight parameters of each column in the weight matrix of the target model are grouped to obtain multiple weight parameter groups, wherein the number of weight parameters contained in each weight parameter group is equal to the number of processes that the processor can process in parallel; based on the significance of the weight parameters contained in each weight parameter group, the significance of each weight parameter group is determined; for the weight parameters in the weight parameter groups ranked in the first predetermined number of significance, it is determined that the significance of the weight parameters meets the high significance condition.

[0047] Through this embodiment, the weight parameters of each column in the weight matrix of the target model are grouped, and the weight parameters whose significance meets the high significance condition are determined by the significance of each group, so that the key weight parameters of the model can be accurately identified. Moreover, this embodiment groups according to the parallel processing capability of the processor and also realizes memory alignment, which is based on the parallel processing capability alignment of the processing, thereby facilitating the storage of the weight parameters after quantization processing and the subsequent calculation of the model.

[0048] As an example, assuming that the parallel processing capability of the processor is determined by the Single Instruction Multiple Data (SIMD) unit, the weight matrix of the target model can be , where T represents the number of rows of the matrix and P represents the number of columns of the matrix, which are divided in a way that aligns with the SIMD unit size.

[0049] As an example, assuming that the number of processes processed in parallel by the SIMD unit is 16, each weight parameter group contains 16 weight parameters after division. Figure 2 The figure shows the weight parameter grouping method. The dotted line in the figure indicates the block operation, the square filled with diagonal lines is the significance weight parameter, and the square filled with vertical lines has an indicator value of 1, indicating that the weight parameter group at the corresponding position is the significance weight parameter group. Figure 2 As shown, you can Divide into 16 rows each, and you can get piece( ),in, , Each column in is a weight parameter group, and then the significance of each weight parameter group can be calculated : (5) in, i Indicates the i A weight parameter group, j Indicates the i The first j The position of the element.

[0050] Then, sort each weight parameter by the significance of the group, select the top K weight parameter groups, and mark the corresponding weight parameters as significant weights on the bitmap (i.e., the identification matrix mentioned above). The bitmap is a Boolean matrix used to mark whether the weight parameters of each weight parameter group are significant weights.

[0051] According to an embodiment of the present disclosure, performing 8-bit quantization on the weight parameters whose significance meets the high significance condition to obtain a first quantization result may include: for each weight parameter group in the first predetermined number of weight parameter groups ranked in the top in significance, performing the following processing: dividing the weight parameters in the current weight parameter group into a second predetermined number of small groups, wherein each small group contains 4 weight parameters; obtaining the average value of the weight parameters in each small group; performing 8-bit quantization on the average value of each small group respectively; determining the quantization result of the current weight parameter group based on the quantization result of each small group; and in response to all weight parameter groups having performed the above processing, determining the quantization results of all weight parameter groups as the first quantization result.

[0052] Through this embodiment, the weight parameters of the weight parameter group are further grouped and the average value is calculated before quantization, which reduces the quantization complexity. Moreover, further grouping in units of 4 weight parameters can support subsequent mixed packaging with non-significant weights, that is, obtaining a complete quantized weight matrix.

[0053] As an example, for each of the top K weight parameter groups, the weight parameter group is divided again with 4 weight parameters as a group, that is, the weight parameters in the weight parameter group are divided into a first predetermined number of groups, and the average of the 4 weight parameters in each group is calculated, and then each average value is quantized to 8 bits to obtain a first predetermined number of 8-bit data.

[0054] As an example, taking a weight parameter group containing 16 weight parameters, with 4 weight parameters as a group, the weight parameter group can be divided into 4 groups, and the average of the 4 weight parameters in each group is calculated to obtain 4 average values, and these 4 average values are quantized to 8 bits respectively to obtain 4 8-bit data.

[0055] In step S103, the weight parameters whose significance does not meet the high significance condition are quantized to 2 bits to obtain a second quantization result.

[0056] The following is an example of an example in which the significance of the group where the weight parameter has a high significance condition is greater than the second preset value.

[0057] According to an embodiment of the present disclosure, performing 2-bit quantization on weight parameters whose significance does not meet the high significance condition to obtain a second quantization result can include: dividing the weight parameters in all weight parameter groups whose significance is not ranked in the first predetermined number into multiple blocks; performing the following processing for each block in the multiple blocks: for each column in the current block, performing 2-bit quantization on the weight parameters of the current column; determining the quantization error of the current column based on the weight parameters after quantization and the weight parameters before quantization of the current column; updating the unquantized weight parameters in the current block based on the quantization error of the current column; in response to completing quantization of all columns of the current block, updating the unquantized blocks in the multiple blocks based on the quantization error of the current block, wherein the quantization error of the current block is determined based on the quantization error of each column in the current block; in response to completing the above processing for all multiple blocks, determining the quantization result of the weight parameter group as the second quantization result.

[0058] In this embodiment, considering that the 2-bit quantization accuracy is much lower than 8-bit quantization, it may cause a large loss of model accuracy. For this reason, this embodiment introduces a block quantization strategy, that is, using the quantization error of the quantized column to update the weight parameter of the unquantized column, compensating for the accuracy loss of the quantized column weight parameter, so as to make up for the accuracy loss caused by low bits.

[0059] As an example, for the remaining non-significant weights, this disclosure uses 2-bit quantization. However, since 2-bit quantization accuracy is much lower than 8-bit quantization accuracy, it may result in a significant loss of model accuracy. Therefore, this embodiment introduces a block quantization strategy, which uses unquantized weights to compensate for the accuracy loss of quantized weights.

[0060] Specifically, the above weight matrix can be divided into piece( ) in each block For example, in determining After the significance weight in The remaining non-significant weights in can be written as Y .

[0061] by For example, for The block quantization strategy can include the following operations: First, Y Divide into multiple blocks evenly, such as Figure 2 As shown, each block contains B weight parameters; Secondly, the Hessian matrix Perform Cholesky decomposition and get Inverse matrix information : (6) Again, 2-bit quantization is performed for each column weight parameter in each Block: (7) in, is a 2-bit quantization function, j Indicates the first j List, Indicates the first j Column weight parameter.

[0062] Once again, after quantizing the weight parameters of each column, calculate the quantization error of each column: (8) in, Indicates the i The first j Quantization error of the column.

[0063] Again, use i The other unquantized weight parameters in the block are used to compensate for the error, that is, the quantization error is used to update the other unquantized weight parameters: (9) in, Indicates the i The weight parameters that have not yet been quantized in a block.

[0064] Finally, when the iAfter all columns in a block have completed 2-bit quantization, you can use All remaining weight parameters in the Block that have not yet started quantization are i The quantization error of each Block is compensated, that is, the weight parameters of other unquantized Blocks are updated using the quantization error: (10) in, express All remaining weight parameters in the Block that have not yet started quantization, For the i The quantization error of the Block is given by i The quantization errors listed in each Block are spliced together to obtain the result.

[0065] It should be noted that here As an example, the non-significant weights are quantified, and the remaining indivual Repeat the above operation to complete the 2-bit quantization operation.

[0066] In step S104 , a quantized weight matrix of the target model is obtained based on the first quantization result and the second quantization result.

[0067] As an example, in order to package the 8-bit quantization result and the 2-bit quantization result into the same weight matrix, that is, the first quantization result and the second quantization result into the same weight matrix, the quantization results of each weight parameter group in the first quantization result and the second quantization result can be spliced into the same number of bytes for storage.

[0068] According to an embodiment of the present disclosure, determining the quantization result of the current weight parameter group based on the quantization result of each group may include: splicing the quantization result of each group into a second predetermined number of bytes for storage, and using the spliced quantization result as the quantization result of the current weight parameter group; wherein, before determining the quantization result of the weight parameter group as the second quantization result, the mixed precision quantization method also includes: for each weight parameter group, splicing the quantized weight parameters in the weight parameter group into a second predetermined number of bytes for storage, and using the spliced weight parameters as the quantization result of the weight parameter group.

[0069] Through this embodiment, the quantization results of each weight parameter group are spliced into the same number of bytes for storage, so that the quantized high-bit width and low-bit width data can be stored in the same weight matrix, which is convenient for subsequent use of the quantized weight matrix.

[0070] For example, when storing the quantized weight matrix, for low-bit-width quantized weight parameters (2 bits), the multiple 2-bit data contained in a weight parameter group are concatenated into a second predetermined number of bytes for storage; while for high-bit-width weight parameters (8 bits), the second predetermined number of 8-bit data contained in a weight parameter group are concatenated into a second predetermined number of bytes for storage, where each 8-bit data actually represents the average of four int8 values. The essential purpose of this design is to be able to store high-bit-width and low-bit-width data in the same weight matrix.

[0071] As an example, still taking a weight parameter group containing 16 weight parameters as an example, when storing the quantized weight matrix, for low-bit-width quantized weight parameters (2 bits), 16 2-bit data can be spliced into 4 bytes for storage; and for high-bit-width weight parameters (8 bits), 4 8-bit data can be spliced into 4 bytes for storage.

[0072] For example, for a weight matrix W of original size T×P, each time the weight parameters are read, the 16 weight parameters in a column are grouped together for calculation using the SIMD 128-bit registers. Therefore, each column requires accessing T / 16 groups. The packed weight matrix has a shape of [T / 16, P], where each element is a 4-byte INT32. This means that the weight matrix has P columns, each containing T / 16 elements, and each element occupies 4 bytes.

[0073] In addition, a bitmap (the aforementioned identification matrix) can be stored to indicate whether each 32-bit data (i.e., a 4-byte INT32) is packed with 2-bit data or 8-bit data. This bitmap is a binary matrix of size [T / 16, P] and has a storage space of [T / 16 / 8, P] bytes.

[0074] In step S105 , the input data of the target model is processed using the quantized weight matrix and the identification matrix, wherein the identification matrix indicates the quantization type of each weight parameter in the quantized weight matrix.

[0075] As an example, this step identifies the quantization type of each weight parameter in the quantized weight matrix recorded by the matrix, so that even if data of different quantization bit widths are mixed and stored in the quantized weight matrix, random access and matrix multiplication calculations can be performed efficiently.

[0076] According to an embodiment of the present disclosure, optionally, processing the input data of the target model using the quantized weight matrix and identification matrix may include: grouping the data of each row of the input data, and quantizing each group of data to 8 bits, wherein the number of each group of data is equal to the number of processes that the processor can process in parallel; for each group of data, performing the following processing: taking out the weight parameters of a second predetermined number of bytes from the column corresponding to the current group of data in the quantized weight matrix; in response to the identification matrix indicating that the quantization type of the weight parameters of the second predetermined number of bytes is 2-bit quantization, converting the weight parameters of the second predetermined number of bytes into 8-bit weight parameters respectively; in response to the identification matrix indicating that the quantization type of the weight parameters of the second predetermined number of bytes is 8-bit quantization, copying the weight parameters of each byte in the second predetermined number of bytes 4 times to obtain an 8-bit weight parameter; and performing multiplication calculations with the current group of data using the 8-bit weight parameter.

[0077] Through this embodiment, the input data and the 2-bit quantized weight parameters are converted into 8-bit data, so that matrix multiplication can be calculated based on 8-bit data, which improves the reasoning efficiency of the model while ensuring the accuracy of the model.

[0078] As an example, still taking a weight parameter group containing 16 weight parameters as an example, first input data Each row of is grouped into 16 data sets, and each set of data is quantized to 8 bits. In each processing, 16 int8 data can be taken from a row of quantized input data, and at the same time, 16 int8 data can be taken from the weight matrix. Extract 4 bytes of data (32-bit data) from the corresponding column and read the bitmap to determine whether the 32-bit data uses 2-bit or 8-bit quantization. If the 32-bit data uses 2-bit quantization, it can be decoded into 16 int8 values through shifting and bitwise AND operations. If the 32-bit data uses 8-bit quantization, the underlying operations replicate each element of the four 8-bit data four times to form 16 int8 values. Finally, SIMD instructions are used to multiply the 16 int8 values of the input data with the corresponding 16 int8 parameters in the weight matrix.

[0079] In order to facilitate understanding of the above embodiments, Figure 3 Provide a description of the system.

[0080] Figure 3 An efficient mixed-precision quantization method for large models is presented. The method includes the following steps: Step S301: Divide the weight parameters based on the weight parameter group significance analysis. Specifically, first divide the parameters in the model's weight matrix W into groups aligned with the SIMD unit size. For example, each weight parameter group contains 16 values. Then, calculate the significance of each weight parameter group. Select the top K weight parameter groups and determine them as the weight parameter groups with high significance.

[0081] Step S302: Processing weight parameters with high significance. The weight parameter groups with high significance have been determined according to step S301. Next, for each weight parameter group with high significance, the average value of every four weight parameters in the weight parameter group can be calculated, and each average value is then quantized to 8 bits.

[0082] Step S303: Block-based low-bit quantization is performed on the non-significant weight parameters. The weight parameter group with higher significance has been determined according to step S301. Next, 2-bit quantization can be performed on the weight parameters in the less significant weight parameter group. Considering that 2-bit quantization may result in loss of model accuracy, a block-based quantization compensation strategy can be introduced to compensate for the resulting loss of accuracy. The specific process has been detailed above and will not be further explained here.

[0083] Step S304: Packing and storing mixed-precision weight parameters. A packing scheme can be used to pack low-bitwidth (2-bit) and high-bitwidth (8-bit) weight parameters into a unified format for easy random access. Specifically, the quantized weight parameters of each weight parameter group can be packaged as a group of 4-byte data, and an additional bitmap is used to record the quantization type of each data block (4-byte data). This allows efficient calculation and access even when data of different bit widths are mixed and stored in the same weight matrix.

[0084] Step S305: Implementation of weight matrix multiplication for packed mixed-precision weight parameters. First, each row of the input data X is grouped into 16 data elements, and these 16 elements are quantized to 8 bits. Next, 4 bytes of data are extracted from the corresponding columns of the quantized weight matrix W. These 4 bytes are decoded into 16 int8 weight parameters according to the data packing method (indicated by the bitmap). Finally, SIMD instructions are used to efficiently perform the matrix multiplication calculation between the 16 8-bit elements of the input data X and the 16 8-bit weight parameters in the weight matrix.

[0085] In summary, this disclosure addresses the conflict between accuracy and efficiency in the quantization process of large models. Through significance analysis, the weight parameters of large models are divided into two categories: highly significant weight parameters are quantized to 8 bits to ensure accuracy in key parts of the model; less significant weight parameters are quantized to 2 bits, while a quantization compensation strategy is introduced to compensate for the precision loss caused by low bit counts. Furthermore, this disclosure designs a unified packaging and storage scheme. By appending a bitmap to record the quantization type of different data blocks (i.e., the aforementioned concatenated 4-byte data), this scheme enables efficient random access and SIMD-based matrix multiplication even when mixed data of different bit widths is stored in the same weight matrix. In summary, how to effectively reduce computational and storage overhead through mixed-precision quantization while ensuring high accuracy for large models is a key technical challenge in the efficient deployment of large models. The disclosed method provides a new and efficient path to address this issue. Specifically, by assigning different bit widths to different parts of the model weight matrix, a better balance between computational efficiency and model accuracy can be achieved.

[0086] Figure 4 is a block diagram illustrating a large model parameter mixed precision quantization device according to an embodiment of the present disclosure, such as Figure 4 As shown, the device includes a first acquisition unit 40 , a second acquisition unit 42 , a third acquisition unit 44 , a fourth acquisition unit 46 and a processing unit 48 .

[0087] The first acquisition unit 40 is configured to obtain the significance of each weight parameter of the target model; the second acquisition unit 42 is configured to perform 8-bit quantization on the weight parameters whose significance meets the high significance condition to obtain a first quantization result; the third acquisition unit 44 is configured to perform 2-bit quantization on the weight parameters whose significance does not meet the high significance condition to obtain a second quantization result; the fourth acquisition unit 46 is configured to obtain a quantized weight matrix of the target model based on the first quantization result and the second quantization result; the processing unit 48 is configured to use the quantized weight matrix and the identification matrix to process the input data of the target model, wherein the identification matrix indicates the quantization type of each weight parameter in the quantized weight matrix.

[0088] According to an embodiment of the present disclosure, the weight parameters whose significance meets the high significance condition are determined in the following manner: the weight parameters of each column in the weight matrix of the target model are grouped to obtain multiple weight parameter groups, wherein the number of weight parameters contained in each weight parameter group is equal to the number of processes that the processor can process in parallel; based on the significance of the weight parameters contained in each weight parameter group, the significance of each weight parameter group is determined; for the weight parameters in the weight parameter groups whose significance is ranked first in a predetermined number, it is determined that the significance of the weight parameters meets the high significance condition.

[0089] According to an embodiment of the present disclosure, the second acquisition unit 42 is also configured to perform the following processing for each weight parameter group in the first predetermined number of weight parameter groups ranked in the top in significance: dividing the weight parameters in the current weight parameter group into a second predetermined number of small groups, wherein each small group contains 4 weight parameters; obtaining the average value of the weight parameters in each small group; performing 8-bit quantization on the average value of each small group respectively; determining the quantization result of the current weight parameter group based on the quantization result of each small group; in response to all weight parameter groups having completed the above processing, determining the quantization results of all weight parameter groups as the first quantization results.

[0090] According to an embodiment of the present disclosure, the third acquisition unit 44 is further configured to divide the weight parameters in all weight parameter groups that are not ranked in the first predetermined number in terms of significance into multiple blocks; for each block in the multiple blocks, perform the following processing: for each column in the current block, perform 2-bit quantization on the weight parameters of the current column; based on the weight parameters after quantization and the weight parameters before quantization of the current column, determine the quantization error of the current column; based on the quantization error of the current column, update the unquantized weight parameters in the current block; in response to the completion of quantization of all columns of the current block, update the unquantized blocks in the multiple blocks based on the quantization error of the current block, wherein the quantization error of the current block is determined based on the quantization error of each column in the current block; in response to the completion of the above processing for multiple blocks, determine the quantization result of the weight parameter group as the second quantization result.

[0091] According to an embodiment of the present disclosure, the second acquisition unit 42 is further configured to splice the quantization results of each group into a second predetermined number of bytes for storage, and use the spliced quantization results as the quantization results of the current weight parameter group; wherein, the third acquisition unit 44 is further configured to, for each weight parameter group, before determining the quantization result of the weight parameter group as the second quantization result, splice the quantized weight parameters in the weight parameter group into a second predetermined number of bytes for storage, and use the spliced weight parameters as the quantization result of the weight parameter group.

[0092] According to an embodiment of the present disclosure, the processing unit 48 is further configured to group the data of each row of the input data and perform 8-bit quantization on each group of data, wherein the number of each group of data is equal to the number of processes that the processor can process in parallel; for each group of data, the following processing is performed: the weight parameters of the second predetermined number of bytes are taken out from the column corresponding to the current group of data in the quantized weight matrix; in response to the identification matrix indicating that the quantization type of the weight parameters of the second predetermined number of bytes is 2-bit quantization, the weight parameters of the second predetermined number of bytes are respectively converted into 8-bit weight parameters; in response to the identification matrix indicating that the quantization type of the weight parameters of the second predetermined number of bytes is 8-bit quantization, the weight parameters of each byte in the second predetermined number of bytes are copied 4 times to obtain an 8-bit weight parameter; and the 8-bit weight parameter is used to perform multiplication calculation with the current group of data.

[0093] According to an embodiment of the present disclosure, a computer-readable storage medium storing instructions is provided, wherein, when the instructions are executed by at least one computing device, the at least one computing device is prompted to execute a large model parameter mixed precision quantization method as described in any of the above embodiments.

[0094] According to an embodiment of the present disclosure, a system is provided comprising at least one computing device and at least one storage device storing instructions, wherein when the instructions are executed by the at least one computing device, the at least one computing device is prompted to execute a large model parameter mixed precision quantization method as described in any of the above embodiments.

[0095] According to an embodiment of the present disclosure, a computer program product is provided, comprising computer instructions, which, when executed by a processor, implement any of the above large model parameter mixed precision quantization methods.

[0096] While some embodiments of the present disclosure have been shown and described, it will be appreciated by those skilled in the art that changes may be made to these embodiments without departing from the principles and spirit of the disclosure, the scope of which is defined by the claims and their equivalents.

Claims

1. A mixed precision quantization method for large model parameters, characterized in that: include: Get the significance of each weight parameter of the target model; 8-bit quantization is performed on the weight parameter whose significance meets the high significance condition to obtain a first quantization result; Performing 2-bit quantization on the weight parameters whose significance does not meet the high significance condition to obtain a second quantization result; Obtaining a quantized weight matrix of the target model based on the first quantization result and the second quantization result; The input data of the target model is processed using the quantized weight matrix and the identification matrix, wherein the identification matrix indicates the quantization type of each weight parameter in the quantized weight matrix.

2. The large model parameter mixed precision quantization method according to claim 1, characterized in that: The weight parameters for the significance to meet the high significance condition are determined as follows: Grouping the weight parameters of each column in the weight matrix of the target model to obtain a plurality of weight parameter groups, wherein the number of weight parameters included in each weight parameter group is equal to the number of processes that can be processed in parallel by the processor; Determining the significance of each weight parameter group based on the significance of the weight parameters included in each weight parameter group; For the weight parameters in the weight parameter groups ranked in the top first predetermined number in terms of significance, it is determined that the significance of the weight parameters satisfies a high significance condition.

3. The large model parameter mixed precision quantization method according to claim 2, characterized in that: The step of performing 8-bit quantization on the weight parameter whose significance satisfies the high significance condition to obtain a first quantization result includes: For each weight parameter group in the first predetermined number of weight parameter groups ranked in the top in terms of significance, the following process is performed: Dividing the weight parameters in the current weight parameter group into a second predetermined number of subgroups, wherein each subgroup includes 4 weight parameters; Get the average value of the weight parameter in each group; The mean value of each group was quantified at 8 bits; Determining the quantization result of the current weight parameter group based on the quantization result of each group; In response to all weight parameter groups having completed the above processing, the quantization results of all weight parameter groups are determined as the first quantization results.

4. The large model parameter mixed precision quantization method according to claim 3, characterized in that: The step of performing 2-bit quantization on the weight parameter whose significance does not meet the high significance condition to obtain a second quantization result includes: Dividing the weight parameters in all weight parameter groups whose significance is not ranked in the first predetermined number into a plurality of blocks; For each of the plurality of blocks, perform the following processing: For each column in the current block, the weight parameter of the current column is quantized to 2 bits; Determining a quantization error of the current column based on a weight parameter of the current column after quantization and a weight parameter before quantization; updating unquantized weight parameters in the current block based on the quantization error of the current column; In response to quantization of all columns of the current block being completed, updating unquantized blocks in the plurality of blocks based on a quantization error of the current block, wherein the quantization error of the current block is determined based on the quantization error of each column in the current block; In response to the plurality of blocks having completed the above-mentioned processing, the quantization result of the weight parameter group is determined as the second quantization result.

5. The large model parameter mixed precision quantization method according to claim 4, characterized in that: The step of determining the quantization result of the current weight parameter group based on the quantization result of each group includes: splicing the quantization results of each group into a second predetermined number of bytes for storage, and using the spliced quantization results as the quantization results of the current weight parameter group; Before determining the quantization result of the weight parameter group as the second quantization result, the mixed precision quantization method further includes: For each weight parameter group, the quantized weight parameters in the weight parameter group are spliced into the second predetermined number of bytes for storage, and the spliced weight parameters are used as the quantization result of the weight parameter group.

6. The large model parameter mixed precision quantization method according to claim 4, characterized in that: The processing of the input data of the target model using the quantized weight matrix and the identification matrix includes: Grouping the data of each row of the input data, and performing 8-bit quantization on each group of data, wherein the number of each group of data is equal to the number of processes that the processor can process in parallel; For each set of data, perform the following processing: Extracting weight parameters of a second predetermined number of bytes from the column corresponding to the current group of data in the quantized weight matrix; In response to the identification matrix indicating that the quantization type of the weight parameters of the second predetermined number of bytes is 2-bit quantization, converting the weight parameters of the second predetermined number of bytes into 8-bit weight parameters respectively; In response to the identification matrix indicating that the quantization type of the weight parameters of the second predetermined number of bytes is 8-bit quantization, copying the weight parameter of each byte in the second predetermined number of bytes four times to obtain an 8-bit weight parameter; The 8-bit weight parameter is multiplied by the current group data.

7. A mixed precision quantization device for large model parameters, characterized in that: include: A first acquisition unit is configured to acquire the significance of each weight parameter of the target model; A second acquisition unit is configured to perform 8-bit quantization on the weight parameter whose significance meets the high significance condition to obtain a first quantization result; a third obtaining unit configured to perform 2-bit quantization on the weight parameter whose significance does not meet the high significance condition to obtain a second quantization result; a fourth acquiring unit, configured to obtain a quantized weight matrix of the target model based on the first quantization result and the second quantization result; A processing unit is configured to process the input data of the target model using the quantized weight matrix and the identification matrix, wherein the identification matrix indicates the quantization type of each weight parameter in the quantized weight matrix.

8. A computer-readable storage medium storing instructions, characterized in that: When the instruction is executed by at least one computing device, the at least one computing device is prompted to perform the large model parameter mixed precision quantization method according to any one of claims 1 to 6.

9. A system comprising at least one computing device and at least one storage device storing instructions, characterized in that: When the instructions are executed by the at least one computing device, they cause the at least one computing device to perform the large model parameter mixed precision quantization method according to any one of claims 1 to 6.

10. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the large model parameter mixed precision quantization method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Convolutional neural network low bit width quantization method based on weight distribution

    CN110222821A

  • Hybrid precision weight processing method, apparatus and device, and computer program product

    CN118378005A

  • Neural Network Parameter Quantization Method and Apparatus

    US20250117637A1