Quantization parameter calibration method and device, equipment and storage medium
By performing Landau distribution fitting on KV Cache data and calculating quantitative parameters, the problems of increased memory overhead and accuracy loss of KV Cache are solved, and more efficient quantitative reasoning is achieved.
Patent Information
- Application Number
- CN202510193739.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-03
AI Technical Summary
During large-scale training and inference, the memory overhead occupied by KV Cache increases, resulting in a degradation of inference performance. The traditional linear mapping quantization method is not good in advance quantization, resulting in accuracy loss.
By inputting the calibration set into the target network model for inference, obtaining the key-value cache data of each network layer, drawing a histogram, fitting the probability density function of the Landau distribution, calculating position parameters and width parameters, and calculating quantization parameters based on the width parameters, which are used to quantify the newly generated key-value cache data.
It effectively reduces the accuracy loss in KV Cache quantitative inference, supports a variety of quantization scenarios and scale scenarios, which is more reasonable than traditional maximum quantization.
Smart Images

Figure CN120086152A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and particularly to a quantization parameter scaling method, apparatus, device, and storage medium. Background Art
[0002] KV Cache (Key-Value Cache) is an optimization technology widely used in large model inference. Its core idea is to use cached keys and values to avoid repeated calculations, thereby improving inference efficiency and performance. However, currently, when building a network model based on a DL (Deep Learning) framework, the data types of the model parameters and intermediate running data stored using the KV Cache technology are usually defaulted to float32 or float16 types. This causes the memory overhead occupied by the KV Cache to continuously increase during large-scale training and inference processes as the batch size and sequence length continue to grow, and it may even exceed the model itself. It can be seen that storing keys and values can improve inference performance, but at the same time, it increases the memory requirements for inference.
[0003] In addition, the loading of the KV Cache causes the computing cores to be idle, thereby restricting the speed of large model inference. Moreover, on mobile embedded devices, due to limitations such as memory size and power consumption, it is necessary to compress the KV Cache to reduce the memory consumption and loading time required for the parameters. Currently, quantization is usually used to reduce the total number of bytes occupied by the KV Cache, that is, the size of the KV Cache. For example, the int8 quantization method is used to quantize 32-bit parameters into 8-bit parameters, reducing the memory consumption to 1 / 4 of the original, thereby greatly reducing the memory consumption during model operation and lowering the model deployment cost.
[0004] However, the current int8 quantization method uses a linear mapping method. For example, the minimum and maximum values [min, max] of FP32 values (32-bit floating-point numbers, i.e., single-precision) in a tensor are mapped to [-127, 127], and the intermediate values are mapped according to a linear relationship. This mapping relationship is a symmetric mapping, that is, an unsaturated (No saturation) mapping. The unsaturated mapping method directly uses the maximum value method to calculate the quantization parameters. This is better for real-time quantization effects, but not good for pre-quantization effects. Since the KV Cache is an intermediate product of inference, calculating quantization parameters in real time will have a great impact on inference performance. For example, calculating quantization parameters during inference will cause a large loss of accuracy. Therefore, it is necessary to calculate and save the quantization parameters in advance, and the process of calculating the quantization parameters in advance is also called scaling. Summary of the Invention
[0005] In view of this, the purpose of this application is to provide a quantization parameter scaling method, device, equipment and storage medium, which can effectively reduce the accuracy loss in KV Cache quantization inference. The specific scheme is as follows:
[0006] In the first aspect, this application discloses a quantization parameter scaling method, including:
[0007] Input the calibration set into the target network model for inference, and obtain the key-value cache data generated by each network layer during the inference process to obtain the current key-value cache data;
[0008] Respectively draw histograms based on the current key-value cache data corresponding to each network layer to obtain the layer histograms corresponding to each network layer;
[0009] Use the probability density function based on the Landau distribution to fit the layer histogram to calculate the position parameter and width parameter of the Landau distribution of the corresponding network layer;
[0010] Calculate the quantization parameters corresponding to each network layer in the target network model based on the width parameter, so that the target network model can perform quantization operations on the newly generated key-value cache data based on the position parameter and using the quantization parameters.
[0011] Optionally, respectively drawing histograms based on the current key-value cache data corresponding to each network layer to obtain the layer histograms corresponding to each network layer includes:
[0012] Respectively draw histograms based on the key cache data and value cache data in the current key-value cache data corresponding to each network layer to obtain the key histogram and value histogram corresponding to each network layer.
[0013] Optionally, respectively drawing histograms based on the key cache data and value cache data in the current key-value cache data corresponding to each network layer to obtain the key histogram and value histogram corresponding to each network layer includes:
[0014] Use the first preset partition as the X-axis and the key cache data in the current key-value cache data corresponding to the network layer as the Y-axis to draw a histogram to obtain the key histogram corresponding to the corresponding network layer;
[0015] Use the second preset partition as the X-axis and the value cache data in the current key-value cache data corresponding to the network layer as the Y-axis to draw a histogram to obtain the value histogram corresponding to the corresponding network layer.
[0016] Optionally, after calculating the quantization parameters corresponding to each network layer in the target network model based on the width parameter, it further includes:
[0017] Obtain the feature information and data type of the current key-value cache data to obtain the key-value feature information and key-value data type;
[0018] Obtain the parameters for grouping and quantifying the current key-value cache data to obtain the grouping and quantization parameters, and establish the corresponding relationship between the grouping and quantization parameters and the corresponding quantization parameters to obtain the parameter correspondence;
[0019] Store the quantization parameters, key-value feature information, key-value data type, and parameter correspondence according to the preset file format to obtain the target file;
[0020] Read the information in the target file and construct a quantization function and an inverse quantization function based on the read information to generate a model inference framework;
[0021] Correspondingly, perform a quantization operation on the newly generated key-value cache data based on the position parameter and using the quantization parameter, including:
[0022] Load the quantization function in the model inference framework, and perform a quantization operation on the newly generated key-value cache data using the loaded quantization function and the position parameter to obtain the quantization result.
[0023] Optionally, after performing a quantization operation on the newly generated key-value cache data using the loaded quantization function and the position parameter to obtain the quantization result, it further includes:
[0024] Load the inverse quantization function in the model inference framework, and perform an inverse quantization operation on the quantization result using the loaded inverse quantization function to obtain the newly generated key-value cache data.
[0025] Optionally, calculate the quantization parameters corresponding to each network layer in the target network model based on the width parameter, including:
[0026] Calculate the product of the width parameter and the preset proportionality coefficient to obtain the quantization parameters corresponding to each network layer in the target network model; the target network model is an 8-bit signed integer type model based on the attention mechanism.
[0027] Optionally, perform a quantization operation on the newly generated key-value cache data based on the position parameter and using the quantization parameter, including:
[0028] Calculate the product of the quantization parameter and the newly generated key-value cache data, and calculate the sum of the product and the position parameter to obtain the quantization result of the newly generated key-value cache data.
[0029] In a second aspect, the present application discloses a quantization parameter scaling device, including:
[0030] An inference and acquisition module, configured to input a calibration set into a target network model for inference, and obtain key-value cache data generated by each network layer during the inference process to obtain current key-value cache data;
[0031] A drawing module, configured to draw histograms respectively based on the current key-value cache data corresponding to each network layer to obtain layer histograms corresponding to each network layer;
[0032] A fitting module, configured to fit the layer histogram by using a probability density function based on a Landau distribution to calculate the position parameter and width parameter of the Landau distribution of the corresponding network layer;
[0033] A calculation module, configured to calculate quantization parameters corresponding to each network layer in the target network model based on the width parameter, so that the target network model performs a quantization operation on newly generated key-value cache data based on the position parameter and by using the quantization parameters.
[0034] In a third aspect, the present application discloses an electronic device, including a processor and a memory; wherein, when the processor executes a computer program stored in the memory, the quantization parameter scaling method described above is implemented.
[0035] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the quantization parameter scaling method described above is implemented.
[0036] It can be seen that in this application, the calibration set is first input into the target network model for inference, and the key-value cache data generated by each network layer during the inference process is obtained to get the current key-value cache data. Then, histograms are respectively drawn based on the current key-value cache data corresponding to each network layer to obtain the layer histograms corresponding to each network layer, and the probability density function based on the Landau distribution is used to fit the layer histograms to calculate the position parameter and width parameter of the Landau distribution of the corresponding network layer. Then, the quantization parameters corresponding to each network layer in the target network model are calculated based on the width parameter, so that the target network model can perform quantization operations on the newly generated key-value cache data based on the position parameter and using the quantization parameters. This application has pre-analyzed the scale of the key-value cache data generated during model inference and found that its distribution has a long-tail effect, which conforms to the characteristics of the Landau distribution. Therefore, when calibrating the quantization parameters, the calibration of the quantization parameters can be realized based on the probability density function of the Landau distribution. Specifically, first draw a histogram based on the key-value cache data generated by each network layer during model inference, and then use the probability density function based on the Landau distribution to fit the histogram, thereby calculating the position parameter and width parameter of the Landau distribution of the corresponding network layer. Then, the quantization parameters corresponding to each network layer are calculated based on the width parameter, so as to perform quantization operations on the newly generated key-value cache data through the quantization parameters and the quantization parameters. Calibrating the quantization parameters through the above Landau fitting method is more reasonable than the traditional maximum quantization, can effectively reduce the accuracy loss in KV Cache quantization inference, and supports multiple quantization scenarios and calibration scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0038] Figure 1 It is a flowchart of a quantization parameter calibration method disclosed in the present application;
[0039] Figure 2 It is a schematic diagram of the distribution of a specific k-cache data disclosed in the present application;
[0040] Figure 3 It is a schematic diagram of the relationship between a specific k-cache data and head_idx disclosed in the present application;
[0041] Figure 4 It is a flowchart of a specific quantization parameter calibration method disclosed in the present application;
[0042] Figure 5Schematic diagram of a quantization parameter scaling device disclosed in this application;
[0043] Figure 6 Structural diagram of an electronic device disclosed in this application. Specific implementation manner
[0044] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0045] The embodiments of the present application disclose a quantization parameter scaling method. Refer to Figure 1 as shown, the method includes:
[0046] Step S11: Input the calibration set into the target network model for inference, and obtain the key-value cache data generated by each network layer during the inference process to obtain the current key-value cache data.
[0047] In this embodiment, considering that KV Cache is an intermediate product of model inference, calculating quantization parameters during inference will cause a large loss of accuracy and have a great impact on the inference performance of the model. Therefore, it is necessary to estimate the quantization parameters in advance, that is, scale the quantization parameters. Specifically, the calibration set can be first input into the target network model for inference, and at the same time, the key-value cache data generated by each network layer of the target network model during the inference process is obtained to obtain the current key-value cache data (i.e., KV Cache data). Among them, the target network model can specifically be a large language model (Large Language Models, LLM) based on the attention mechanism, such as LLaMA 3 (Large Language Model Meta AI, the third generation version of a series of large-scale pre-trained language models released by Meta).
[0048] Furthermore, the obtained current key-value cache data (i.e., KV Cache data) can also be stored in a preset file format. For example, the key cache data (i.e., K Cache data) and value cache data (i.e., V Cache data) in the current key-value cache data (i.e., KV Cache data) are stored in the pth (a binary file format dedicated to saving models in PyTorch) file format. The code for specifically implementing the saving of KV Cache data is as follows:
[0049] a) torch.save(key_stats, out_dir / 'key_stats.pth')
[0050] b) torch.save(value_stats, out_dir / 'value_stats.pth')。
[0051] Step S12: Draw histograms based on the current key - value cache data corresponding to each network layer respectively, to obtain the layer histograms corresponding to each network layer.
[0052] In this embodiment, after collecting the current key - value cache data generated by each network layer during the model inference process, further, histograms can be drawn respectively based on the current key - value cache data corresponding to each network layer, so as to obtain the layer histograms corresponding to each network layer.
[0053] In this embodiment, drawing histograms respectively based on the current key - value cache data corresponding to each network layer to obtain the layer histograms corresponding to each network layer may include: drawing histograms respectively based on the key cache data and value cache data in the current key - value cache data corresponding to each network layer, to obtain the key histograms and value histograms corresponding to each network layer. It can be understood that the key - value cache data (i.e., KV Cache data) includes two parts of data, namely the key cache data (i.e., K Cache data) and the value cache data (i.e., V Cache data). Therefore, when drawing a histogram based on the key - value cache data, histograms can be drawn respectively based on the key cache data and value cache data corresponding to each network layer, so as to obtain the key histograms and value histograms corresponding to each network layer. By drawing histograms respectively for the key cache data and value cache data in the key - value cache data, the distribution of the key - value cache data can be observed more intuitively, and data support can be provided for the scale of the subsequent quantization parameters, so as to realize the pre - calculation of the quantization parameters.
[0054] Specifically, histograms are respectively drawn based on the key cache data and value cache data in the current key-value cache data corresponding to each network layer, and the key histogram and value histogram corresponding to each network layer are obtained, which may include: taking the first preset partition as the X-axis and taking the key cache data in the current key-value cache data corresponding to the network layer as the Y-axis to draw a histogram, so as to obtain the key histogram corresponding to the corresponding network layer; taking the second preset partition as the X-axis and taking the value cache data in the current key-value cache data corresponding to the network layer as the Y-axis to draw a histogram, so as to obtain the value histogram corresponding to the corresponding network layer. In this embodiment, the first preset partition (bin1) can be taken as the X-axis, and the key cache data (i.e., K Cache data) in the key-value cache data (i.e., KV Cache data) corresponding to the network layer can be taken as the Y-axis to draw a histogram, so as to obtain the key histogram corresponding to the corresponding network layer. Refer to Figure 2 shown Figure 2 Fig. Figure 2 shows the one-dimensional distribution diagram of the key-cache values generated by network layer 0 (i.e., layer 0) during the inference process of the llama3-8b model using the ceval_val_cmcc dataset. By observing the Figure 2 distribution of the key-cache values in Fig. Figure 2 , it can be seen that the distribution of the key-cache values has a long-tail effect, similar to the Landau distribution. Therefore, it is unreasonable to directly use the traditional maximum quantization method to pre-compute the quantization parameters, which will cause inaccurate scaling. Instead, it is more reasonable to calculate the quantization parameters through the density function of the Landau distribution. Similarly, the value cache data (i.e., V Cache data) in the key-value cache data (i.e., KV Cache data) corresponding to the network layer can be taken as the Y-axis to draw a histogram, so as to obtain the value histogram corresponding to the corresponding network layer.
[0055] Specifically, the key histogram of the key-cache value distribution in Fig. Figure 2 can be implemented through the following code: Figure 2 The key histogram of the key-cache value distribution in Fig. Figure 2 :
[0056] a) def plot_per_value(t: np, png_name, quant_group):
[0057] b) t = np.transpose(t, (1, 0, 2, 3))
[0058] c) t = t.reshape(n_layer, -1, kv_head * head_size / / quant_group)
[0059] d) for i in range(t.shape[0]):
[0060] e) print("Ploting %s layer %i " % (png_name, i))
[0061] f) y = t[:, i, :]
[0062] g) y = y.tolist()
[0063] h) plt.subplot(4, 8, i + 1)
[0064] i) sns.histplot(y, bins = 100, legend = False)
[0065] j) format(i,'scaling factor', 'count bin', png_name).
[0066] In addition, before generating the Figure 2 histogram, the distribution of KV Cache data can be further analyzed. For example, call the dataset ceval_val_cmcc to train the llama3 - 8b model, and obtain the KV Cache values generated during the training process. Then, obtain the feature information of the KV Cache values, and analyze the distribution of the KV Cache values based on the feature information. For example, obtain the feature information of key - cache (i.e., key value) and value - cache (i.e., value value) in the KV Cache values, such as the shape information of key - cache (e.g., input_length, KV_head, head_size, etc.), where input_length is the length of the input data, kv_head is the number of attention heads in the llama3 - 8b model, and head_size is the dimension of the feature vector of each attention head. The elements corresponding to head_size can be named head_idx. Specifically, kv_head = 8 and head_size = 128. Then, count the relationship between the KV Cache values generated during the training process and head_idx of the llama3 - 8b model, and obtain a statistical relationship diagram as shown in Figure 3 ; where the x - axis is head_idx and the Y - axis is the key value. By analyzing the Figure 3 statistical relationship diagram, it can be seen that the distribution function is similar to a uniform distribution along the head_size dimension. In addition, from the distribution function, it can be seen that along the head_size dimension, the key values have no obvious characteristics. Therefore, it is feasible to calculate the quantization parameters of a single layer using all head_sizes together.
[0067] Step S13: Fit the layer histogram using the probability density function based on the Landau distribution to calculate the position parameter and width parameter of the Landau distribution of the corresponding network layer.
[0068] In this embodiment, after obtaining the layer histograms corresponding to each network layer, the layer histogram can be fitted using the probability density function based on the Landau distribution, so as to calculate the position parameter and width parameter of the Landau distribution of the corresponding network layer.
[0069] Among them, the calculation formula corresponding to the probability density function p is:
[0070] ;
[0071] where x is the statistical partition bin of the x-axis, is the position parameter, is the width parameter, is the probability density function based on the Landau distribution.
[0072] In this embodiment, in order to ensure the accuracy of the quantization parameters of the final scale, before fitting the layer histogram using the probability density function based on the Landau distribution, it can be determined whether the key histogram and value histogram corresponding to each network layer conform to the Landau distribution. If both the key histogram and value histogram conform to the Landau distribution, then execute the step of fitting the layer histogram using the probability density function based on the Landau distribution. This is because the K Cache data and V Cache data have the same probability distribution characteristics (which can be seen through Figure 2 ). If the K Cache data conforms to the Landau distribution, then the V Cache data should also conform to the Landau distribution. If the distribution characteristics of the two are different, it indicates that the KV Cache data has an abnormality. At this time, reasoning can be performed again, or reasoning can be performed again based on a new calibration set to obtain new KV Cache data until the new key histogram and the new value histogram of the corresponding network layer both conform to the Landau distribution.
[0073] Step S14: Calculate the quantization parameters corresponding to each network layer in the target network model based on the width parameter, so that the target network model can perform quantization operations on the newly generated key-value cache data based on the position parameter and using the quantization parameters.
[0074] In this embodiment, after calculating the position parameter and width parameter of the Landau distribution of each network layer, the quantization parameters (i.e., scaling factor, scale factor) corresponding to each network layer in the target network model can be directly calculated based on the width parameter , so as to facilitate the target network model to be based on this position parameter And perform quantization operations on the newly generated key-value cache data during the real-time inference process using quantization parameters. Among them, the position parameter is equivalent to a bias (such as the zero point position, zero point).
[0075] In a specific embodiment, calculating the quantization parameters corresponding to each network layer in the target network model based on the width parameter may include: calculating the product of the width parameter and a preset proportionality coefficient to obtain the quantization parameters corresponding to each network layer in the target network model; the target network model is a model of 8-bit signed integer type based on the attention mechanism.
[0076] In this embodiment, the target network model is a model of 8-bit signed integer type (i.e., int8) based on the attention mechanism, and can perform int8 quantization on the KV Cache data generated during the inference process. Specifically, based on the width parameter when calculating the quantization parameters of each network layer, the product of the width parameter and the preset proportionality coefficient can be used as the quantization parameter (i.e., the scaling factor value) of the corresponding network layer. The preset proportionality coefficient can be selected according to actual application requirements (such as 2 or 3), and the corresponding quantization parameter is or .
[0077] In this embodiment, performing quantization operations on the newly generated key-value cache data based on the position parameter and using the quantization parameter may specifically include: calculating the product of the quantization parameter and the newly generated key-value cache data, and calculating the sum of the product and the position parameter to obtain the quantization result of the newly generated key-value cache data. That is, when calculating the quantization parameter of the real-time generated key-value cache data, first calculate the product of the quantization parameter (i.e., the scaling factor) and the key-value cache data, and then calculate the sum value of the product and the position parameter and use this sum value as the quantization result; where the position parameter can be 0 by default.
[0078] It can be seen that in the embodiment of the present application, the calibration set is first input into the target network model for inference, and the key-value cache data generated by each network layer during the inference process is obtained to get the current key-value cache data. Then, histograms are respectively drawn based on the current key-value cache data corresponding to each network layer to obtain the layer histograms corresponding to each network layer. The probability density function based on the Landau distribution is used to fit the layer histograms to calculate the position parameter and width parameter of the Landau distribution of the corresponding network layer. Then, the quantization parameters corresponding to each network layer in the target network model are calculated based on the width parameter, so that the target network model can perform quantization operations on the newly generated key-value cache data based on the position parameter and using the quantization parameters. In the embodiment of the present application, a scale analysis is pre-performed on the key-value cache data generated during model inference, and it is found that its distribution has a long-tail effect, which conforms to the characteristics of the Landau distribution. Therefore, when calibrating the quantization parameters, the calibration of the quantization parameters can be realized based on the probability density function of the Landau distribution. Specifically, first, a histogram is drawn based on the key-value cache data generated by each network layer during model inference, and then the probability density function based on the Landau distribution is used to fit the histogram, so as to calculate the position parameter and width parameter of the Landau distribution of the corresponding network layer. Then, the quantization parameters corresponding to each network layer are calculated based on the width parameter, so as to perform quantization operations on the newly generated key-value cache data through the quantization parameters and the quantization parameters. Calibrating the quantization parameters through the above Landau fitting method is more reasonable than the traditional maximum quantization, can effectively reduce the accuracy loss in KV Cache quantization inference, and supports multiple quantization scenarios and calibration scenarios.
[0079] The embodiment of the present application discloses a specific method for calibrating quantization parameters. Refer to Figure 4 as shown, the method includes:
[0080] Step S21: Input the calibration set into the target network model for inference, and obtain the key-value cache data generated by each network layer during the inference process to get the current key-value cache data.
[0081] Step S22: Respectively draw histograms based on the current key-value cache data corresponding to each network layer to obtain the layer histograms corresponding to each network layer.
[0082] Step S23: Use the probability density function based on the Landau distribution to fit the layer histograms to calculate the position parameter and width parameter of the Landau distribution of the corresponding network layer.
[0083] Step S24: Calculate the quantization parameters corresponding to each network layer in the target network model based on the width parameter, and obtain the feature information and data type of the current key-value cache data to get the key-value feature information and key-value data type.
[0084] In this embodiment, the quantization parameter (scaling factor) corresponding to each network layer in the target network model can be calculated based on the product of the width parameter and the preset proportionality coefficient. Then, the feature information (such as attribute information, KV_head, head_size, etc.) and data type of the current key-value cache data (i.e., KV Cache data) are obtained to get the key-value feature information and key-value data type (kv-cacha dtype).
[0085] Step S25: Obtain the parameters for grouping and quantizing the current key-value cache data to get the grouped quantization parameters, and establish the corresponding relationship between the grouped quantization parameters and the corresponding quantization parameters to obtain the parameter correspondence relationship.
[0086] In this embodiment, the parameters for grouping and quantizing the current key-value cache data are obtained to get the grouped quantization parameters (i.e., quant_group), and the corresponding relationship between the grouped quantization parameters (i.e., quant_group) and the corresponding quantization parameters (scaling factor) is established to obtain the parameter correspondence relationship, such as k_scale, v_scale (layer_i: value).
[0087] Step S26: Store the quantization parameters, key-value feature information, key-value data type, and parameter correspondence relationship in a preset file format to obtain the target file.
[0088] In this embodiment, the quantization parameters (scaling factor), key-value feature information, key-value data type (kv-cacha dtype), and parameter correspondence relationship are stored in a preset file format, such as the json format, to obtain a json file, and the json file is saved.
[0089] Step S27: Read the information in the target file and construct a quantization function and an inverse quantization function based on the read information to generate a model inference framework.
[0090] Then, read the saved json file containing the scaling factor and transfer it to the underlying layer so that the underlying layer can construct a quantization function and an inverse quantization function based on the read information, thereby generating a model inference framework. Specifically, the construction of the quantization function and the inverse quantization function can be implemented by the following code:
[0091] i. / / KV CACHE int8
[0092] ii. static inline __device__ float int8_to_float(uint8_t x, const float scale, const float zero_point) {
[0093] iii. int8_t a = x - 128;
[0094] iv. float res = a * scale + zero_point;
[0095] v. return res;
[0096] vi.}
[0097] vii.
[0098] viii. static inline __device__ uint8_t float_to_int8(float x, const float scale, const float zero_point) {
[0099] ix. int8_t fx = roundf(max(-128.f, min(127.f, (x - zero_point) / scale)));
[0100] x. uint8_t res = fx + 128;
[0101] xi. return res;
[0102] xii.}。
[0103] Step S28: Load the quantization function in the model inference framework, and perform a quantization operation on the newly generated key-value cache data by using the loaded quantization function and the position parameters to obtain a quantization result.
[0104] In this embodiment, during the implementation inference of the target network model, the quantization function in the model inference framework can be loaded, and then the loaded quantization function and the position parameters are used to perform a quantization operation on the newly generated key-value cache data to obtain the corresponding quantization result.
[0105] Further, after performing a quantization operation on the newly generated key-value cache data by using the loaded quantization function and position parameters, and obtaining the quantization result, the following steps may further be included: loading the dequantization function in the model inference framework, and performing a dequantization operation on the quantization result by using the loaded dequantization function to obtain the newly generated key-value cache data. In this embodiment, after quantizing the KV Cache data, a dequantization operation may further be performed on the quantization result. Specifically, the dequantization function in the model inference framework may be loaded first, and then the loaded dequantization function may be used to perform a dequantization operation on the quantization result.
[0106] Among them, for a more specific processing procedure of the above steps S21 to S23, reference may be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated herein.
[0107] It can be seen that in the embodiment of the present application, the layer histograms drawn based on the current key-value cache data corresponding to each network layer are fitted by using the probability density function of the Landau distribution, so as to calculate the position parameter and width parameter of the Landau distribution of the corresponding network layer, and calculate the quantization parameters corresponding to each network layer in the target network model based on the width parameter. After that, the feature information and data type of the current key-value cache data are obtained to get the key-value feature information and key-value data type, and the grouping quantization parameters when grouping and quantizing the current key-value cache data are obtained. Then, the corresponding relationship between the grouping quantization parameters and the corresponding quantization parameters is established to obtain the parameter corresponding relationship. Then, the quantization parameters, key-value feature information, key-value data type, and parameter corresponding relationship are stored in accordance with a preset file format to obtain a target file, and the information in the target file is read to construct a quantization function and a dequantization function based on the read information to generate a model inference framework, so as to implement the quantization of real-time key-value cache data through the model inference framework. In the embodiment of the present application, the distribution of the KV Cache data of each network layer is obtained by drawing a histogram, and the scale of the quantization parameters is implemented by using the probability density function of the Landau distribution based on the distribution characteristics of the KV Cache data. During the process of scaling, a model inference framework is constructed based on the corresponding relationship between the quantization parameters, key-value feature information, key-value data type, and parameters, and quantization and dequantization are implemented based on this framework. The present application pre-obtains quantization parameters in a Landau fitting manner based on the distribution characteristics of the KV Cache data itself, has a complete theoretical basis, is more reasonable than the traditional maximum quantization method, can effectively reduce the precision loss in KV Cache quantization inference, and can improve the efficiency of real-time model inference by calculating quantization parameters in advance.
[0108] Correspondingly, the embodiment of the present application also discloses a quantization parameter scaling device, as shown in Figure 5 shown, the device includes:
[0109] The inference and acquisition module 11 is configured to input a calibration set into a target network model for inference, and obtain key-value cache data generated by each network layer during the inference process, so as to obtain the current key-value cache data;
[0110] The drawing module 12 is configured to draw histograms respectively based on the current key-value cache data corresponding to each network layer, so as to obtain layer histograms corresponding to each network layer;
[0111] The fitting module 13 is configured to fit the layer histogram by using a probability density function based on the Landau distribution, so as to calculate the position parameter and width parameter of the Landau distribution of the corresponding network layer;
[0112] The calculation module 14 is configured to calculate quantization parameters corresponding to each network layer in the target network model based on the width parameter, so that the target network model can perform a quantization operation on newly generated key-value cache data based on the position parameter and by using the quantization parameter.
[0113] Among them, for the specific working processes of the above-mentioned various modules, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details are not described herein again.
[0114] It can be seen that in the embodiment of the present application, first, a calibration set is input into a target network model for inference, and key-value cache data generated by each network layer during the inference process is obtained to obtain the current key-value cache data. Then, histograms are respectively drawn based on the current key-value cache data corresponding to each network layer to obtain layer histograms corresponding to each network layer, and the layer histograms are fitted by using a probability density function based on the Landau distribution to calculate the position parameter and width parameter of the Landau distribution of the corresponding network layer. Then, quantization parameters corresponding to each network layer in the target network model are calculated based on the width parameter, so that the target network model can perform a quantization operation on newly generated key-value cache data based on the position parameter and by using the quantization parameter. In the embodiment of the present application, scale analysis is pre-performed on the key-value cache data generated during model inference, and it is found that its distribution has a long-tail effect and conforms to the characteristics of the Landau distribution. Therefore, when calibrating the quantization parameter, the calibration of the quantization parameter can be realized based on the probability density function of the Landau distribution. Specifically, first, a histogram is drawn based on the key-value cache data generated by each network layer during model inference, and then the histogram is fitted by using a probability density function based on the Landau distribution, so as to calculate the position parameter and width parameter of the Landau distribution of the corresponding network layer. Then, quantization parameters corresponding to each network layer are calculated based on the width parameter, so as to perform a quantization operation on newly generated key-value cache data through the quantization parameter and the quantization parameter. Calibrating the quantization parameter through the above-mentioned Landau fitting method is more reasonable than the traditional maximum quantization, can effectively reduce the precision loss in KV Cache quantization inference, and supports multiple quantization scenarios and calibration scenarios.
[0115] In some specific embodiments, the drawing module 12 may specifically include:
[0116] A histogram drawing unit, configured to draw histograms respectively based on the key cache data and value cache data in the current key-value cache data corresponding to each network layer, so as to obtain a key histogram and a value histogram corresponding to each network layer.
[0117] In some specific embodiments, the histogram drawing unit may specifically include:
[0118] A first drawing unit, configured to use a first preset partition as the X-axis and the key cache data in the current key-value cache data corresponding to the network layer as the Y-axis to draw a histogram, so as to obtain a key histogram corresponding to the corresponding network layer;
[0119] A second drawing unit, configured to use a second preset partition as the X-axis and the value cache data in the current key-value cache data corresponding to the network layer as the Y-axis to draw a histogram, so as to obtain a value histogram corresponding to the corresponding network layer.
[0120] In some specific embodiments, after the calculation module 14, it may further include:
[0121] An information acquisition unit, configured to acquire the feature information and data type of the current key-value cache data, so as to obtain key-value feature information and key-value data type;
[0122] A parameter acquisition unit, configured to acquire the parameters for grouping and quantifying the current key-value cache data, so as to obtain grouping quantization parameters;
[0123] A relationship establishment unit, configured to establish a correspondence between the grouping quantization parameters and the corresponding quantization parameters, so as to obtain a parameter correspondence;
[0124] A unit, configured to store the quantization parameters, key-value feature information, key-value data type, and parameter correspondence in accordance with a preset file format, so as to obtain a target file;
[0125] A reading unit, configured to read the information in the target file;
[0126] A construction unit, configured to construct a quantization function and an inverse quantization function based on the read information, so as to generate a model inference framework;
[0127] Correspondingly, the calculation module 14 may specifically include:
[0128] A quantization operation unit, configured to load the quantization function in the model inference framework, and perform a quantization operation on the newly generated key-value cache data by using the loaded quantization function and position parameters, so as to obtain a quantization result.
[0129] In some specific embodiments, after the quantization operation unit, it may further include:
[0130] The dequantization operation unit is used to load the dequantization function in the model inference framework, and perform a dequantization operation on the quantization result by using the loaded dequantization function to obtain newly generated key-value cache data.
[0131] In some specific embodiments, the calculation module 14 may specifically include:
[0132] The first calculation unit is used to calculate the product of the width parameter and the preset proportionality coefficient to obtain the quantization parameters corresponding to each network layer in the target network model; the target network model is an 8-bit signed integer type model based on the attention mechanism.
[0133] In some specific embodiments, the calculation module 14 may specifically include:
[0134] The second calculation unit is used to calculate the product of the quantization parameter and the newly generated key-value cache data, and calculate the sum of the product and the position parameter to obtain the quantization result of the newly generated key-value cache data.
[0135] Furthermore, the embodiment of the present application also discloses an electronic device. Figure 6 It is a structural diagram of the electronic device 20 shown according to an exemplary embodiment, and the content in the figure cannot be considered as any limitation to the scope of use of the present application.
[0136] Figure 6 It is a schematic structural diagram of an electronic device 20 provided by the embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the quantization parameter scaling method disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0137] In this embodiment, the power supply 23 is used to provide a working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.
[0138] In addition, as a carrier for storing resources, the memory 22 can be a read-only memory, a random access memory, a magnetic disk, an optical disc, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0139] Among them, the operating system 221 is used to manage and control each hardware device and the computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the quantization parameter scaling method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 can further include computer programs that can be used to complete other specific tasks.
[0140] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the quantization parameter scaling method disclosed above is implemented. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated herein.
[0141] Furthermore, the embodiments of the present application also disclose a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the quantization parameter scaling method disclosed above are implemented.
[0142] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and reference can be made to the description in the method part for related parts.
[0143] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0144] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented directly in hardware, in software modules executed by a processor, or in a combination thereof. The software modules may be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0145] Finally, it should also be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0146] The above has introduced in detail a quantization parameter scaling method, apparatus, device and storage medium provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A quantitative parameter calibration method, characterized in that: include: Input the calibration set into the target network model for inference, and obtain the key-value cache data generated by each network layer during the inference process to obtain the current key-value cache data; Draw a histogram based on the current key-value cache data corresponding to each network layer to obtain a layer histogram corresponding to each network layer; Fitting the layer histogram using a probability density function based on Landau distribution to calculate the location parameter and width parameter of the Landau distribution of the corresponding network layer; The quantization parameters corresponding to each network layer in the target network model are calculated based on the width parameter, so that the target network model performs a quantization operation on the newly generated key-value cache data based on the position parameter and using the quantization parameter.
2. The quantization parameter calibration method according to claim 1, characterized in that: The drawing of histograms based on the current key-value cache data corresponding to each network layer to obtain a layer histogram corresponding to each network layer includes: Histograms are drawn based on the key cache data and the value cache data in the current key-value cache data corresponding to each network layer, respectively, to obtain a key histogram and a value histogram corresponding to each network layer.
3. The quantization parameter calibration method according to claim 2, characterized in that: The step of drawing a histogram based on the key cache data and the value cache data in the current key-value cache data corresponding to each network layer to obtain a key histogram and a value histogram corresponding to each network layer includes: Using the first preset partition as the X-axis and the key cache data in the current key-value cache data corresponding to the network layer as the Y-axis to draw a histogram, so as to obtain a key histogram corresponding to the corresponding network layer; A histogram is drawn using the second preset partition as the X-axis and the value cache data in the current key-value cache data corresponding to the network layer as the Y-axis to obtain a value histogram corresponding to the corresponding network layer.
4. The quantization parameter calibration method according to claim 1, characterized in that: After calculating the quantization parameters corresponding to each network layer in the target network model based on the width parameter, the method further includes: Acquire characteristic information and data type of the current key-value cache data to obtain key-value characteristic information and key-value data type; Acquire parameters for grouping and quantizing the current key-value cache data to obtain grouping quantization parameters, and establish a corresponding relationship between the grouping quantization parameters and the corresponding quantization parameters to obtain a parameter corresponding relationship; The quantization parameter, the key-value feature information, the key-value data type and the parameter correspondence are stored according to a preset file format to obtain a target file; Reading the information in the target file, and constructing a quantization function and a dequantization function based on the read information to generate a model reasoning framework; Accordingly, the quantizing operation on the newly generated key-value cache data based on the position parameter and using the quantization parameter includes: The quantization function in the model inference framework is loaded, and the newly generated key-value cache data is quantized using the loaded quantization function and the position parameter to obtain a quantization result.
5. The quantization parameter calibration method according to claim 4, characterized in that: After the quantization operation is performed on the newly generated key-value cache data using the loaded quantization function and the position parameter to obtain the quantization result, the method further includes: The dequantization function in the model inference framework is loaded, and the dequantization operation is performed on the quantization result by using the loaded dequantization function to obtain the newly generated key-value cache data.
6. The quantization parameter calibration method according to claim 1, characterized in that: The calculating, based on the width parameter, the quantization parameter corresponding to each network layer in the target network model includes: The product of the width parameter and the preset proportional coefficient is calculated to obtain the quantization parameters corresponding to each network layer in the target network model; the target network model is an 8-bit signed integer type model based on the attention mechanism.
7. The quantization parameter calibration method according to any one of claims 1 to 6, characterized in that: The step of performing a quantization operation on the newly generated key-value cache data based on the position parameter and using the quantization parameter includes: The product of the quantization parameter and the newly generated key-value cache data is calculated, and the sum of the product and the position parameter is calculated to obtain a quantization result of the newly generated key-value cache data.
8. A quantization parameter calibration device, characterized in that: include: The inference and acquisition module is used to input the calibration set into the target network model for inference, and obtain the key-value cache data generated by each network layer during the inference process to obtain the current key-value cache data; A drawing module, used for drawing histograms based on the current key-value cache data corresponding to each network layer, to obtain a layer histogram corresponding to each network layer; A fitting module, used for fitting the layer histogram using a probability density function based on Landau distribution to calculate the location parameter and width parameter of the Landau distribution of the corresponding network layer; A calculation module is used to calculate the quantization parameters corresponding to each network layer in the target network model based on the width parameter, so that the target network model performs a quantization operation on the newly generated key-value cache data based on the position parameter and using the quantization parameter.
9. An electronic device, characterized in that: The method comprises a processor and a memory; wherein, when the processor executes the computer program stored in the memory, the quantization parameter calibration method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: Used to store a computer program; wherein, when the computer program is executed by a processor, the quantization parameter calibration method according to any one of claims 1 to 7 is implemented.