Model quantification method and related device
By dividing the large language model into multiple network structures and performing multiple rounds of quantization processes, the problem of large resource occupancy during the deployment of a large number of parameters is solved, and efficient model quantification and deployment are achieved.
Patent Information
- Application Number
- CN202311626819.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-05-30
AI Technical Summary
The existing technology is difficult to quantify large language models with huge parameters, which makes the model occupy a large amount of memory and computing resources during deployment, making it difficult to deploy effectively.
By dividing the target model into multiple network structures connected in sequence and performing multiple rounds of quantization processes, the number of parameters loaded in a single time is reduced to ensure the smooth execution of the quantization process. Each round of quantization process performs quantification on at least two consecutive network structures in multiple network structures, and partially overlapping network structures are maintained during multiple rounds of quantization to establish joint optimization relationships.
Effective quantification of a model with huge parameters is achieved, the memory and computing resources are occupied, the model deployment efficiency is ensured, and the quantized model accuracy is ensured through joint optimization relationships.
Smart Images

Figure CN120068975A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of Artificial Intelligence (AI), and particularly to a model quantization method and related devices. Background Art
[0002] With the continuous development of AI, neural network models are widely used in different fields, achieving far better results than before. At the same time, the number of parameters of neural network models has become larger and larger, resulting in a large amount of memory resources and computing resources being occupied when neural network models run, severely restricting the deployment of neural network models in various scenarios.
[0003] To solve the problem that it is difficult to deploy neural network models due to the excessive number of parameters, the concept of model quantization has been proposed in related technologies. Model quantization is actually a compression technology for neural network models. Specifically, it uses a lower-bit width to represent the weight parameters and feature data in the neural network model, thereby saving memory space and reducing the amount of computation. Generally, in related technologies, the entire neural network model is loaded onto a processing device, and then the entire neural network model is quantized to ensure the accuracy of the quantized neural network model.
[0004] However, with the rise of large language models, large language models with huge numbers of parameters cannot be fully loaded onto a processing device at one time. Therefore, the model quantization methods in related technologies are difficult to perform quantization on models with huge numbers of parameters such as large language models. Summary of the Invention
[0005] This application provides a model quantization method that can perform quantization on models with huge numbers of parameters and ensure the accuracy of the quantized models.
[0006] The first aspect of this application provides a model quantization method, which is applied to perform quantization on models with relatively large numbers of parameters in the AI field. The method includes: obtaining a target model, where the target model is a neural network model that needs to be quantized, and the target model is a model trained based on the training data in a training set. Generally, the target model is a neural network model with a relatively large number of parameters, and it is difficult for a conventional AI processing device to load the entire target model at one time to perform the quantization process.
[0007] After obtaining the target model, the target model can be divided into multiple network structures connected in sequence. That is, the target model itself includes multiple network structures connected in sequence. By dividing the target model into multiple parts to perform the subsequent quantization process, each part corresponds to a network structure. Among the multiple network structures obtained by division, the output of the previous network structure is the input of the next network structure, that is, there is an input-output dependency relationship between the front and back network structures.
[0008] Finally, a multi-round quantization process is performed to obtain a quantized target model, and the quantized target model includes multiple quantized network structures. Among them, the multi-round quantization process is sequentially performed based on the connection order of the multiple network structures. Moreover, in the multi-round quantization process, each round of quantization process quantizes at least two consecutive network structures among the multiple network structures, and each round of quantization process quantizes a part of the network structures among the multiple network structures. In addition, for any two consecutive quantization processes in the multi-round quantization process, there is at least one repeated network structure between the network structures quantized in the latter round of quantization process and the network structures quantized in the former round of quantization process. It should be noted that for the network structures repeatedly quantized in the two consecutive quantization processes (i.e., the aforementioned at least one repeated network structure), the latter round of quantization process is actually performed based on the former round of quantization process, that is, the quantization parameters used when the latter round of quantization process just starts quantization are the quantization parameters determined in the former round of process.
[0009] In this solution, the target model is divided into multiple parts based on the connection relationship of the network structures in the target model, and a multi-round quantization process is performed on the target model. The multi-round quantization process is sequentially performed based on the network structures of the model, and each round of quantization process quantizes a part of the network structures in the target model, so as to reduce the number of parameters loaded onto the processing device at one time and ensure that the quantization process of the target model can be smoothly performed. Moreover, in the quantization process, there are partially overlapping network structures in any two consecutive quantization processes, fully considering the dependency relationship between the network structures, ensuring that a joint optimization relationship can be established between the network structures in the multi-round quantization process, and effectively avoiding the situation where each round of quantization process independently quantizes different network structures and cannot achieve the global optimum, and guaranteeing the accuracy of the finally quantized target model.
[0010] In a possible implementation manner, the first N network structures quantized in the latter round of quantization process are the same N network structures as the last N network structures quantized in the former round of quantization process, where N is an integer greater than or equal to 1. For example, assume that both the former round of quantization process and the latter round of quantization process quantize M network structures. Then, the last N network structures among the M network structures quantized in the former round of quantization process are actually the same N network structures as the first N network structures among the M network structures quantized in the latter round of quantization process, that is, there are N overlapping network structures in the two consecutive quantization processes.
[0011] In this solution, during the quantization processes of two consecutive rounds, there is a partially overlapping network structure that jointly performs quantization with another part of the network structure in the previous quantization process and also jointly performs quantization with yet another part of the network structure in the next quantization process, thereby establishing a connection relationship between the quantization processes of different rounds and effectively avoiding the situation where each quantization process independently quantifies different network structures and fails to achieve the global optimum.
[0012] In a possible implementation, any quantization process except the first quantization process among multiple rounds of quantization processes is regarded as the target round quantization process. Then, the target round quantization process specifically includes: First, input the first input data into at least two network structures to obtain the first output data. The at least two network structures are the network structures quantized in the target round quantization process, and the first input data is the output of the network structure before the at least two network structures. Since the current at least two network structures are not at the very beginning of the target model, the input of these at least two network structures is actually the feature data (i.e., the first input data) output by the previous network structure.
[0013] Then, input the second input data into the at least two network structures quantized based on the quantization parameters to obtain the second output data. The second input data is the output after quantization of the network structure before the at least two network structures.
[0014] Secondly, based on the difference between the first output data and the second output data, update the quantization parameters of the at least two network structures. For example, use the method of gradient backpropagation to update the quantization parameters of the at least two network structures, so that when the at least two network structures are quantized based on the updated quantization parameters, the output of the at least two network structures can be closer to the output before quantization.
[0015] In this solution, during the process of determining the quantization parameters of the network structure, by using the output of the original network structure as the supervision signal to construct a difference with the output of the network structure quantized based on the quantization parameters, and then updating the quantization parameters with the goal of minimizing the output difference, accurate quantization parameters can be obtained, ensuring the accuracy of the network structure quantized based on the quantization parameters.
[0016] In a possible implementation, the quantization parameters of the weight parameters include the quantization step size and the weight compensation matrix. The quantization step size is used to update the weight parameters in the network structure, and the weight compensation matrix is used to compensate the updated weight parameters.
[0017] In addition, the weight compensation matrix is decomposed into multiple matrices by low-rank decomposition. The total number of parameters of the multiple matrices is less than that of the weight compensation matrix, and the update of the weight compensation matrix is achieved by updating the multiple matrices. Among them, low-rank decomposition is a method of decomposing a high-dimensional matrix into low-rank matrices, which can effectively reduce the dimension of data and reduce redundant information.
[0018] In this solution, by decomposing the weight compensation matrix with a large number of parameters into multiple matrices with smaller numbers of parameters through low-rank decomposition, it is possible to learn multiple matrices with smaller numbers of parameters during the quantization process without having to learn the weight compensation matrix with a large number of parameters, greatly reducing the number of parameters to be learned during the quantization process and reducing the number of iterations and resource occupancy during the quantization process. Especially for large language models with a huge number of parameters, by performing low-rank decomposition, the weight compensation matrix with a huge number of parameters can be decomposed into multiple matrices with very small numbers of parameters, thus greatly reducing the cost of learning the weight compensation matrix.
[0019] In a possible implementation, the difference between the first output data and the second output data is the weighted average between the first difference value and the second difference value. The first difference value is the Euclidean distance between the first output data and the second output data, and the second difference value is the relative entropy between the first output data and the second output data.
[0020] In this solution, by constructing the difference between the outputs of the network structure based on the weighted average of the Euclidean distance and the relative entropy between the outputs, it is possible to effectively handle the outliers in the feature data and improve the robustness of subsequent optimization of the network structure based on the difference.
[0021] In a possible implementation, during each round of quantization process, the above model quantization method further includes: selecting one or more outliers from the values of multiple weight parameters of the network structure. That is, determining one or more outliers among the values of the multiple weight parameters. One or more outliers refer to one or more values among the values of the multiple weight parameters that are significantly different from other values.
[0022] Then, set the values of the weight parameters corresponding to the one or more outliers to a preset threshold. Among them, the preset threshold can be a preset value determined based on the distribution range of the multiple weight parameters, such as 0 or any value within the distribution range of the other weight parameters except the outliers among the multiple weight parameters.
[0023] In this solution, by determining the outliers in the weight parameters and adjusting the values of the weight parameters belonging to the outliers to a preset threshold, the distribution range of the weight parameters can be narrowed, thereby reducing the difficulty of learning quantization parameters and facilitating the learning of better quantization parameters, and improving the accuracy of the target model after quantization. Moreover, since the outliers of the activation values are to some extent caused by the outliers of the weight parameters, after adjusting the outliers of the weight parameters, the outliers of the activation values can be improved simultaneously, which is beneficial to learning better quantization parameters of the activation values.
[0024] In a possible implementation manner, one or more outliers are determined based on the values of multiple weight parameters in the network structure. Specifically, it may include: first, based on the values of the multiple weight parameters, a coarse-grained outlier interval is determined, where the values of the weight parameters located in the coarse-grained outlier interval are all greater than the first threshold or all less than the second threshold, and the first threshold and the second threshold are determined based on the distribution of the values of the multiple weight parameters. That is, first, based on the distribution of the values of the multiple weight parameters, the first threshold or the second threshold is determined, and the range where the values of the weight parameters greater than the first threshold are located is determined as the coarse-grained outlier interval, or the range where the values of the weight parameters less than the second threshold are located is determined as the coarse-grained outlier interval. Specifically, the first threshold can be a positive value, and the second threshold is a negative value. Generally speaking, the values in the coarse-grained outlier interval are values far from 0.
[0025] Then, in the coarse-grained outlier interval, a fine-grained outlier interval is further determined. Among them, the fine-grained outlier interval is within the range of the coarse-grained outlier interval, and the distribution of the values within the fine-grained outlier interval satisfies a preset condition. After determining the fine-grained outlier interval, the values of the weight parameters located in the fine-grained outlier interval can be determined as the above one or more outliers.
[0026] In this solution, by first determining a threshold based on the distribution of the values of the weight parameters, determining the coarse-grained outlier interval based on the threshold, and then further searching for the fine-grained interval that satisfies the preset condition within the coarse-grained outlier interval, the efficiency of determining outliers can be effectively improved.
[0027] In a possible implementation manner, during each round of quantization process, the above model quantization method further includes: selecting at least one outlier from the values of multiple activation values in the network structure; scaling the values of the activation values corresponding to the at least one outlier according to a target ratio.
[0028] In this solution, by determining the outliers in the activation values and scaling the activation values belonging to the outliers, the distribution range of the activation values can be narrowed, thereby reducing the difficulty of learning quantization parameters and facilitating the learning of better quantization parameters, and improving the accuracy of the target model after quantization.
[0029] In a possible implementation, the above-mentioned target model is a large language model.
[0030] In a possible implementation, each of the multiple network structures of the target model includes one or more Transformer network blocks.
[0031] The second aspect of this application provides a model quantization device, including: an acquisition module, configured to acquire a target model, where the target model includes multiple network structures connected in sequence; a processing module, configured to perform multiple rounds of quantization processes to obtain a quantized target model, where the quantized target model includes quantized multiple network structures, and the multiple rounds of quantization processes are executed sequentially based on the connection order of the multiple network structures; wherein, in the multiple rounds of quantization processes, each round of quantization process quantizes at least two consecutive network structures among the multiple network structures, and each round of quantization process quantizes only a part of the network structures among the multiple network structures. For any two consecutive quantization processes in the multiple rounds of quantization processes, there is at least one repeated network structure between the network structures quantized in the latter round of quantization process and the network structures quantized in the former round of quantization process.
[0032] In a possible implementation, the first N network structures quantized in the latter round of quantization process are the same N network structures as the last N network structures quantized in the former round of quantization process, where N is an integer greater than or equal to 1.
[0033] In a possible implementation, the processing module is further configured to: input first input data into at least two network structures to obtain first output data, where the at least two network structures are the network structures quantized in the target round of quantization process, the first input data is the output of the network structures before the at least two network structures, and the target round of quantization process is one round of quantization process in the multiple rounds of quantization processes; input second input data into the at least two network structures quantized based on quantization parameters to obtain second output data, where the second input data is the output after quantization of the network structures before the at least two network structures; and update the quantization parameters of the at least two network structures based on the difference between the first output data and the second output data.
[0034] In a possible implementation, the quantization parameters include a quantization step size and a weight compensation matrix. The quantization step size is used to update the weight parameters in the network structure, and the weight compensation matrix is used to compensate the updated weight parameters; wherein, the weight compensation matrix is low-rank decomposed into multiple matrices, the total number of parameters of the multiple matrices is less than the number of parameters of the weight compensation matrix, and the update of the weight compensation matrix is achieved by updating the multiple matrices.
[0035] In a possible implementation, the difference is the weighted average between the first difference value and the second difference value, the first difference value is the Euclidean distance between the first output data and the second output data, and the second difference value is the relative entropy between the first output data and the second output data.
[0036] In a possible implementation, during each round of quantization process, the processing module is further configured to: select one or more outliers from the values of multiple weight parameters of the network structure; set the values of the weight parameters corresponding to the one or more outliers to a preset threshold.
[0037] In a possible implementation, the processing module is further configured to: determine a coarse-grained outlier interval based on the values of multiple weight parameters, where the values of the weight parameters located in the coarse-grained outlier interval are all greater than the first threshold or all less than the second threshold, and the first threshold or the second threshold is determined based on the distribution of the values of multiple weight parameters; determine a fine-grained outlier interval based on the coarse-grained outlier interval, the fine-grained outlier interval is within the range of the coarse-grained outlier interval, and the distribution of the values within the fine-grained outlier interval satisfies a preset condition, and the values of the weight parameters located in the fine-grained outlier interval belong to one or more outliers.
[0038] In a possible implementation, during each round of quantization process, the processing module is further configured to: select at least one outlier from the values of multiple activation values of the network structure; scale the values of the activation values corresponding to the at least one outlier according to a target ratio.
[0039] In a possible implementation, the target model is a large language model.
[0040] In a possible implementation, each of the multiple network structures includes one or more Transformer network blocks.
[0041] The third aspect of the present application provides a model quantization device, which may include a processor, the processor is coupled with a memory, and the memory stores program instructions. When the program instructions stored in the memory are executed by the processor, the method of the first aspect or any implementation manner of the first aspect is implemented. For the steps executed by the processor in each possible implementation manner of the first aspect, reference may be specifically made to the first aspect, and details are not described herein again.
[0042] The fourth aspect of the present application provides a computer-readable storage medium, in which a computer program is stored. When it runs on a computer, the computer is enabled to execute the method of any implementation manner of the first aspect.
[0043] The fifth aspect of the present application provides a circuit system, the circuit system includes a processing circuit, and the processing circuit is configured to execute the method of any implementation manner of the first aspect.
[0044] The sixth aspect of the present application provides a computer program product, which, when running on a computer, enables the computer to execute the method according to any implementation manner of the first aspect above.
[0045] The seventh aspect of the present application provides a chip system, which includes a processor for supporting an electronic device to implement the functions involved in any implementation manner of the first aspect above. For example, the processor is used to process the data and / or information involved in the above method. In a possible design, the chip system further includes a memory for storing necessary program instructions and data of the electronic device. The chip system can be composed of chips or include chips and other discrete devices.
[0046] For the beneficial effects of the second to seventh aspects above, reference can be made to the introduction of the first aspect above, and details are not described herein again. Description of the Drawings
[0047] Figure 1 It is a schematic diagram of a system architecture 100 provided by an embodiment of the present application;
[0048] Figure 2 It is a schematic flowchart of a model quantization method provided by an embodiment of the present application;
[0049] Figure 3A It is a schematic diagram of a process of performing multiple rounds of quantization on a target model provided by an embodiment of the present application;
[0050] Figure 3B It is another schematic diagram of a process of performing multiple rounds of quantization on a target model provided by an embodiment of the present application;
[0051] Figure 3C It is still another schematic diagram of a process of performing multiple rounds of quantization on a target model provided by an embodiment of the present application;
[0052] Figure 4 It is a schematic flowchart of a process of performing quantization on a target model provided by an embodiment of the present application;
[0053] Figure 5 It is a schematic diagram of performing low-rank decomposition on a weight compensation matrix provided by an embodiment of the present application;
[0054] Figure 6 It is a schematic diagram for comparing the distribution of weight parameters provided by an embodiment of the present application;
[0055] Figure 7 It is another schematic diagram for comparing the distribution of weight parameters provided by an embodiment of the present application;
[0056] Figure 8 It is a schematic flowchart of performing preprocessing on weight parameters provided by an embodiment of the present application;
[0057] Figure 9 A comparative schematic diagram of the distribution of an activation value provided by an embodiment of the present application;
[0058] Figure 10 Another comparative schematic diagram of the distribution of an activation value provided by an embodiment of the present application;
[0059] Figure 11 A schematic structural diagram of a model quantization device provided by an embodiment of the present application;
[0060] Figure 12 A schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0061] Figure 13 A schematic structural diagram of a chip provided by an embodiment of the present application;
[0062] Figure 14 A schematic structural diagram of a computer-readable storage medium provided by an embodiment of the present application. Detailed implementation manners
[0063] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the embodiments of the present application will be described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Those of ordinary skill in the art will know that with the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0064] In the description and claims of this application and the above-mentioned drawings, terms such as "first" and "second" are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such descriptions can be interchanged under appropriate circumstances so that the embodiments can be implemented in an order other than that shown or described in this application. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules does not have to be limited to those steps or modules clearly listed, but may include other steps or modules not clearly listed or inherent to these processes, methods, products, or devices. In this application, the naming or numbering of steps does not mean that the steps in the method flow must be executed in the time / logical sequence indicated by the naming or numbering. The named or numbered process steps can be changed in the execution order according to the technical purpose to be achieved, as long as the same or similar technical effects can be achieved. The division of units in this application is a logical division, and there may be other division methods in actual implementation. For example, multiple units can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection between units can be electrical or other similar forms, which are not limited in this application. And the units or subunits described as separate components may or may not be physically separated, may or may not be physical units, or may be distributed to multiple circuit units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this application.
[0065] For ease of understanding, some technical terms related to the embodiments of this application are introduced below.
[0066] (1) Model quantization
[0067] Model quantization is actually a compression technology for neural network models. Specifically, it uses a lower-bit width to represent the weight parameters and feature data in the neural network model, thereby saving memory space and reducing the amount of computation. Generally, model quantization refers to the process of mapping the parameters of a neural network model from single-precision floating-point numbers (32-bit floating-point numbers, FP32) to n-bit positions. Simply put, it is to establish a data mapping relationship between fixed-point numbers and floating-point numbers and other data, so as to obtain better benefits at the cost of a smaller accuracy loss. For example, by mapping the parameters of a neural network model from FP32 to 8-bit signed integers (INT8), 4-fold parameter compression can be achieved, enabling faster calculations while compressing memory, thereby effectively improving the performance of the model.
[0068] Generally speaking, model quantization is usually divided into two types: Post-training Quantization (PTQ) and Quantization-aware Training (QAT). PTQ is a process that can quantize a pre-trained neural network model with only a small amount of unlabeled calibration data set, and is often used to compress neural network models that use a relatively high bit width to represent parameters. QAT, on the other hand, requires a complete data set to train the neural network model and simulates quantization operations during the training process, so that the quantized model (i.e., the model after quantization) can further converge to the optimal point, and is often used in the accuracy recovery process after a large accuracy loss occurs in the quantized model.
[0069] In addition, during the model quantization process, generally the weight parameters and activation values in the neural network model are quantized. The weight parameters refer to the parameters in each neural network layer of the neural network model that process the input data, and are variables that can be adjusted during the training process. The activation value usually refers to the feature data obtained after each neural network layer processes the input data. Since the output of the previous neural network layer in the neural network model will be used as the input of the next neural network layer, the activation value of the previous neural network layer is usually used as the input of the next neural network layer. By quantizing the weight parameters and activation values, the number of parameters and the amount of computation of the neural network model can be greatly reduced, thereby improving the running performance of the neural network model.
[0070] (2) Neural network
[0071] A neural network can be composed of neural units, and a neural unit can refer to an arithmetic unit with x s (i.e., input data) and intercept 1 as inputs, and the output of this arithmetic unit can be:
[0072]
[0073] where s = 1, 2,... n, and n is a natural number greater than 1, and W s is x sThe weight parameter is \(w\), \(b\) is the bias of the neuron. \(f\) is the activation function of the neuron, which is used to introduce non - linear characteristics into the neural network to convert the input signal in the neuron into an output signal. The output signal of this activation function can be used as the input of the next convolutional layer. The activation function can be the sigmoid function. A neural network is a network formed by connecting multiple such single neurons together, that is, the output of one neuron can be the input of another neuron. The input of each neuron can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, and the local receptive field can be a region composed of several neurons.
[0074] (3) Deep Neural Network (DNN)
[0075] A deep neural network, also known as a multi - layer neural network, can be understood as a neural network with many hidden layers. Here, "many" does not have a specific measurement standard. Dividing the DNN according to the positions of different layers, the neural network inside the DNN can be divided into three categories: the input layer, the hidden layer, and the output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the middle layers are all hidden layers. The layers are fully connected, that is, any neuron in the \(i\) - th layer must be connected to any neuron in the \((i + 1)\) - th layer. Although the DNN looks very complex, in terms of the work of each layer, it is actually not complex. Simply put, it is the following linear relationship expression: Among them, \(\mathbf{x}\) is the input vector, \(\mathbf{y}\) is the output vector, \(\mathbf{b}\) is the offset vector, \(W\) is the weight matrix (also known as the coefficient), and \(\alpha(\cdot)\) is the activation function. Each layer is just a simple operation on the input vector \(\mathbf{x}\) to obtain the output vector Since the DNN has many layers, the number of coefficients \(W\) and offset vectors \(\mathbf{b}\) is also very large. The definitions of these parameters in the DNN are as follows: Taking the coefficient \(W\) as an example: Suppose in a three - layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer where the coefficient \(W\) is located, and the subscript corresponds to the index 2 of the output third layer and the index 4 of the input second layer. In summary: The coefficient from the \(k\) - th neuron in the \((L - 1)\) - th layer to the \(j\) - th neuron in the \(L\) - th layer is defined as It should be noted that there is no W parameter in the input layer. In a deep neural network, more hidden layers enable the network to better depict complex situations in the real world. Theoretically, the more parameters a model has, the higher its complexity and the greater its "capacity", which means it can complete more complex learning tasks. Training a deep neural network is a process of learning the weight matrix, and its ultimate goal is to obtain the weight matrices of all layers of the trained deep neural network (the weight matrix formed by vectors W of many layers).
[0076] (4) Convolutional Neural Network (CNN)
[0077] A convolutional neural network is a deep neural network with a convolutional structure. A convolutional neural network contains a feature extractor composed of convolutional layers and subsampling layers. This feature extractor can be regarded as a filter, and the convolution process can be regarded as convolving a trainable filter with an input feature map. A convolutional layer refers to the layer of neural units in a convolutional neural network that performs convolution processing on the input signal. In the convolutional layer of a convolutional neural network, a neural unit can only be connected to some neighboring layer neural units. In a convolutional layer, there are usually several feature planes, and each feature plane can be composed of some rectangularly arranged neural units. The neural units in the same feature plane share weights, and the shared weight here is the convolution kernel.
[0078] The convolution kernel can be initialized in the form of a matrix of random size, and during the training process of the convolutional neural network, the convolution kernel can learn reasonable weights. Additionally, the direct benefit of sharing weights is to reduce the connections between layers of the convolutional neural network while also reducing the risk of overfitting.
[0079] (5) Attention Network
[0080] An attention network is a network model that uses the attention mechanism to improve the model training speed. Currently, typical attention networks include Transformer. A model applying the attention mechanism can assign different weights to each part of the input sequence, thereby extracting more important feature information from the input sequence and enabling the model to finally obtain a more accurate output.
[0081] Transformer is a neural network architecture based on the self-attention mechanism. Transformer is usually composed of an encoder and a decoder, and each part is composed of multiple layers of stacked self-attention and feed-forward neural networks. The self-attention mechanism allows the model to consider relevant information at all positions simultaneously when processing the sequence, thus solving the problem of long-distance dependencies.
[0082] (6) Large language model (LLM)
[0083] A large language model refers to a deep learning model trained with a large amount of text data, which can generate natural language text or understand the meaning of language text. Large language models can handle various natural language tasks, such as text classification, question answering, dialogue, etc., and are an important approach to artificial intelligence.
[0084] (7) Outlier
[0085] Outliers, usually also known as extreme values, refer to those values in the data that are significantly different from other values when there are one or several values with large differences compared to other values.
[0086] (8) Quartile
[0087] Quartiles, also known as quartile points, refer to the values at the three splitting points when all values are arranged in ascending order in statistics. Quartiles divide all data into four equal parts through three points, with each part containing 25% of the data.
[0088] There are three quartiles. The first quartile is called the lower quartile, the second quartile is the median, and the third quartile is called the upper quartile, denoted by Q1, Q2, and Q3 respectively.
[0089] The first quartile (Q1), also known as the "lower quartile", is equal to the 25th percentile of all the values in the sample arranged in ascending order.
[0090] The second quartile (Q2), also known as the "median", is equal to the 50th percentile of all the values in the sample arranged in ascending order.
[0091] The third quartile (Q3), also known as the "upper quartile", is equal to the 75th percentile of all the values in the sample arranged in ascending order.
[0092] The difference between the third quartile and the first quartile is also called the interquartile range (IQR).
[0093] (9) Euclidean distance
[0094] Euclidean distance, also known as Euclidean metric or L2 distance, is the most common distance measure, which measures the absolute distance between two points in a multi-dimensional space.
[0095] (10) Relative entropy
[0096] Relative entropy, also known as Kullback-Leibler divergence (KL divergence) or information divergence, is an asymmetric measure of the difference between two probability distributions. In information theory, relative entropy is equivalent to the difference in the Shannon entropy of two probability distributions.
[0097] Relative entropy is the loss function of some optimization algorithms, such as the Expectation-Maximization algorithm (EM). At this time, one of the probability distributions participating in the calculation is the true distribution, and the other is the theoretical (fitted) distribution. Relative entropy represents the information loss generated when using the theoretical distribution to fit the true distribution.
[0098] Through research by the applicant, it is found that the model quantization technology in the related art usually loads the entire neural network model onto the processing device and then performs quantization on the entire neural network model to ensure the accuracy of the quantized neural network model. However, with the rise of large-scale models such as large language models, large language models with a huge number of parameters cannot be fully loaded onto the processing device at one time. Therefore, the model quantization methods in the related art are difficult to perform quantization on models with a large number of parameters such as large language models.
[0099] Based on this, the embodiments of the present application provide a model quantization method, which divides the target model into multiple parts based on the connection relationship of the network structure in the target model, and performs multiple rounds of quantization processes on the target model. The multiple rounds of quantization processes are sequentially performed based on the network structure of the model. Each round of quantization process performs quantization on a part of the network structure in the target model, thereby reducing the number of parameters loaded onto the processing device at one time and ensuring that the quantization process of the target model can be smoothly executed. Moreover, during the quantization process, there are partially overlapping network structures in any two consecutive rounds of quantization processes, fully considering the dependency relationship between the network structures, ensuring that a joint optimization relationship can be established between the network structures during the multiple rounds of quantization processes, effectively avoiding the situation where each round of quantization process independently performs quantization on different network structures and cannot achieve the global optimum, and ensuring the accuracy of the finally quantized target model.
[0100] Please refer to Figure 1 , Figure 1 which is a schematic diagram of a system architecture 100 provided by the embodiments of the present application. As Figure 1As shown, in the system architecture 100, the execution device 110 can be implemented by one or more servers. Optionally, the execution device 110 cooperates with other computing devices, such as devices for data storage, routers, load balancers, etc.; the execution device 110 can be arranged on one physical site or distributed across multiple physical sites. The execution device 110 can use the data in the data storage system 120 or call the program code in the data storage system 120 to implement the model quantization method provided by the embodiments of the present application.
[0101] Users can operate their respective user devices (such as local device 101 and local device 102) to interact with the execution device 110. Each local device can represent any computing device, such as a personal computer, computer workstation, smartphone, tablet computer, laptop computer, and intelligent vehicle, etc.
[0102] The local device of each user can interact with the execution device 110 through a communication network of any communication mechanism / communication standard. The communication network can be a wide area network, local area network, point-to-point connection, etc., or any combination thereof.
[0103] In one implementation, the execution device 110 is used to implement the model quantization method provided by the embodiments of the present application to obtain a quantized model. And, during the process where the local device 101 and the local device 102 need to use the model to perform inference, the execution device 110 processes the data provided by the user based on the quantized model, and then returns the corresponding processing results to the local device 101 and the local device 102.
[0104] In another implementation, the execution device 110 is used to implement the model quantization method provided by the embodiments of the present application, and send the obtained quantized model to the local device 101 and the local device 102. In this way, the local device 101 and the local device 102 can deploy the quantized model locally, and thus perform data processing based on the quantized model.
[0105] In another implementation, one or more aspects of the execution device 110 can be implemented by each local device. For example, the local device 101 can provide local data or feedback calculation results for the execution device 110, or implement the model quantization method provided by the embodiments of the present application.
[0106] Generally speaking, the model quantization method provided by the embodiments of the present application can be applied to electronic devices, such as the above-mentioned execution device 110, local device 101, or local device 102.
[0107] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a model quantization method provided by the embodiments of the present application. AsFigure 2 As shown in Figure 2 , the model quantization method provided by the embodiment of the present application includes the following steps 201-203.
[0108] Step 201, obtain a target model.
[0109] In this embodiment, the target model is a neural network model that needs to be quantized, and the target model is a model trained based on the training data in the training set.
[0110] Exemplarily, the target model is, for example, a model of types such as a large language model, a convolutional neural network, an attention network, or a recurrent neural network, and this embodiment does not make specific limitations thereto. Moreover, in the case where the target model is a neural network model of different types, the target model can be, for example, applied to perform natural language processing (NLP) tasks or computer vision tasks, and this embodiment does not limit the tasks to which the target model is applied.
[0111] It should be noted that the target model is a neural network model with a large number of parameters, and it is difficult for a conventional AI processing device to load the entire target model at one time to perform the quantization process. For example, in the case where the target model is a large language model, the target model often includes hundreds of billions of parameters, and it is difficult for existing graphics processing unit (GPU) clusters to load the entire target model at one time.
[0112] Step 202, divide the target model into a plurality of network structures connected in sequence.
[0113] In this embodiment, based on the network structure of the target model itself, the target model can be divided into multiple parts, and each part corresponds to a network structure, so as to realize dividing the target model into a plurality of network structures connected in sequence. Among them, in the plurality of network structures obtained by division, the output of the previous network structure is the input of the next network structure, that is, there is an input-output dependency relationship between the front and rear network structures. In addition, the output of the previous network structure is usually feature data (i.e., a feature matrix), and the feature data output by the previous network structure may be one or more (such as outputting one or more feature matrices), and specific limitations are not made herein.
[0114] For example, in the case where the target model is a large language model, the target model is usually composed of multiple Transformer network blocks connected together. For example, the target model includes 20-30 sequentially connected Transformer network blocks. Therefore, when performing network structure partitioning on the target model, one or more Transformer network blocks can be used as a network structure, so as to obtain multiple sequentially connected network structures.
[0115] That is to say, for the multiple sequentially connected network structures obtained by partitioning the target model, each network structure in the multiple network structures includes one or more Transformer network blocks. For example, for a target model including 30 sequentially connected Transformer network blocks, the target model can be partitioned into 30 network structures, and each network structure corresponds to a Transformer network block.
[0116] Step 203: Perform multiple rounds of quantization processes to obtain a quantized target model, and the quantized target model includes multiple quantized network structures.
[0117] In this embodiment, the multiple rounds of quantization processes are sequentially performed based on the connection order of the multiple network structures, that is, one round of quantization process is performed each time. After sequentially performing multiple rounds of quantization processes, it is possible to perform quantization on all network structures in the target model, so as to obtain a quantized target model, and this quantized target model is actually composed of multiple quantized network structures. That is, during the process of performing multiple rounds of quantization on the target model, it can be started from the first network structure in the target model and then sequentially perform quantization on other network structures until the quantization of the last network structure in the target model is completed.
[0118] During the multiple rounds of quantization processes performed on the target model, each round of quantization process quantizes at least two consecutive network structures among the multiple network structures. Each round of quantization process quantizes some of the network structures among the multiple network structures, and the network structures quantized in different rounds of quantization processes are not exactly the same. In addition, for any two consecutive rounds of quantization processes in the multiple rounds of quantization processes, there is at least one repeated network structure between the network structures quantized in the latter round of quantization process and the network structures quantized in the former round of quantization process.
[0119] That is to say, for the part of the network structure that has been quantized in the previous round of quantization, it will still be quantized again in the subsequent round of quantization. In this way, for the two rounds of quantization before and after, there will be a part of the overlapping network structure that is jointly quantized with another part of the network structure in the previous round of quantization and is also jointly quantized with yet another part of the network structure in the subsequent round of quantization, thereby establishing a connection relationship between different rounds of quantization processes and effectively avoiding the situation where each round of quantization process independently quantizes different network structures and fails to achieve the global optimum.
[0120] Exemplarily, for any two consecutive rounds of quantization processes in the multi-round quantization process, the first N network structures quantized in the subsequent round of quantization are the same N network structures as the last N network structures quantized in the previous round of quantization, where N is an integer greater than or equal to 1. Among them, the total number of network structures quantized in the subsequent round of quantization and the total number of network structures quantized in the previous round of quantization can be the same or different. For example, assume that both the previous round of quantization process and the subsequent round of quantization process are quantizing M network structures. Then, the last N network structures among the M network structures quantized in the previous round of quantization are actually the same N network structures as the first N network structures among the M network structures quantized in the subsequent round of quantization, that is, there are N overlapping network structures in the two rounds of quantization processes before and after.
[0121] Among them, the number of network structures quantized in each round of quantization process can be determined according to the capabilities of the processing device, for example, the number is 2 - 6. Similarly, the number of overlapping network structures in the two rounds of quantization processes before and after can also be determined according to the capabilities of the processing device and the accuracy of the model after quantization. In the case where the capabilities of the processing device are strong and the accuracy of the model after quantization is high, the number of overlapping network structures in the two rounds of quantization processes before and after can be set to be relatively large.
[0122] In a possible example, please refer to Figure 3A , Figure 3A which is a schematic diagram of a multi-round quantization process for a target model provided by an embodiment of the present application. In Figure 3A the example shown, when each round of quantization process for the target model is to quantize two consecutive network structures among multiple network structures, the first network structure quantized in the subsequent round of quantization is the same network structure as the last network structure quantized in the previous round of quantization.
[0123] As Figure 3AAs shown, the target model includes n network structures connected in sequence, namely Network Structure 1 - Network Structure n. In the first round of quantization, Network Structure 1 and Network Structure 2 are quantized; in the second round of quantization, Network Structure 2 and Network Structure 3 are quantized. That is, after Network Structure 2 finishes quantization jointly with Network Structure 1 in the first round of quantization, it continues to perform quantization jointly with Network Structure 3 in the second round of quantization. Similarly, in the third round of quantization, Network Structure 3 and Network Structure 4 are quantized; in the fourth round of quantization, Network Structure 4 and Network Structure 5 are quantized... and so on, until the quantization of Network Structure n is completed.
[0124] That is to say, in this embodiment, the quantization of the network structure is actually performed in a sliding window manner. Each round of quantization process is actually to quantize multiple network structures included in the sliding window. By sliding the sliding window, different network structures are included in the sliding window to trigger different rounds of quantization processes. And, in two adjacent sliding windows, there will be overlapping network structures, so as to ensure the joint quantization between network structures.
[0125] In another possible example, please refer to Figure 3B , Figure 3B which is a schematic diagram of another multi-round quantization process for the target model provided by the embodiment of the present application. In Figure 3B the example shown, when each round of quantization process for the target model is to quantize three consecutive network structures among multiple network structures, the first network structure quantized in the subsequent round of quantization process is the same network structure as the last network structure quantized in the previous round of quantization process.
[0126] As Figure 3B shown, for the target model including Network Structure 1 - Network Structure n. In the first round of quantization, Network Structure 1, Network Structure 2 and Network Structure 3 are quantized; in the second round of quantization, Network Structure 3, Network Structure 4 and Network Structure 5 are quantized, that is, after Network Structure 3 finishes quantization jointly with Network Structure 1 and Network Structure 2 in the first round of quantization, it continues to perform quantization jointly with Network Structure 4 and Network Structure 5 in the second round of quantization.
[0127] In yet another possible example, please refer to Figure 3C , Figure 3C which is another schematic diagram of a multi-round quantization process for the target model provided by the embodiment of the present application. As Figure 3BAs shown, for the target model including network structure 1 - network structure n. In the first round of quantization process, network structure 1, network structure 2, network structure 3, and network structure 4 are quantized; in the second round of quantization process, network structure 3, network structure 4, network structure 5, and network structure 6 are quantized, that is, after network structure 3 and network structure 4 complete quantization jointly with network structure 1 and network structure 2 in the first round of quantization process, they continue to perform quantization jointly with network structure 5 and network structure 6 in the second round of quantization process.
[0128] Generally speaking, for any two consecutive quantization processes in the multi-round quantization process, the number of network structures quantized in the previous round of quantization process and the number of network structures quantized in the next round of quantization process can be the same or different, and the number of network structures that are repeatedly quantized in the previous and next rounds of quantization processes can be one or more.
[0129] The above introduces the process of performing multi-round quantization on the target model to achieve quantization of the target model. For ease of understanding, the following will detail the specific process of each round of quantization performed on the target model.
[0130] It can be understood that in the multi-round quantization process performed on the target model, each round of quantization process is actually to determine the quantization parameters corresponding to the network structures in the target model, so as to be able to transform the parameters in the network structure based on the quantization parameters (for example, transform the parameters in the network structure from FP32 to INT8) to obtain the network structure after parameter transformation (i.e., the quantized network structure). Among them, when determining the quantization parameters corresponding to the network structures in the target model, it is necessary to ensure that the output of the network structure after parameter transformation based on the determined quantization parameters is as close as possible to the output of the original network structure, that is, the output of the quantized network structure is as close as possible to the output of the network structure before quantization.
[0131] Based on this, in this embodiment, in each round of quantization process, the quantization parameters corresponding to the network structure are constrained by comparing the output difference between the network structure adjusted based on the quantization parameters and the network structure before quantization, so that the output difference between the network structure adjusted based on the finally determined quantization parameters and the network structure before quantization is as small as possible.
[0132] Exemplarily, for the first round of quantization process in the multi-round quantization process, first determine multiple network structures to be quantized in the first round of quantization process. Then, input calibration data (e.g., part of the training data in the training set of the target model) into the multiple network structures to be quantized in the first round of quantization process, and obtain corresponding output data; and, after performing parameter transformation on the multiple network structures to be quantized in the first round of quantization process based on quantization parameters, input the same calibration data into the multiple network structures after performing parameter transformation to obtain output data. In this way, by calculating the difference between the output data of the multiple network structures before parameter transformation (i.e., the original multiple network structures) and the output data of the multiple network structures after parameter transformation, the quantization parameters of the multiple network structures can be updated based on this difference (e.g., by backpropagation of gradients to update the quantization parameters), and finally the output of the multiple network structures after performing parameter transformation based on the quantization parameters can meet the requirements.
[0133] In addition, consider any round of quantization process other than the first round of quantization process in the multi-round quantization process as the target round quantization process. Then, the target round quantization process specifically includes: First, input the first input data into at least two network structures to obtain the first output data. Among them, the at least two network structures into which the first input data is input are the network structures quantized in the target round quantization process, and the first input data is the output of the network structures before the at least two network structures. Since the current at least two network structures are not at the beginning position in the target model, the input of these at least two network structures is actually the feature data (i.e., the first input data) output by the previous network structures. By inputting the first input data into at least two network structures, the first output data output by these at least two network structures before quantization can be obtained.
[0134] Then, input the second input data into at least two network structures after quantization based on quantization parameters to obtain the second output data. The second input data is the output after quantization of the network structures before the at least two network structures. Since in the case where the entire target model is quantized, the input received by the network structures in the middle position is the output of the quantized network structures in the previous position, in this embodiment, the output data of the quantized network structures in the previous position (i.e., the second input data) is used as the input of the middle network structures (i.e., the above at least two network structures).
[0135] Secondly, based on the difference between the first output data and the second output data, the quantization parameters of at least two network structures are updated. Specifically, the quantization parameters of at least two network structures can be updated by means of backpropagation of gradients, so that when at least two network structures are quantized based on the updated quantization parameters, the outputs of the at least two network structures can be closer to the outputs before quantization. That is, the process of updating the quantization parameters based on the difference between the two output data is similar to the way of updating the weight parameters during the model training process. It can be understood that the quantization parameters are updated by constructing a loss function indicating the difference between the output data, and the update target of the quantization parameters is to make the loss function as small as possible (that is, the difference between the two output data is as small as possible).
[0136] In practical applications, by repeatedly executing the above steps of updating the quantization parameters with different input data, the update of the quantization parameters can be continuously performed until the output of the network structure after parameter transformation based on the quantization parameters meets the requirements, or the number of times of repeatedly updating the quantization parameters reaches a certain number, and finally the quantization of the network structure is completed.
[0137] In this solution, during the process of determining the quantization parameters of the network structure, by using the output of the original network structure as a supervision signal to construct a difference with the output of the network structure quantized based on the quantization parameters, and then updating the quantization parameters with the goal of minimizing the output difference, accurate quantization parameters can be obtained, ensuring the accuracy of the network structure quantized based on the quantization parameters.
[0138] Optionally, the difference between the first output data and the second output data is, for example, the weighted average between the first difference value and the second difference value. The first difference value is the Euclidean distance (i.e., L2 distance) between the first output data and the second output data, and the second difference value is the relative entropy (i.e., KL divergence) between the first output data and the second output data. In this solution, by constructing the difference between the outputs of the network structure based on the weighted average of the Euclidean distance and the relative entropy between the outputs, outliers in the feature data can be effectively dealt with, improving the robustness of subsequent optimization of the network structure based on the difference.
[0139] Exemplarily, please refer to Figure 4 , Figure 4 which is a schematic flowchart of a process for quantizing a target model provided by an embodiment of the present application. As Figure 4 shown, during the process of quantizing the target model, the same calibration data needs to be input into the original target model and the target model after parameter transformation using the quantization parameters at the same time (i.e., Figure 4the target model in quantization as shown in , and compare the differences between the outputs of the network structures at the same positions of the original target model and the target model in quantization, and then update the quantization parameters corresponding to the network structure based on the differences in the outputs.
[0140] Specifically, in the first round of quantization process, first perform parameter transformation on the target model based on random quantization parameters to obtain the target model in quantization. Then, input the same calibration data into the original target model and the target model in quantization respectively, and obtain the output of network structure 2 in the original target model and the output of network structure 2 with parameters transformed based on the quantization parameters in the target model in quantization. Secondly, by comparing the output of network structure 2 with the output of network structure 2 after parameter transformation, difference 1 can be obtained. In this way, taking minimizing difference 1 as the goal, the quantization parameters corresponding to network structure 1 and network structure 2 can be updated by means of gradient backpropagation until the final outputs of network structure 1 and network structure 2 after parameter transformation based on the updated quantization parameters are close to the outputs of network structure 1 and network structure 2 without parameter transformation, or the number of times of updating the quantization parameters reaches the preset number of times, so as to obtain the quantized network structure 1 and network structure 2.
[0141] In the second round of quantization process, with the calibration data unchanged, input the calibration data into the quantized network structure 1, and input the output of the quantized network structure 1 into network structure 2 and network structure 3 in the target model in quantization to obtain the output of network structure 3 after parameter transformation. Then, by comparing the output of the original network structure 3 with the output of network structure 3 after parameter transformation, difference 2 can be obtained. In this way, taking minimizing difference 2 as the goal, the quantization parameters corresponding to network structure 2 and network structure 3 can be updated by means of gradient backpropagation until the final outputs of network structure 2 and network structure 3 after parameter transformation based on the updated quantization parameters are close to the outputs of network structure 2 and network structure 3 without parameter transformation, or the number of times of updating the quantization parameters reaches the preset number of times, so as to obtain the quantized network structure 2 and network structure 3.
[0142] It should be noted that network structure 2 actually undergoes two rounds of quantization process. In the first round of quantization process, network structure 2 is jointly executed with network structure 1, and in the second round of quantization process, network structure 2 is jointly executed with network structure 3. Only when network structure 2 has completed two rounds of quantization process in its entirety will the quantization parameters corresponding to network structure 2 be fixed, that is, network structure 2 is considered to have completed the quantization process.
[0143] Similarly, based on a process similar to the second-round quantization process, quantization can continue to be performed on network structure 3 and network structure 4, network structure 4 and network structure 5, and subsequent network structures until the quantization of the last network structure in the target model is completed, thereby obtaining the quantized target model.
[0144] Generally speaking, the quantization parameters of a network structure can include two types. One is the quantization parameter of the weight parameter, and the other is the quantization parameter of the activation value. Based on the quantization parameter of the weight parameter, parameter transformation can be performed on the weight parameter in the network structure to obtain the quantized weight parameter; similarly, based on the quantization parameter of the activation value, parameter transformation can be performed on the activation value in the network structure to obtain the quantized activation value.
[0145] Exemplarily, the quantized weight parameter can be shown as the following formula.
[0146]
[0147] Among them, represents the quantized weight parameter; Δ W represents the quantization step of the weight parameter; W represents the weight parameter; represents rounding down the tensor to the nearest integer.
[0148] In addition, the quantized activation value can be shown as the following formula.
[0149]
[0150] Among them, represents the quantized activation value; Δ x represents the quantization step of the activation value; X represents the activation value.
[0151] Since each round of quantization process actually performs quantization on at least two network structures, the process of optimizing the quantization parameters of the network structure in any round of quantization process can be shown as the following formula.
[0152]
[0153] Among them, argmin(f(x)) represents finding x that makes the function f(x) take the minimum value; Recon() represents finding the difference value; B l,k (W l,k , X l,k ) represents the output of network structure 1 - network structure k; W l,k represents the weight parameter of network structure 1 - network structure k; X l,k represents the activation value of network structure 1 - network structure k; Denote the outputs of network structures 1 - k after performing parameter transformation based on quantization parameters; Denote the quantized weight parameters in network structures 1 - k; Denote the quantized activation values in network structures 1 - k.
[0154] Specifically, the definition of Recon() can be shown as the following formula.
[0155] R = con(f 1 , f 2 ) = ||f1 - f2|| + KLD(softmax(f 1 ), softmax(f 2 ))
[0156] Wherein, KLD() represents calculating the KL divergence; softmax() represents normalization.
[0157] In some embodiments, the quantization parameters of the weight parameters may further include a weight compensation matrix in addition to the above quantization step size. Among them, the quantization step size is used to update the weight parameters in the network structure, while the weight compensation matrix is used to compensate the updated weight parameters. Then, in the process of performing quantization on the target model, it is actually a process of learning the quantization step size and the weight compensation matrix. Since there will be a certain error after updating the weight parameters using the quantization step size, by adding the weight compensation matrix to the updated weight parameters, the error caused by rounding down during weight fixed-point quantization can be compensated. Generally, the values of the elements in the weight compensation matrix are usually 0 and 1.
[0158] Optionally, the above weight compensation matrix is low-rank decomposed into multiple matrices, the total number of parameters of the multiple matrices is less than the number of parameters of the weight compensation matrix, and the update of the weight compensation matrix is achieved by updating the multiple matrices. Among them, low-rank decomposition is a way to decompose a high-dimensional matrix into low-rank matrices, which can effectively reduce the dimension of the data and reduce redundant information.
[0159] Exemplarily, please refer to Figure 5 , Figure 5 which is a schematic diagram of performing low-rank decomposition on the weight compensation matrix provided by an embodiment of the present application. As Figure 5 shown, the dimension of the weight compensation matrix A is d×k. After performing low-rank decomposition on the weight compensation matrix A, the weight compensation matrix A can be decomposed into matrix A 1 and matrix A 2 . Among them, the dimension of matrix A 1 is d×r, the dimension of matrix A 1 is r×k, and r is much smaller than the minimum value of d and k.
[0160] It should be noted that Figure 5 The low-rank decomposition method shown in Figure 5 decomposes a weight compensation matrix into two matrices in a low-rank manner. In practical applications, it can also be to decompose a weight compensation matrix into more than two matrices, which is specifically determined by the low-rank decomposition method, and this embodiment does not make specific limitations on this.
[0161] That is to say, by performing low-rank decomposition on a weight compensation matrix with a large number of parameters, it is possible to represent a weight compensation matrix with a large number of parameters using multiple matrices with a small number of parameters, thereby avoiding the need to store and process a weight compensation matrix with a huge number of parameters during the quantization process.
[0162] In this solution, by decomposing a weight compensation matrix with a large number of parameters into multiple matrices with a small number of parameters, it is possible to learn multiple matrices with a small number of parameters during the quantization process without learning a weight compensation matrix with a large number of parameters, greatly reducing the number of parameters to be learned during the quantization process, reducing the number of iterations and resource occupancy during the quantization process. Especially for large language models with a huge number of parameters, by performing low-rank decomposition, a weight compensation matrix with a huge number of parameters can be decomposed into multiple matrices with a very small number of parameters, thereby greatly reducing the cost of learning the weight compensation matrix.
[0163] Generally speaking, in a network structure, different neural network layers correspond to different quantization parameters. For example, the quantization parameter of the weight parameter is actually to convert the weight parameter in the current neural network layer from a floating-point number to a fixed-point number, and the distribution range of the weight parameters in the current neural network layer will affect the actual value of the quantization parameter, and ultimately affect the specific value of each weight parameter after being converted into a fixed-point number. When the values of some weight parameters in the neural network layer are quite different from the values of most other weight parameters, it will lead to a large distribution range of the weight parameters, and it is difficult to determine effective quantization parameters.
[0164] Based on this, in this embodiment, during each round of quantization process, one or more outliers can be selected from the values of multiple weight parameters in the same neural network layer in the network structure, and the values of the multiple weight parameters include one or more outliers. That is, one or more outliers are determined among the values of the multiple weight parameters. One or more outliers refer to one or more values that are quite different from other values among the values of the multiple weight parameters. For example, assume that there are currently 1000 values of weight parameters, among which the values of 998 weight parameters are distributed in the range of [0, 10], while the values of the other two weight parameters are 150 and 160 respectively, then the values of these two weight parameters can be determined as outliers.
[0165] Then, set the values of the weight parameters corresponding to one or more outliers to a preset threshold, that is, modify the values of the weight parameters belonging to the outliers to the preset threshold. Here, the preset threshold can be a preset value determined based on the distribution range of multiple weight parameters, such as 0 or any value within the distribution range of the other weight parameters excluding the outliers among the multiple weight parameters.
[0166] It should be noted that the process of determining the outliers of the weight parameters is performed independently for each neural network layer in the network structure, that is, different neural network layers may determine different outliers. When determining the outliers of any neural network layer, the above steps can be used for execution, which will not be elaborated in this embodiment.
[0167] In this solution, by determining the outliers in the weight parameters and adjusting the values of the weight parameters belonging to the outliers to the preset threshold, the distribution range of the weight parameters can be narrowed, thereby reducing the difficulty of learning the quantization parameters and facilitating learning better quantization parameters, and improving the accuracy of the target model after quantization. Moreover, since the outliers of the activation values are to some extent caused by the outliers of the weight parameters, after adjusting the outliers of the weight parameters, the outliers of the activation values can be improved simultaneously, which is beneficial to learning better quantization parameters of the activation values.
[0168] Exemplarily, please refer to Figure 6 and Figure 7 , Figure 6 which is a comparison schematic diagram of the distribution of a kind of weight parameters provided by an embodiment of the present application; Figure 7 which is another comparison schematic diagram of the distribution of weight parameters provided by an embodiment of the present application. As Figure 6 shown, before the outliers are removed, the distribution range of the weight parameters is large, while after the outliers are removed, the distribution range of the weight parameters becomes smaller. As Figure 7 shown, before the outliers are removed, the distribution range of the weight parameters is [0, 4], and after the outliers are removed, the distribution range of the weight parameters becomes [0, 2.2], effectively narrowing the distribution range of the weight parameters and facilitating the subsequent determination of the quantization parameters of the weight parameters.
[0169] Optionally, the process of selecting outliers from the values of multiple weight parameters in the network structure may specifically include: First, based on the values of the multiple weight parameters, a coarse-grained outlier interval is determined, where the values of the weight parameters located in the coarse-grained outlier interval are all greater than the first threshold or all less than the second threshold, and the first threshold or the second threshold is determined based on the distribution of the values of the multiple weight parameters. That is, first, based on the distribution of the values of the multiple weight parameters, the first threshold or the second threshold is determined, and the range where the values of the weight parameters greater than the first threshold are located is determined as the coarse-grained outlier interval, or the range where the values of the weight parameters less than the second threshold are located is determined as the coarse-grained outlier interval. Specifically, the first threshold may be a positive value, and the second threshold is a negative value. Generally speaking, the values in the coarse-grained outlier interval are values far from 0. Among them, the values of the above-mentioned multiple weight parameters are all positive or negative. For the weight parameters in the same neural network layer, in this embodiment, a coarse-grained outlier interval is determined for the multiple weight parameters with positive values, and another coarse-grained outlier interval is determined for the multiple weight parameters with negative values, so as to further determine outliers in each coarse-grained outlier interval.
[0170] Then, in the coarse-grained outlier interval, a fine-grained outlier interval is further determined. Among them, the coarse-grained outlier interval is relative to the fine-grained outlier interval. The fine-grained outlier interval is within the range of the coarse-grained outlier interval, and the distribution of the values within the fine-grained outlier interval satisfies a preset condition. After determining the fine-grained outlier interval, it can be determined that the values of the weight parameters located in the fine-grained outlier interval belong to one or more of the above-mentioned outliers. That is, if the value of a certain weight parameter falls within the fine-grained outlier interval, it can be considered that the value of this weight parameter is an outlier. Generally, the coarse-grained outlier interval can be determined to be divided into a reserved interval and a fine-grained outlier interval. Then, the preset condition satisfied by the distribution of the values within the fine-grained outlier interval may specifically be: the distance between the fine-grained outlier interval and the reserved interval is as large as possible, and the numerical distribution within the reserved interval is as close as possible.
[0171] Among them, there are various ways to determine the first threshold and the second threshold for determining the coarse-grained outlier interval. For example, taking the first threshold as an example, the first threshold can be determined based on the third quartile and the interquartile range among the multiple weight parameters. Or, the first threshold can be a numerical value at a specific position among the multiple weight parameters. For example, the first threshold is the numerical value at the 90% position after arranging the multiple weight parameters from small to large. Generally speaking, this embodiment does not limit the way to determine the first threshold and the second threshold.
[0172] Exemplarily, please refer to Figure 8 , Figure 8 which is a schematic flowchart of a process for performing preprocessing on weight parameters provided by an embodiment of the present application. As Figure 8As shown in the figure, the process of performing preprocessing on the weight parameters in the network structure includes the following steps 801-805.
[0173] Step 801: Calculate the quartiles and the interquartile range based on the distribution of the weight parameters.
[0174] Specifically, for multiple weight parameters in a neural network layer, the multiple weight parameters can be arranged in ascending order, and the first quartile Q 1 and the third quartile Q 3 are determined. The first quartile Q 1 is specifically the value of the 25% of the multiple weight parameters arranged in ascending order; the third quartile Q 3 is specifically the value of the 75% of the multiple weight parameters arranged in ascending order. The interquartile range (IQP) is the difference between the third quartile Q 3 and the first quartile Q 1 .
[0175] Step 802: Determine the interval threshold based on the quartiles and the interquartile range.
[0176] Among them, the interval threshold can specifically be: T = Q 3 + λ 1 IQR. Among them, λ 1 is a hyperparameter, and the value of λ 1 can specifically be 1.5.
[0177] Step 803: Determine that the interval to which the weight parameters greater than the interval threshold belong is the coarse-grained outlier interval.
[0178] Assume that the coarse-grained outlier interval is O, then the coarse-grained outlier interval O can specifically be expressed as: O = {x|x > T, x ∈ X}. Among them, x is the value of the weight parameter, and X is the distribution range of the values of the weight parameter.
[0179] Step 804: Search for a threshold within the coarse-grained outlier interval to divide the coarse-grained outlier interval into a fine-grained outlier interval and a retention interval, and the distance between the fine-grained outlier interval and the retention interval is as large as possible, and the numerical distribution within the retention interval is as close as possible.
[0180] Specifically, it can be to traverse one by one the values of the weight parameters included in the coarse-grained outlier interval, and then determine that the value of one of the weight parameters is the searched threshold, so as to divide the coarse-grained outlier interval into a fine-grained outlier interval and a retention interval based on the searched threshold. For example, within the coarse-grained outlier interval, the range where the value of the weight parameter less than or equal to the searched threshold is located is the retention interval, and the range where the value of the weight parameter greater than the searched threshold is located is the fine-grained outlier interval.
[0181] To maximize the distance between the fine-grained outlier interval and the retention interval and make the numerical distribution within the retention interval as close as possible, the following formula can be set to search for the corresponding threshold. Specifically, assume the fine-grained outlier interval is O outlier , and the retention interval is O reserved . Then, the distance between the fine-grained outlier interval O outlier and the retention interval O reserved , and the numerical distribution within the retention interval O reserved satisfy the following formula.
[0182] M intra = var(O reserved )
[0183] M inter = (min(O outlier ) - max(O rcserved )) 2
[0184] M = M inter - λ 2 M intra
[0185] Among them, M intra represents the variance of the values within the retention interval O reserved ; M inter represents the distance between the fine-grained outlier interval O outlier and the retention interval O reserved ; λ 2 represents a hyperparameter, such as 0.1. By traversing each value in the coarse-grained outlier interval, the value that maximizes M is determined as the above-mentioned threshold, and then the final fine-grained outlier interval O outlier and the retention interval O reserved are determined.
[0186] Step 805: Set the values of the weight parameters within the fine-grained outlier interval to 0.
[0187] After determining the fine-grained outlier interval, the values of each weight parameter within the fine-grained outlier interval can be set to 0, so as to remove the outlier values of the weight parameters.
[0188] The above introduces the process of determining the outlier values of the weight parameters and processing the outlier values of the weight parameters. Similarly, before performing quantization on the activation values in the network structure, it is also possible to determine the outlier values of the activation values and process the outlier values of the activation values.
[0189] Exemplarily, during each round of quantization, at least one outlier can be selected from multiple activation values within the same neural network layer of the network structure. Herein, at least one outlier refers to one or more values among the multiple activation values that are significantly different from other values.
[0190] Then, the activation values corresponding to at least one outlier are scaled according to a target ratio, so that the scaled activation values are the same as the largest activation value among the multiple activation values excluding the outliers. It should be noted that for activation values with different numerical values, the target ratios for scaling are also different. For example, assume that there are currently 1000 activation values, among which 998 activation values are distributed within the range of [0, 10], and the largest activation value among these 998 activation values is 10; while the other two activation values belonging to the outliers are 80 and 100 respectively. In this way, for the activation value with a value of 80, this activation value can be scaled to 10 according to a ratio of 8:1; for the activation value with a value of 100, this activation value can be scaled to 10 according to a ratio of 10:1.
[0191] Exemplarily, please refer to Figure 9 and Figure 10 , Figure 9 which is a comparative schematic diagram of the distribution of activation values provided by an embodiment of the present application; Figure 10 which is another comparative schematic diagram of the distribution of activation values provided by an embodiment of the present application. As Figure 9 shown, before outlier removal, the distribution range of the activation values is relatively large, while after scaling the outliers, the distribution range of the activation values becomes smaller. As Figure 10 shown, before outlier removal, the distribution range of the activation values is [0, 280], and after outlier removal, the distribution range of the activation values becomes [0, 40], effectively narrowing the distribution range of the activation values, which is convenient for subsequently determining the quantization parameters of the activation values.
[0192] After determining the outliers corresponding to the activation values, different from setting the outliers of the weight parameters to a preset threshold, in this embodiment, the activation values are scaled. It can be understood that since the activation values are often the outputs of the neural network layer, if the activation values are directly set to a preset threshold (such as 0), it will directly have a greater impact on the output of the neural network layer and easily affect the accuracy of the model. Therefore, in this embodiment, the outliers of the activation values are scaled. For the outliers of the weight parameters, since the weight parameters indirectly affect the output of the neural network layer, setting the weight parameters to a preset threshold will not have too much impact on the accuracy of the model and can effectively reduce the difficulty of learning the quantization parameters.
[0193] In this solution, by determining the outliers in the activation values and performing scaling on the activation values belonging to the outliers, the distribution range of the activation values can be narrowed, thereby reducing the difficulty of learning quantization parameters and facilitating learning better quantization parameters, and improving the accuracy of the target model after quantization.
[0194] It should be noted that the method for determining outliers based on the values of multiple activation values in this embodiment may be similar to the method for determining outliers based on the values of multiple weight parameters described above. For details, please refer to the above embodiments and will not be elaborated here.
[0195] The method provided in the embodiments of the present application has been introduced in detail above. Next, the device provided in the embodiments of the present application for executing the above method will be introduced.
[0196] Please refer to Figure 11 , Figure 11 , which is a schematic structural diagram of a model quantization device provided in an embodiment of the present application. As Figure 11 shown, the model quantization device provided in the embodiments of the present application includes: an acquisition module 1101, configured to acquire a target model, where the target model includes a plurality of network structures connected in sequence; a processing module 1102, configured to perform multiple rounds of quantization processes to obtain a quantized target model, where the quantized target model includes quantized multiple network structures, and the multiple rounds of quantization processes are sequentially performed based on the connection order of the multiple network structures; wherein, in the multiple rounds of quantization processes, each round of quantization process quantizes at least two consecutive network structures among the multiple network structures, and each round of quantization process quantizes only part of the network structures among the multiple network structures. For any two consecutive quantization processes in the multiple rounds of quantization processes, there is at least one repeated network structure between the network structures quantized in the latter round of quantization process and the network structures quantized in the former round of quantization process.
[0197] In a possible implementation manner, the first N network structures quantized in the latter round of quantization process are the same N network structures as the last N network structures quantized in the former round of quantization process, and N is an integer greater than or equal to 1.
[0198] In a possible implementation manner, the processing module 1102 is further configured to: input the first input data into at least two network structures to obtain first output data, where the at least two network structures are the network structures quantized in the target round of quantization process, the first input data is the output of the network structures before the at least two network structures, and the target round of quantization process is one round of quantization process in the multiple rounds of quantization processes; input the second input data into the at least two network structures quantized based on the quantization parameters to obtain second output data, where the second input data is the output of the network structures before the at least two network structures after quantization; and update the quantization parameters of the at least two network structures based on the difference between the first output data and the second output data.
[0199] In a possible implementation, the quantization parameters include a quantization step size and a weight compensation matrix. The quantization step size is used to update the weight parameters in the network structure, and the weight compensation matrix is used to compensate the updated weight parameters. Among them, the weight compensation matrix is low-rank decomposed into multiple matrices, the total number of parameters of the multiple matrices is less than the number of parameters of the weight compensation matrix, and the update of the weight compensation matrix is achieved by updating the multiple matrices.
[0200] In a possible implementation, the difference is the weighted average between a first difference value and a second difference value. The first difference value is the Euclidean distance between the first output data and the second output data, and the second difference value is the relative entropy between the first output data and the second output data.
[0201] In a possible implementation, during each round of quantization process, the processing module 1102 is further configured to: select one or more outliers from the values of multiple weight parameters of the network structure; set the values of the weight parameters corresponding to the one or more outliers to a preset threshold.
[0202] In a possible implementation, the processing module 1102 is further configured to: determine a coarse-grained outlier interval based on the values of multiple weight parameters, where the values of the weight parameters located in the coarse-grained outlier interval are all greater than a first threshold or all less than a second threshold, and the first threshold or the second threshold is determined based on the distribution of the values of multiple weight parameters; determine a fine-grained outlier interval based on the coarse-grained outlier interval, the fine-grained outlier interval is within the range of the coarse-grained outlier interval, and the distribution of the values within the fine-grained outlier interval satisfies a preset condition, and the values of the weight parameters located in the fine-grained outlier interval belong to one or more outliers.
[0203] In a possible implementation, during the execution of multiple rounds of quantization process and before quantizing the activation values in the network structure, the processing module 1102 is further configured to: select at least one outlier from the values of multiple activation values of the network structure, and the values of the multiple activation values include at least one outlier; scale the values of the activation values corresponding to the at least one outlier according to a target ratio.
[0204] In a possible implementation, the target model is a large language model.
[0205] In a possible implementation, each of the multiple network structures includes one or more Transformer network blocks.
[0206] Please refer to Figure 12 , Figure 12Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. The electronic device 1200 may specifically be embodied as a mobile phone, a tablet computer, a laptop computer, a smart wearable device, a server, etc., which is not limited herein. Specifically, the electronic device 1200 includes: a receiver 1201, a transmitter 1202, a processor 1203, and a memory 1204 (where the number of processors 1203 in the electronic device 1200 may be one or more, Figure 12 and one processor is taken as an example here). Among them, the processor 1203 may include an application processor 12031 and a communication processor 12032. In some embodiments of the present application, the receiver 1201, the transmitter 1202, the processor 1203, and the memory 1204 may be connected through a bus or other means.
[0207] The memory 1204 may include a read-only memory and a random access memory, and provide instructions and data to the processor 1203. A part of the memory 1204 may further include a non-volatile random access memory (NVRAM). The memory 1204 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof. Among them, the operation instructions may include various operation instructions for implementing various operations.
[0208] The processor 1203 controls the operation of the electronic device. In a specific application, the various components of the electronic device are coupled together through a bus system. Among them, the bus system may further include a power bus, a control bus, a status signal bus, etc. in addition to the data bus. However, for the sake of clear illustration, all kinds of buses are referred to as the bus system in the figure.
[0209] The method disclosed in the above embodiments of the present application may be applied to the processor 1203 or implemented by the processor 1203. The processor 1203 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method may be completed by the integrated logic circuit in the hardware of the processor 1203 or the instructions in software form. The above-mentioned processor 1203 may be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and may further include an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate, or transistor logic devices, discrete hardware components.
[0210] The processor 1203 may implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed and completed by a hardware decoding processor, or may be executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 1204, and the processor 1203 reads the information in the memory 1204 and combines its hardware to complete the steps of the above method.
[0211] The receiver 1201 may be used to receive input digital or character information, and generate a signal input related to the relevant settings and function control of the electronic device. The transmitter 1202 may be used to output digital or character information through the first interface; the transmitter 1202 may also be used to send instructions to the disk group through the first interface to modify the data in the disk group; the transmitter 1202 may also include a display device such as a display screen.
[0212] The electronic device provided in the embodiments of the present application may specifically be a chip, and the chip includes: a processing unit and a communication unit. The processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin, or a circuit, etc. The processing unit may execute the computer execution instructions stored in the storage unit to enable the chip in the electronic device to execute the classification method of multimedia data described in the above embodiments, or to enable the chip in the training device to execute the model quantization method described in the above embodiments. Optionally, the storage unit is a storage unit inside the chip, such as a register, a cache, etc., and the storage unit may also be a storage unit outside the chip located in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0213] Specifically, please refer to Figure 13 , Figure 13 which is a schematic structural diagram of a chip provided in the embodiments of the present application. The chip may be represented as a neural network processor NPU 1300. The NPU 1300 is mounted on the main CPU (Host CPU) as a coprocessor, and tasks are assigned by the Host CPU. The core part of the NPU is the arithmetic circuit 1303, and the arithmetic circuit 1303 is controlled by the controller 1304 to extract matrix data from the memory and perform multiplication operations.
[0214] In some implementations, the arithmetic circuit 1303 internally includes multiple processing units (Process Engine, PE). In some implementations, the arithmetic circuit 1303 is a two-dimensional systolic array. The arithmetic circuit 1303 can also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1303 is a general matrix processor.
[0215] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit fetches the corresponding data of matrix B from the weight memory 1302 and caches it on each PE in the arithmetic circuit. The arithmetic circuit fetches the data of matrix A from the input memory 1301 and performs matrix operations with matrix B, and the partial results or final results of the obtained matrix are stored in the accumulator 1308.
[0216] The unified memory 1306 is used to store input data and output data. The weight data is directly transported through the Direct Memory Access Controller (DMAC) 1305 and is carried to the weight memory 1302. The input data is also carried to the unified memory 1306 through the DMAC.
[0217] The BIU is the Bus Interface Unit, that is, the bus interface unit 1310, which is used for the interaction between the AXI bus and the DMAC and the Instruction Fetch Buffer (IFB) 1309.
[0218] The bus interface unit 1310 (Bus Interface Unit, BIU) is used for the instruction fetch memory 1309 to obtain instructions from the external memory, and is also used for the storage unit access controller 1305 to obtain the original data of the input matrix A or the weight matrix B from the external memory.
[0219] The DMAC is mainly used to transport the input data in the external memory DDR to the unified memory 1306, or transport the weight data to the weight memory 1302, or transport the input data to the input memory 1301.
[0220] The vector calculation unit 1307 includes multiple arithmetic processing units, and in case of need, further processes the output of the arithmetic circuit 1303, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolution / full connection layer network calculations in neural networks, such as Batch Normalization, pixel-level summation, upsampling of the feature plane, etc.
[0221] In some implementations, the vector computing unit 1307 can store the processed output vectors into the unified memory 1306. For example, the vector computing unit 1307 can apply a linear function; or, a non-linear function to the output of the arithmetic circuit 1303, such as performing linear interpolation on the feature planes extracted by the convolutional layer, or, for another example, vectors of accumulated values, to generate activation values. In some implementations, the vector computing unit 1307 generates normalized values, pixel-level summation values, or both. In some implementations, the processed output vectors can be used as activation inputs to the arithmetic circuit 1303, such as for use in subsequent layers in a neural network.
[0222] The instruction fetch buffer 1309 connected to the controller 1304 is used to store the instructions used by the controller 1304;
[0223] The unified memory 1306, the input memory 1301, the weight memory 1302, and the instruction fetch buffer 1309 are all On-Chip memories. The external memory is private to this NPU hardware architecture.
[0224] Wherein, the processor mentioned anywhere above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above programs.
[0225] Reference can be made to Figure 14 , Figure 14 which is a schematic structural view of a computer-readable storage medium provided by an embodiment of the present application. The present application also provides a computer-readable storage medium. In some embodiments, the above Figure 2 disclosed method can be implemented as computer program instructions encoded in a computer-readable storage medium in a machine-readable format or encoded on other non-transitory media or articles.
[0226] Figure 14 Schematically shows a conceptual partial view of an example computer-readable storage medium arranged according to at least some of the embodiments shown here. The example computer-readable storage medium includes a computer program for executing a computer process on a computing device.
[0227] In one embodiment, the computer-readable storage medium 1400 is provided using a signal-bearing medium 1401. The signal-bearing medium 1401 can include one or more program instructions 1402, which when run by one or more processors can provide the functions or partial functions described above for Figure 2 description.
[0228] In some examples, the signal-bearing medium 1401 can include a computer-readable medium 1403, such as but not limited to, a hard disk drive, a compact disc (CD), a digital video disc (DVD), a digital tape, a memory, a ROM, or a RAM, and so on.
[0229] In some embodiments, the signal-bearing medium 1401 can include a computer-recordable medium 1404, such as but not limited to, a memory, a read / write (R / W) CD, a R / W DVD, and so on. In some embodiments, the signal-bearing medium 1401 can include a communication medium 1405, such as but not limited to, digital and / or analog communication media (e.g., fiber optic cables, waveguides, wired communication links, wireless communication links, and so on). Thus, for example, the signal-bearing medium 1401 can be conveyed by a wireless form of the communication medium 1405 (e.g., a wireless communication medium compliant with the IEEE 802.X standard or other transmission protocols).
[0230] One or more program instructions 1402 can be, for example, computer-executable instructions or logic-implemented instructions. In some examples, a computing device of the computing device can be configured to provide various operations, functions, or actions in response to the program instructions 1402 communicated to the computing device through one or more of the computer-readable medium 1403, the computer-recordable medium 1404, and / or the communication medium 1405.
[0231] It should be further noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the accompanying drawings of the device embodiments provided in this application, the connection relationships between the modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines.
[0232] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits, etc. However, for this application, software program implementation is a better embodiment in more cases. Based on such an understanding, the technical solution of this application, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, etc., and includes several instructions to enable a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods of various embodiments of this application.
[0233] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0234] A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of this application are generated in whole or in part. The computer can be a general computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training device, or data center to another website, computer, training device, or data center in a wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
Claims
1. A model quantization method, characterized in that, it includes: obtaining a target model, where the target model includes a plurality of network structures connected in sequence; performing multiple rounds of quantization processes to obtain a quantized target model, where the quantized target model includes the quantized plurality of network structures, and the multiple rounds of quantization processes are sequentially performed based on the connection order of the plurality of network structures; wherein, in the multiple rounds of quantization processes, each round of quantization process quantizes at least two consecutive network structures among the plurality of network structures, and each round of quantization process quantizes a part of the network structures among the plurality of network structures. For any two consecutive quantization processes in the multiple rounds of quantization processes, there is at least one repeated network structure between the network structures quantized by the latter round of quantization process and the network structures quantized by the former round of quantization process.
2. The method according to claim 1, characterized in that, the first N network structures quantized by the latter round of quantization process are the same N network structures as the last N network structures quantized by the former round of quantization process, and N is an integer greater than or equal to 1.
3. The method according to claim 1 or 2, characterized in that, the target round quantization process in the multiple rounds of quantization processes includes: inputting first input data into at least two network structures to obtain first output data, where the at least two network structures are the network structures quantized by the target round quantization process, the first input data is the output of the network structure before the at least two network structures, and the target round quantization process is one round of quantization process in the multiple rounds of quantization processes; inputting second input data into the at least two network structures quantized based on quantization parameters to obtain second output data, where the second input data is the output after quantization of the network structure before the at least two network structures; updating the quantization parameters of the at least two network structures based on the difference between the first output data and the second output data.
4. The method according to claim 3, characterized in that, the quantization parameters include a quantization step size and a weight compensation matrix, the quantization step size is used to update the weight parameters in the network structure, and the weight compensation matrix is used to compensate the updated weight parameters; wherein, the weight compensation matrix is low-rank decomposed into a plurality of matrices, the total number of parameters of the plurality of matrices is less than the number of parameters of the weight compensation matrix, and the update of the weight compensation matrix is achieved by updating the plurality of matrices.
5. The method according to claim 3 or 4, characterized in that, the difference is the weighted average between a first difference value and a second difference value, the first difference value is the Euclidean distance between the first output data and the second output data, and the second difference value is the relative entropy between the first output data and the second output data.
6. The method according to any one of claims 1-5, characterized in that, during the execution of each round of quantization process, the method further includes: selecting one or more outliers from the values of a plurality of weight parameters of the network structure; Set the value of the weight parameter corresponding to the one or more outliers to a preset threshold.
7. The method according to claim 6, wherein, selecting one or more outliers from the values of multiple weight parameters of the network structure includes: Based on the values of the multiple weight parameters, determining a coarse-grained outlier interval, wherein the values of the weight parameters located in the coarse-grained outlier interval are all greater than a first threshold or all less than a second threshold, and the first threshold or the second threshold is determined based on the distribution of the values of the multiple weight parameters; Based on the coarse-grained outlier interval, determining a fine-grained outlier interval, the fine-grained outlier interval is within the range of the coarse-grained outlier interval, and the distribution of the values within the fine-grained outlier interval satisfies a preset condition, and the values of the weight parameters located in the fine-grained outlier interval belong to the one or more outliers.
8. The method according to any one of claims 1-7, wherein, During each round of quantization process, the method further includes: Selecting at least one outlier from the values of multiple activation values in the network structure; Scaling the value of the activation value corresponding to the at least one outlier according to a target ratio.
9. The method according to any one of claims 1-8, wherein, The target model is a large language model.
10. The method according to claim 9, wherein, Each of the multiple network structures includes one or more Transformer network blocks.
11. A model quantization device, wherein, includes: An acquisition module for acquiring a target model, the target model includes a plurality of network structures connected in sequence; A processing module for performing multiple rounds of quantization process to obtain a quantized target model, the quantized target model includes the quantized multiple network structures, and the multiple rounds of quantization process are sequentially performed based on the connection order of the multiple network structures; Wherein, during the multiple rounds of quantization process, each round of quantization process performs quantization on at least two consecutive network structures among the multiple network structures, and each round of quantization process performs quantization on a part of the network structures among the multiple network structures. For any two consecutive quantization processes among the multiple rounds of quantization process, there is at least one repeated network structure between the network structures quantized by the latter round of quantization process and the network structures quantized by the former round of quantization process.
12. The device according to claim 11, wherein, The first N network structures quantized by the latter round of quantization process are the same N network structures as the last N network structures quantized by the former round of quantization process, and N is an integer greater than or equal to 1.
13. The device according to claim 11 or 12, wherein, The processing module is further used for: Input the first input data into at least two network structures to obtain first output data. The at least two network structures are the network structures quantized in the target round quantization process. The first input data is the output of the network structure before the at least two network structures. The target round quantization process is one round quantization process in the multi-round quantization process. Input the second input data into the at least two network structures quantized based on quantization parameters to obtain second output data. The second input data is the output after quantization of the network structure before the at least two network structures. Update the quantization parameters of the at least two network structures based on the difference between the first output data and the second output data.
14. The apparatus according to claim 13, wherein, the quantization parameters include a quantization step size and a weight compensation matrix. The quantization step size is used to update the weight parameters in the network structure, and the weight compensation matrix is used to compensate the updated weight parameters; wherein, the weight compensation matrix is low-rank decomposed into multiple matrices, the total number of parameters of the multiple matrices is less than the number of parameters of the weight compensation matrix, and the update of the weight compensation matrix is achieved by updating the multiple matrices.
15. The apparatus according to claim 13 or 14, wherein, the difference is the weighted average between a first difference value and a second difference value. The first difference value is the Euclidean distance between the first output data and the second output data, and the second difference value is the weighted average of the relative entropy between the first output data and the second output data.
16. The apparatus according to any one of claims 11-15, wherein, during each round of quantization process, the processing module is further configured to: select one or more outliers from the values of multiple weight parameters of the network structure; set the values of the weight parameters corresponding to the one or more outliers to a preset threshold.
17. The apparatus according to claim 16, wherein, the processing module is further configured to: determine a coarse-grained outlier interval based on the values of the multiple weight parameters, where the values of the weight parameters located in the coarse-grained outlier interval are all greater than a first threshold or all less than a second threshold, and the first threshold or the second threshold is determined based on the distribution of the values of the multiple weight parameters; determine a fine-grained outlier interval based on the coarse-grained outlier interval. The fine-grained outlier interval is within the range of the coarse-grained outlier interval, and the distribution of the values within the fine-grained outlier interval satisfies a preset condition. The values of the weight parameters located in the fine-grained outlier interval belong to the one or more outliers.
18. The apparatus according to any one of claims 11-17, wherein, during each round of quantization process, the processing module is further configured to: select at least one outlier from the values of multiple activation values of the network structure, and the values of the multiple activation values include the at least one outlier; scale the values of the activation values corresponding to the at least one outlier according to a target ratio.
19. The device according to any one of claims 11 - 18, wherein, the target model is a large language model.
20. The device according to claim 19, wherein, each of the plurality of network structures includes one or more Transformer network blocks.
21. A model quantization device, wherein, it includes a memory and a processor; the memory stores code, and the processor is configured to execute the code, and when the code is executed, the device executes the method according to any one of claims 1 to 10.
22. A computer storage medium, wherein, the computer storage medium stores instructions, and when the instructions are executed by a computer, the computer implements the method according to any one of claims 1 to 10.
Citation Information
Cited By
Model obtaining method and device and electronic equipment
CN120975164A
KV cache data quantification device and method
CN121031682A
Model quantization method and related apparatus
EP4797159A1
Model quantization method and related apparatus
WO2025112718A1