Model quantization method and related apparatus

By dividing the large language model into multiple network structures and performing multiple rounds of quantization processes, the problem of difficult to quantify models with huge parameters is solved, and the smooth deployment of the model on the processing equipment is achieved.

WO2025112718A1PCT designated stage expired Publication Date: 2025-06-05HUAWEI TECH CO LTD

Patent Information

Application Number
PCT/CN2024/114829
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-29
Filing Date
2024-08-27
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

The prior art is difficult to quantify large language models with huge parameters, making it difficult to deploy on processing equipment.

Method used

By dividing the target model into multiple network structures connected in sequence and performing multiple rounds of quantization processes, each round of quantization process performs quantization on at least two continuous network structures in multiple network structures to ensure the continuity and dependence of the quantization process.

Benefits of technology

It effectively reduces the amount of parameters loaded onto the processing device in a single time, ensures that the quantization process of the target model can be performed smoothly, and ensures the quantized model accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024114829_05062025_PF_FP_ABST
    Figure CN2024114829_05062025_PF_FP_ABST
Patent Text Reader

Abstract

A model quantization method, applied to quantization of a model in the technical field of artificial intelligence. In the method, a target model is divided into a plurality of parts on the basis of the connection relationship of network structures in the target model, and multiple rounds of quantization processes are executed on the target model, the multiple rounds of quantization processes being sequentially executed on the basis of the network structures of the model, and each round of quantization process being to quantize some of the network structures in the target model, thereby reducing the amount of parameters loaded onto a processing device a single time, and ensuring that the quantization processes of the target model can be smoothly executed. In the quantization processes, there are partially overlapping network structures in any two rounds of consecutive quantization processes, and the dependency relationship between the network structures is fully considered, ensuring that a joint optimization relationship can be established between the network structures in the multiple rounds of quantization processes, avoiding the situation in which global optimization cannot be achieved due to different network structures being independently quantized in each round of quantization process, and ensuring the precision of the finally quantized target model.
Need to check novelty before this filing date? Find Prior Art

Description

A model quantization method and related device

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on November 29, 2023, with application number 202311626819.4 and application name “A model quantization method and related device”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of artificial intelligence (AI) technology, and in particular to a model quantization method and related devices. Background Art

[0003] With the continuous development of AI, neural network models are widely used in various fields, achieving far greater results than ever before. At the same time, the number of parameters in neural network models is increasing, resulting in the need for large amounts of memory and computing resources when running neural network models, which seriously restricts their deployment in various scenarios.

[0004] To address the issue of neural network model deployment being difficult due to the large number of parameters, the concept of model quantization has been proposed in related technologies. Model quantization is essentially a compression technology for neural network models. Specifically, it uses a lower bit width to represent the weight parameters and feature data in the neural network model, thereby saving memory space and reducing the amount of computation. Generally speaking, related technologies load the entire neural network model onto a processing device, and then quantize the entire neural network model to ensure the accuracy of the quantized neural network model.

[0005] However, with the rise of large language models, large language models with huge parameter amounts cannot be loaded onto processing devices all at once. Therefore, the model quantization methods in related technologies are difficult to perform quantization on models with huge parameter amounts, such as large language models.

[0006] Summary of the Invention

[0007] The present application provides a model quantization method that can perform quantization on a model with a large number of parameters and ensure the accuracy of the quantized model.

[0008] In a first aspect, the present application provides a model quantization method for performing quantization on models with a large number of parameters in the field of AI. The method comprises: obtaining a target model, wherein the target model is a neural network model that needs to be quantized, and the target model is a model trained based on training data in a training set. Generally, the target model is a neural network model with a large number of parameters, and conventional AI processing equipment has difficulty loading the entire target model at once to perform the quantization process.

[0009] After obtaining the target model, it can be divided into multiple sequentially connected network structures. That is, the target model itself consists of multiple sequentially connected network structures. The subsequent quantization process is performed by dividing the target model into multiple parts, each corresponding to a network structure. Among the multiple network structures obtained by division, the output of the previous network structure serves as the input of the next network structure, that is, there is an input-output dependency relationship between the previous and next network structures.

[0010] Finally, multiple rounds of quantization are performed to obtain a quantized target model, where the quantized target model includes multiple quantized network structures. The multiple rounds of quantization are performed sequentially based on the connection order of the multiple network structures. Moreover, in the multiple rounds of quantization, each round of quantization is performed on at least two consecutive network structures among the multiple network structures, and each round of quantization is performed on part of the network structures among the multiple network structures. In addition, for any two consecutive rounds of quantization in the multiple rounds of quantization, there is at least one repeated network structure between the network structure quantized by the latter round of quantization and the network structure quantized by the previous round of quantization. It should be noted that for the network structure that is repeatedly quantized in the two previous and next rounds of quantization (i.e., the at least one repeated network structure mentioned above), the latter round of quantization is actually performed on the basis of the previous round of quantization, i.e., the quantization parameter used when the latter round of quantization just starts quantization is the quantization parameter determined in the previous round of quantization.

[0011] In this solution, the target model is divided into multiple parts based on the connection relationship of the network structure in the target model, and multiple rounds of quantization are performed on the target model. The multiple rounds of quantization are performed sequentially based on the network structure of the model. Each round of quantization quantizes a part of the network structure in the target model, thereby reducing the number of parameters loaded onto the processing device at a time and ensuring that the quantization process of the target model can be executed smoothly. In addition, during the quantization process, any two consecutive rounds of quantization have partially overlapping network structures, fully considering the dependency between network structures, ensuring that a joint optimization relationship can be established between network structures during the multiple rounds of quantization, effectively avoiding the situation where each round of quantization independently quantizes different network structures and cannot achieve the global optimum, thereby ensuring the accuracy of the target model obtained by the final quantization.

[0012] In one possible implementation, the first N network structures quantized in the subsequent quantization process are the same N network structures as the last N network structures quantized in the previous quantization process, where N is an integer greater than or equal to 1. For example, assuming that both the previous quantization process and the subsequent quantization process quantize M network structures, then the last N network structures among the M network structures quantized in the previous quantization process are actually the same N network structures as the first N network structures among the M network structures quantized in the subsequent quantization process, that is, there are N overlapping network structures in the previous and next quantization processes.

[0013] In this scheme, in the two rounds of quantization, there will be a part of overlapping network structures that will be jointly quantized with another part of the network structure in the previous round of quantization, and will be jointly quantized with another part of the network structure in the next round of quantization, thereby establishing a connection relationship between the quantization processes of different rounds, effectively avoiding the situation where each round of quantization process independently quantizes different network structures and cannot reach the global optimality.

[0014] In one possible implementation, any quantization round other than the first quantization round in the multi-round quantization process is considered a target quantization round. The target quantization round specifically includes: first, inputting first input data into at least two network structures to obtain first output data, wherein the at least two network structures are network structures quantized by the target quantization round, and the first input data is the output of the network structure before the at least two network structures. Since the current at least two network structures are not located at the very beginning of the target model, the input of the at least two network structures is actually the feature data (i.e., the first input data) output by the previous network structure.

[0015] Then, the second input data is input into at least two network structures after quantization based on the quantization parameter to obtain second output data, where the second input data is the output of the network structure before the at least two network structures after quantization.

[0016] Secondly, based on the difference between the first output data and the second output data, quantization parameters of the at least two network structures are updated. For example, the quantization parameters of the at least two network structures are updated by gradient backpropagation, so that when the at least two network structures are quantized based on the updated quantization parameters, the outputs of the at least two network structures are closer to the outputs before quantization.

[0017] In this scheme, in the process of determining the quantization parameters of the network structure, the output of the original network structure is used as a supervisory signal, and the difference is constructed with the output of the network structure after quantization based on the quantization parameters. Then, the quantization parameters are updated with the goal of minimizing the output difference, so that accurate quantization parameters can be obtained, ensuring the accuracy of the network structure after quantization based on the quantization parameters.

[0018] In a possible implementation, the quantization parameters of the weight parameters include a quantization step size and a weight compensation matrix. The quantization step size is used to update the weight parameters in the network structure, and the weight compensation matrix is ​​used to compensate for the updated weight parameters.

[0019] In addition, the weight compensation matrix is ​​decomposed into multiple matrices using a low-rank decomposition. The sum of the parameters of the multiple matrices is smaller than the parameter number of the weight compensation matrix, and the weight compensation matrix is ​​updated by updating multiple matrices. Among them, low-rank decomposition is a method of decomposing a high-dimensional matrix into a low-rank matrix, which can effectively reduce the dimensionality of the data and reduce redundant information.

[0020] In this solution, by performing a low-rank decomposition of a weight compensation matrix with a large number of parameters into multiple matrices with smaller parameters, multiple matrices with smaller parameters can be learned during the quantization process, eliminating the need to learn a weight compensation matrix with a large number of parameters. This significantly reduces the number of parameters to be learned during the quantization process, as well as the number of iterations and resource usage. Especially for large language models with a large number of parameters, performing a low-rank decomposition can decompose the weight compensation matrix with a large number of parameters into multiple matrices with very small parameters, significantly reducing the cost of learning the weight compensation matrix.

[0021] In one possible implementation, the difference between the first output data and the second output data is a weighted average of the first difference value and the second difference value, the first difference value is the Euclidean distance between the first output data and the second output data, and the second difference value is the relative entropy between the first output data and the second output data.

[0022] In this scheme, by constructing the differences between the outputs of the network structure based on the Euclidean distance between the outputs and the weighted average of the relative entropy, it is possible to effectively deal with outliers in the feature data and improve the robustness of the subsequent optimization of the network structure based on the differences.

[0023] In one possible implementation, during each round of quantization, the model quantization method further includes selecting one or more outliers from the values ​​of multiple weight parameters of the network structure. That is, one or more outliers are determined from the values ​​of the multiple weight parameters. The one or more outliers are values ​​in the multiple weight parameters that differ significantly from other values.

[0024] Then, the values ​​of the weight parameters corresponding to the one or more outliers are set to a preset threshold value. The preset threshold value can be a preset value determined based on the distribution range of the multiple weight parameters, such as 0 or any value within the distribution range of the other weight parameters excluding the outliers in the multiple weight parameters.

[0025] In this solution, by identifying outliers in the weight parameters and adjusting the values ​​of the weight parameters belonging to the outliers to a preset threshold, the distribution range of the weight parameters can be narrowed, thereby reducing the difficulty of learning the quantization parameters and facilitating the learning of better quantization parameters, thereby improving the accuracy of the quantized target model. In addition, since outliers in the activation values ​​are also caused to a certain extent by outliers in the weight parameters, adjusting the outliers in the weight parameters can also improve the outliers in the activation values, which is conducive to learning better quantization parameters for the activation values.

[0026] In one possible implementation, one or more outliers are determined based on the values ​​of multiple weight parameters in the network structure. Specifically, the method may include: first determining a coarse-grained outlier interval based on the values ​​of multiple weight parameters, wherein the values ​​of the weight parameters in the coarse-grained outlier interval are all greater than a first threshold or are all less than a second threshold, and the first threshold and the second threshold are determined based on the distribution of the values ​​of the multiple weight parameters. That is, first determine the first threshold or the second threshold based on the distribution of the values ​​of the multiple weight parameters, and determine the range of the values ​​of the weight parameters greater than the first threshold as the coarse-grained outlier interval, or determine the range of the values ​​of the weight parameters less than the second threshold as the coarse-grained outlier interval. Specifically, the first threshold can be a positive value, and the second threshold can be a negative value. In general, the values ​​in the coarse-grained outlier interval are values ​​far from 0.

[0027] Then, within the coarse-grained outlier interval, a fine-grained outlier interval is further determined. The fine-grained outlier interval is located within the coarse-grained outlier interval, and the distribution of values ​​within the fine-grained outlier interval meets a preset condition. After the fine-grained outlier interval is determined, it can be determined whether the value of the weight parameter within the fine-grained outlier interval belongs to one or more of the aforementioned outliers.

[0028] In this scheme, by first determining a threshold based on the distribution of the values ​​of the weight parameter, and then determining the coarse-grained outlier interval based on the threshold, and then further searching for fine-grained intervals that meet the preset conditions within the coarse-grained outlier interval, the efficiency of determining outliers can be effectively improved.

[0029] In one possible implementation, during each round of quantization, the above-mentioned model quantization method further includes: selecting at least one outlier from multiple activation values ​​of the network structure; and scaling the activation value corresponding to the at least one outlier according to a target ratio.

[0030] In this scheme, by identifying outliers in the activation values ​​and performing scaling on the activation values ​​belonging to the outliers, the distribution range of the activation values ​​can be narrowed, thereby reducing the difficulty of learning quantization parameters and facilitating the learning of better quantization parameters, thereby improving the accuracy of the quantized target model.

[0031] In a possible implementation, the target model is a large language model.

[0032] In a possible implementation, each of the multiple network structures of the target model includes one or more Transformer network blocks.

[0033] The second aspect of the present application provides a model quantization device, including: an acquisition module, used to acquire a target model, the target model includes multiple network structures connected in sequence; a processing module, used to perform multiple rounds of quantization process to obtain a quantized target model, the quantized target model includes multiple quantized network structures, and the multiple rounds of quantization process are performed in sequence based on the connection order of the multiple network structures; wherein, in the multiple rounds of quantization process, each round of quantization process is performed on at least two consecutive network structures in the multiple network structures, and each round of quantization process is performed on part of the network structures in the multiple network structures. For any two consecutive rounds of quantization process in the multiple rounds of quantization process, there is at least one repeated network structure between the network structure quantized by the latter round of quantization process and the network structure quantized by the previous round of quantization process.

[0034] In a possible implementation, the first N network structures quantized in the next round of quantization process and the last N network structures quantized in the previous round of quantization process are the same N network structures, where N is an integer greater than or equal to 1.

[0035] In one possible implementation, the processing module is further used to: input first input data into at least two network structures to obtain first output data, where the at least two network structures are network structures quantized by a target round quantization process, the first input data are outputs of the network structure before at least two network structures, and the target round quantization process is one round of quantization process in a multi-round quantization process; input second input data into at least two network structures after quantization based on quantization parameters to obtain second output data, where the second input data are outputs of the network structure before at least two network structures after quantization; and update the quantization parameters of the at least two network structures based on the difference between the first output data and the second output data.

[0036] In one possible implementation, the quantization parameters include a quantization step and a weight compensation matrix, the quantization step is used to update the weight parameters in the network structure, and the weight compensation matrix is ​​used to compensate for the updated weight parameters; wherein, the weight compensation matrix is ​​decomposed into multiple matrices by low rank, the sum of the parameter quantities of the multiple matrices is less than the parameter quantity of the weight compensation matrix, and the update of the weight compensation matrix is ​​achieved by updating multiple matrices.

[0037] In one possible implementation, the difference is a weighted average of the first difference value and the second difference value, the first difference value is the Euclidean distance between the first output data and the second output data, and the second difference value is the relative entropy between the first output data and the second output data.

[0038] In a possible implementation, during each round of quantization, the processing module is further used to: select one or more outliers from the values ​​of multiple weight parameters of the network structure; and set the values ​​of the weight parameters corresponding to the one or more outliers to a preset threshold.

[0039] In one possible implementation, the processing module is further used to: determine a coarse-grained outlier interval based on the values ​​of multiple weight parameters, wherein the values ​​of the weight parameters located in the coarse-grained outlier interval are all greater than a first threshold or are all less than a second threshold, and the first threshold or the second threshold is determined based on the distribution of the values ​​of the multiple weight parameters; determine a fine-grained outlier interval based on the coarse-grained outlier interval, wherein the fine-grained outlier interval is within the range of the coarse-grained outlier interval, and the distribution of values ​​within the fine-grained outlier interval meets preset conditions, and the values ​​of the weight parameters located in the fine-grained outlier interval belong to one or more outlier values.

[0040] In a possible implementation, during each round of quantization, the processing module is further configured to: select at least one outlier from multiple activation values ​​of the network structure; and scale the activation value corresponding to the at least one outlier according to a target ratio.

[0041] In one possible implementation, the target model is a large language model.

[0042] In a possible implementation, each of the multiple network structures includes one or more Transformer network blocks.

[0043] In a third aspect, the present application provides a model quantization device, which may include a processor coupled to a memory, wherein the memory stores program instructions. When the program instructions stored in the memory are executed by the processor, the method of the first aspect or any implementation of the first aspect is implemented. For details of the steps in each possible implementation of the first aspect executed by the processor, please refer to the first aspect and will not be repeated here.

[0044] In a fourth aspect, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer executes the method of any implementation of the first aspect.

[0045] A fifth aspect of the present application provides a circuit system, the circuit system including a processing circuit, and the processing circuit is configured to execute a method of any implementation manner of the above-mentioned first aspect.

[0046] In a sixth aspect, the present application provides a computer program product, which, when executed on a computer, enables the computer to execute the method of any implementation manner of the first aspect.

[0047] In a seventh aspect, the present application provides a chip system, which includes a processor for supporting an electronic device to implement the functions involved in any implementation of the first aspect, for example, processing the data and / or information involved in the above method. In one possible design, the chip system also includes a memory for storing program instructions and data necessary for the electronic device. The chip system can be composed of a chip or a chip and other discrete devices.

[0048] The beneficial effects of the second to seventh aspects mentioned above can be referred to the introduction of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] FIG1 is a schematic diagram of a system architecture 100 provided in an embodiment of the present application;

[0050] FIG2 is a schematic diagram of a flow chart of a model quantization method provided in an embodiment of the present application;

[0051] FIG3A is a schematic diagram of a multi-round quantization process for a target model according to an embodiment of the present application;

[0052] FIG3B is a schematic diagram of another embodiment of the present application for performing multiple rounds of quantization on a target model;

[0053] FIG3C is a schematic diagram of another embodiment of the present application providing a process of performing multiple rounds of quantization on a target model;

[0054] FIG4 is a schematic diagram of a process for performing quantization on a target model according to an embodiment of the present application;

[0055] FIG5 is a schematic diagram of performing low-rank decomposition on a weight compensation matrix according to an embodiment of the present application;

[0056] FIG6 is a comparative diagram of the distribution of weight parameters provided in an embodiment of the present application;

[0057] FIG7 is a comparative diagram of another distribution of weight parameters provided in an embodiment of the present application;

[0058] FIG8 is a schematic diagram of a process for performing preprocessing on weight parameters according to an embodiment of the present application;

[0059] FIG9 is a comparative schematic diagram of the distribution of activation values ​​provided in an embodiment of the present application;

[0060] FIG10 is a comparative schematic diagram of another distribution of activation values ​​provided in an embodiment of the present application;

[0061] FIG11 is a schematic structural diagram of a model quantization device provided in an embodiment of the present application;

[0062] FIG12 is a schematic structural diagram of an electronic device provided in an embodiment of the present application;

[0063] FIG13 is a schematic diagram of the structure of a chip provided in an embodiment of the present application;

[0064] FIG14 is a schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of the present application. DETAILED DESCRIPTION

[0065] In order to make the purpose, technical solutions and advantages of this application more clear, the embodiments of this application are described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only embodiments of a part of this application, rather than all embodiments. It is known to those skilled in the art that with the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0066] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the descriptions used in this way can be interchangeable where appropriate so that the embodiments can be implemented in a sequence other than that illustrated or described in this application. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or modules is not necessarily limited to those steps or modules clearly listed, but may include other steps or modules that are not clearly listed or that are inherent to these processes, methods, products or devices. The naming or numbering of steps in this application does not mean that the steps in the method flow must be executed in the time / logical sequence indicated by the naming or numbering. The named or numbered process steps can change the execution order according to the technical purpose to be achieved, as long as the same or similar technical effects can be achieved. The division of units in this application is a logical division. In actual application, there may be other division methods. For example, multiple units can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, and the indirect coupling or communication connection between units can be electrical or other similar forms, which are not limited in this application. Moreover, the units or sub-units described as separate components may or may not be physically separated, may or may not be physical units, or may be distributed into multiple circuit units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this application.

[0067] To facilitate understanding, some technical terms involved in the embodiments of this application are first introduced below.

[0068] (1) Model quantization

[0069] Model quantization is a compression technology for neural network models. Specifically, it uses a lower bit width to represent the weight parameters and feature data in the neural network model, thereby saving memory space and reducing the amount of computation. Generally speaking, model quantization refers to the process of mapping the parameters of the neural network model from single-precision floating point numbers (32-bit floating point numbers, FP32) to n bits. Simply put, it establishes a data mapping relationship between fixed-point numbers and floating-point numbers, achieving better results at the cost of minimal precision loss. For example, by mapping the parameters of the neural network model from FP32 to 8-bit signed integers (INT8), a 4x parameter compression can be achieved, which not only compresses memory but also enables faster computation, effectively improving model performance.

[0070] Generally speaking, model quantization is usually divided into two types: post-training quantization (PTQ) and quantization-aware training (QAT). PTQ is a process that only requires a small amount of unlabeled calibration data set to complete the quantization of a pre-trained neural network model. It is often used to compress neural network models that use a higher bit width to represent parameters. QAT, on the other hand, requires a complete data set to train the neural network model and simulates quantization operations during the training process so that the quantized model (i.e., the quantized model) can further converge to the optimal point. It is often used to recover accuracy after a large loss of accuracy in the quantized model.

[0071] Furthermore, during model quantization, the weight parameters and activation values ​​within the neural network model are generally quantized. Weight parameters refer to the parameters used to process input data within each neural network layer of the neural network model and are variables that can be adjusted during training. Activation values ​​typically refer to the feature data obtained after each neural network layer processes the input data. Since the output of the previous neural network layer in a neural network model serves as the input for the next layer, the activation value of the previous neural network layer typically serves as the input for the next layer. By quantizing the weight parameters and activation values, the number of parameters and computational complexity within the neural network model can be significantly reduced, thereby improving the model's operational performance.

[0072] (2) Neural Network

[0073] A neural network can be composed of neural units, which can be represented by x s (i.e. input data) and intercept 1 as input operation unit, the output of the operation unit can be:

[0074] Where, s = 1, 2, ... n, n is a natural number greater than 1, W s is x s The weight parameter of the neural unit, b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce nonlinear characteristics into the neural network to convert the input signal in the neural unit into the output signal. The output signal of the activation function can be used as the input of the next convolutional layer, and the activation function can be a sigmoid function. A neural network is a network formed by connecting multiple single neural units mentioned above, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be an area composed of several neural units.

[0075] (3) Deep Neural Network (DNN)

[0076] Deep neural networks, also known as multi-layer neural networks, can be understood as neural networks with many hidden layers. There is no special metric for "many" here. Based on the position of different layers in DNN, the neural network inside DNN can be divided into three categories: input layer, hidden layer, and output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the layers in between are all hidden layers. The layers are fully connected, that is, any neuron in the i-th layer must be connected to any neuron in the i+1-th layer. Although DNN looks complicated, the work of each layer is actually not complicated. Simply put, it is the following linear relationship expression: in, is the input vector, is the output vector, is the offset vector, W is the weight matrix (also called coefficient), and α() is the activation function. Each layer is just an input vector After such a simple operation, the output vector Since there are many DNN layers, the coefficient W and the offset vector The definition of these parameters in DNN is as follows: Take the coefficient W as an example: Assume that in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer number of the coefficient W, while the subscript corresponds to the output of the third layer index 2 and the input of the second layer index 4. In summary, the coefficient from the kth neuron in the L-1th layer to the jth neuron in the Lth layer is defined as It's important to note that the input layer has no W parameter. In deep neural networks, more hidden layers allow the network to better capture complex real-world situations. Theoretically, a model with more parameters has higher complexity and greater "capacity," meaning it can handle more complex learning tasks. Training a deep neural network is essentially the process of learning the weight matrix, with the ultimate goal of obtaining the weight matrices for all layers of a trained deep neural network (a weight matrix formed by the vectors W across many layers).

[0077] (4) Convolutional Neural Network (CNN)

[0078] A convolutional neural network is a deep neural network with a convolutional structure. It consists of a feature extractor consisting of a convolutional layer and a subsampling layer. This feature extractor can be thought of as a filter, and the convolution process can be thought of as convolving a trainable filter with an input feature map. A convolutional layer refers to the layer of neural units in a convolutional neural network that performs convolution processing on the input signal. In a convolutional layer of a convolutional neural network, a neural unit can only be connected to some of the neural units in adjacent layers. A convolutional layer typically contains several feature planes, each of which can be composed of a number of neural units arranged in a rectangular pattern. Neural units in the same feature plane share weights, which are referred to as convolution kernels.

[0079] Convolution kernels can be initialized as matrices of random size, and during the training process of the convolutional neural network, the convolution kernels can be learned to obtain reasonable weights. In addition, the direct benefit of shared weights is that they reduce the number of connections between the layers of the convolutional neural network, while also reducing the risk of overfitting.

[0080] (5) Attention Network

[0081] Attention networks are network models that utilize the attention mechanism to accelerate model training. Currently, typical attention networks include the Transformer. Models using the attention mechanism assign different weights to each part of the input sequence, thereby extracting more important features from the input sequence and ultimately achieving more accurate output.

[0082] The Transformer is a neural network architecture based on the self-attention mechanism. It typically consists of an encoder and a decoder, each of which is composed of multiple layers of stacked self-attention and feedforward neural networks. The self-attention mechanism allows the model to simultaneously consider relevant information at all positions when processing a sequence, thus addressing the problem of long-range dependencies.

[0083] (6) Large language model (LLM)

[0084] Large language models are deep learning models trained using large amounts of text data. They can generate natural language text or understand the meaning of text. Large language models can handle a variety of natural language tasks, such as text classification, question-answering, and conversation, and are an important path to artificial intelligence.

[0085] (7) Outliers

[0086] Outliers are also commonly called escape values. They refer to the situation where there is one or several values ​​in the data that are significantly different from other values. These values ​​that are significantly different from other values ​​are outliers.

[0087] (8)Quartile

[0088] Quartiles, also known as quartiles, are the three points in statistics where all values ​​are arranged from smallest to largest and divided into four equal parts. Quartiles divide all data into four equal parts using three points, with each part containing 25% of the data.

[0089] There are three quartiles. The first quartile is called the lower quartile, the second quartile is the median, and the third quartile is called the upper quartile, which are represented by Q1, Q2, and Q3 respectively.

[0090] The first quartile (Q1), also known as the "lower quartile", is equal to the 25th percentile of all values ​​in the sample arranged from small to large.

[0091] The second quartile (Q2), also known as the "median", is equal to the 50th percentile of all values ​​in the sample arranged from small to large.

[0092] The third quartile (Q3), also known as the "upper quartile", is equal to the 75th percentile of all the values ​​in the sample arranged from small to large.

[0093] The gap between the third quartile and the first quartile is also called the interquartile range (IQR).

[0094] (9) Euclidean distance

[0095] Euclidean distance, also known as Euclidean distance or L2 distance, is the most common distance metric, which measures the absolute distance between two points in multidimensional space.

[0096] (10) Relative entropy

[0097] Relative entropy, also known as the Kullback-Leibler divergence (KL divergence) or information divergence, is a measure of the asymmetry between two probability distributions. In information theory, relative entropy is equivalent to the difference in the Shannon entropy of two probability distributions.

[0098] Relative entropy is a loss function used in some optimization algorithms, such as the Expectation-Maximization algorithm (EM). In this case, one of the probability distributions involved in the calculation is the true distribution, and the other is the theoretical (fitted) distribution. The relative entropy represents the information loss incurred when fitting the true distribution to the theoretical distribution.

[0099] The applicant's research has revealed that model quantization techniques in related art typically load the entire neural network model onto a processing device and then quantize the entire neural network model to ensure the accuracy of the quantized neural network model. However, with the rise of large-scale models such as large language models, these models, with their enormous number of parameters, cannot be fully loaded onto a processing device at once. Therefore, model quantization methods in related art have difficulty quantizing models with large parameter counts, such as large language models.

[0100] Based on this, a model quantization method is provided in an embodiment of the present application. The target model is divided into multiple parts based on the connection relationship of the network structure in the target model, and multiple rounds of quantization processes are performed on the target model. The multiple rounds of quantization processes are performed in sequence based on the network structure of the model. Each round of quantization process quantizes a part of the network structure in the target model, thereby reducing the amount of parameters loaded onto the processing device at a single time, ensuring that the quantization process of the target model can be executed smoothly. Moreover, in the quantization process, any two consecutive rounds of quantization processes have partially overlapping network structures, fully considering the dependency relationship between the network structures, ensuring that a joint optimization relationship can be established between the network structures in the multiple rounds of quantization process, effectively avoiding the situation where each round of quantization process independently quantizes different network structures and cannot achieve the global optimum, thereby ensuring the accuracy of the target model obtained by the final quantization.

[0101] Please refer to Figure 1, which is a schematic diagram of a system architecture 100 provided in an embodiment of the present application. As shown in Figure 1, in the system architecture 100, the execution device 110 can be implemented by one or more servers. Optionally, the execution device 110 cooperates with other computing devices, such as data storage, routers, load balancers, and other devices; the execution device 110 can be deployed at a single physical site or distributed across multiple physical sites. The execution device 110 can use data in the data storage system 120, or call program code in the data storage system 120 to implement the model quantization method provided in an embodiment of the present application.

[0102] Users can operate their respective user devices (such as local device 101 and local device 102) to interact with execution device 110. Each local device can represent any computing device, such as a personal computer, a computer workstation, a smart phone, a tablet computer, a laptop computer, and a smart car.

[0103] Each user's local device can interact with the execution device 110 through a communication network of any communication mechanism / communication standard. The communication network can be a wide area network, a local area network, a point-to-point connection, etc., or any combination thereof.

[0104] In one implementation, execution device 110 is configured to implement the model quantization method provided in the embodiments of the present application to obtain a quantized model. Furthermore, when local device 101 or local device 102 needs to use the model for inference, execution device 110 processes user-provided data based on the quantized model and returns the corresponding processing results to local device 101 or local device 102.

[0105] In another implementation, the execution device 110 is used to implement the model quantization method provided in the embodiment of the present application, and send the obtained quantized model to the local device 101 and the local device 102. In this way, the local device 101 and the local device 102 can locally deploy the quantized model and perform data processing based on the quantized model.

[0106] In another implementation, one or more aspects of the execution device 110 can be implemented by each local device. For example, the local device 101 can provide local data or feedback calculation results to the execution device 110, or execute the model quantization method provided in the embodiment of the present application.

[0107] In general, the model quantization method provided in the embodiments of the present application can be applied to electronic devices, such as the aforementioned execution device 110, local device 101 or local device 102.

[0108] Please refer to Figure 2, which is a flow chart of a model quantization method provided in an embodiment of the present application. As shown in Figure 2, the model quantization method provided in an embodiment of the present application includes the following steps 201-203.

[0109] Step 201: Obtain a target model.

[0110] In this embodiment, the target model is a neural network model that needs to be quantized, and the target model is a model trained based on training data in a training set.

[0111] For example, the target model may be a large language model, a convolutional neural network, an attention network, or a recurrent neural network, and this embodiment does not specifically limit this. Furthermore, when the target model is a different type of neural network model, the target model may be applied to, for example, a natural language processing (NLP) task or a computer vision task, and this embodiment does not limit the tasks to which the target model is applied.

[0112] It should be noted that the target model is a neural network model with a large number of parameters, making it difficult for conventional AI processing devices to load the entire target model at once to perform the quantization process. For example, in the case of a large language model, the target model often includes hundreds of billions of parameters, which makes it difficult for existing graphics processing unit (GPU) clusters to load the entire target model at once.

[0113] Step 202: Divide the target model into multiple network structures connected in sequence.

[0114] In this embodiment, based on the network structure of the target model itself, the target model can be divided into multiple parts, each part corresponding to a network structure, thereby achieving the division of the target model into multiple network structures connected in sequence. Among them, in the multiple network structures obtained by division, the output of the previous network structure is the input of the next network structure, that is, there is an input-output dependency relationship between the previous and next network structures. In addition, the output of the previous network structure is usually feature data (i.e., a feature matrix), and the feature data output by the previous network structure may be one or more (for example, one or more feature matrices are output), which is not specifically limited here.

[0115] For example, if the target model is a large language model, it is typically composed of multiple connected Transformer blocks, such as 20-30 sequentially connected Transformer blocks. Therefore, when performing network structure partitioning on the target model, one or more Transformer blocks can be used as a single network structure, resulting in multiple sequentially connected network structures.

[0116] That is, for the target model, each of the multiple sequentially connected network structures obtained by partitioning the target model includes one or more transformer network blocks. For example, for a target model that includes 30 sequentially connected transformer network blocks, the target model can be divided into 30 network structures, each corresponding to a transformer network block.

[0117] Step 203 , performing multiple rounds of quantization process to obtain a quantized target model, where the quantized target model includes multiple quantized network structures.

[0118] In this embodiment, the multi-round quantization process is performed sequentially based on the connection order of the multiple network structures, that is, one round of quantization is performed each time. After the multi-round quantization process is performed sequentially, all network structures in the target model can be over-quantized, thereby obtaining a quantized target model. The quantized target model is actually composed of the multiple quantized network structures. That is, during the multi-round quantization process of the target model, quantization can be performed on the other network structures in sequence starting from the first network structure in the target model until the quantization of the last network structure in the target model is completed.

[0119] During the multiple rounds of quantization performed on the target model, each round of quantization is performed on at least two consecutive network structures among the multiple network structures. Each round of quantization is performed on a portion of the multiple network structures, and the network structures quantized in different rounds of quantization are not exactly the same. In addition, for any two consecutive rounds of quantization during the multiple rounds of quantization, there is at least one duplicate network structure between the network structure quantized in the subsequent round of quantization and the network structure quantized in the previous round of quantization.

[0120] In other words, parts of the network structure that have already been quantized in the previous round of quantization will be quantized again in the next round. This way, for both rounds of quantization, some overlapping network structures will be jointly quantized with another part of the network structure in the previous round, and then jointly quantized with yet another part of the network structure in the next round. This establishes a connection between the quantization processes of different rounds, effectively avoiding the situation where each round of quantization independently quantizes different network structures and fails to achieve the global optimality.

[0121] Exemplarily, for any two consecutive quantization processes in a multi-round quantization process, the first N network structures quantized by the latter quantization process and the last N network structures quantized by the former quantization process are the same N network structures, and N is an integer greater than or equal to 1. The total number of network structures quantized by the latter quantization process and the total number of network structures quantized by the former quantization process may be the same or different. For example, assuming that both the former quantization process and the latter quantization process quantize M network structures, then the last N network structures among the M network structures quantized by the former quantization process are actually the same N network structures as the first N network structures among the M network structures quantized by the latter quantization process, that is, there are N overlapping network structures in the former and latter quantization processes.

[0122] The number of network structures quantized during each round of quantization can be determined based on the capabilities of the processing device, for example, 2-6. Similarly, the number of network structures overlapped between the two rounds of quantization can also be determined based on the capabilities of the processing device and the accuracy of the model after quantization. If the processing device is more capable and the accuracy of the model after quantization is higher, the number of network structures overlapped between the two rounds of quantization can be set larger.

[0123] In one possible example, please refer to FIG3A , which is a schematic diagram of a multi-round quantization process performed on a target model according to an embodiment of the present application. In the example shown in FIG3A , when each round of quantization of the target model is performed on two consecutive network structures among multiple network structures, the first network structure quantized in the latter round of quantization is the same network structure as the last network structure quantized in the previous round of quantization.

[0124] As shown in Figure 3A, the target model consists of n sequentially connected network structures, namely, Network Structure 1 through Network Structure n. In the first round of quantization, Network Structure 1 and Network Structure 2 are quantized; in the second round, Network Structure 2 and Network Structure 3 are quantized. That is, after Network Structure 2 is quantized in conjunction with Network Structure 1 in the first round, it is then quantized in conjunction with Network Structure 3 in the second round. Similarly, in the third round, Network Structure 3 and Network Structure 4 are quantized; in the fourth round, Network Structure 4 and Network Structure 5 are quantized, and so on, until Network Structure n is quantized.

[0125] In other words, this embodiment actually performs network structure quantization using a sliding window approach. Each round of quantization actually quantizes multiple network structures within the sliding window. By sliding the window, different network structures are included within the window, triggering different rounds of quantization. Furthermore, adjacent sliding windows may contain overlapping network structures, thereby ensuring joint quantization of the network structures.

[0126] In another possible example, please refer to FIG3B , which is a schematic diagram of another embodiment of the present application for performing multiple rounds of quantization on a target model. In the example shown in FIG3B , when each round of quantization for the target model is performed on three consecutive network structures among multiple network structures, the first network structure quantized in the subsequent round of quantization is the same network structure as the last network structure quantized in the previous round of quantization.

[0127] As shown in Figure 3B, for the target model including network structures 1 to n, in the first round of quantization, network structures 1, 2, and 3 are quantized; in the second round of quantization, network structures 3, 4, and 5 are quantized. That is, after network structure 3 is quantized in conjunction with network structures 1 and 2 in the first round of quantization, network structure 3 is quantized in conjunction with network structures 4 and 5 in the second round of quantization.

[0128] In another possible example, please refer to Figure 3C, which is a schematic diagram of another embodiment of the present application for performing multiple rounds of quantization on a target model. As shown in Figure 3B, for a target model including network structure 1-network structure n. In the first round of quantization, network structure 1, network structure 2, network structure 3 and network structure 4 are quantized; in the second round of quantization, network structure 3, network structure 4, network structure 5 and network structure 6 are quantized, that is, after network structure 3 and network structure 4 are quantized in conjunction with network structure 1 and network structure 2 in the first round of quantization, they continue to be quantized in conjunction with network structure 5 and network structure 6 in the second round of quantization.

[0129] In general, for any two consecutive rounds of quantization in a multi-round quantization process, the number of network structures quantized in the previous round of quantization can be the same as or different from the number of network structures quantized in the next round of quantization, and the number of network structures repeatedly quantized in the two rounds of quantization can be one or more.

[0130] The above describes the process of performing multiple rounds of quantization on the target model to achieve quantization of the target model. For ease of understanding, the following details the specific process of each round of quantization performed on the target model.

[0131] It can be understood that in the multiple rounds of quantization performed on the target model, each round of quantization is actually to determine the quantization parameters corresponding to the network structure in the target model, so that the parameters in the network structure can be transformed based on the quantization parameters (for example, the parameters in the network structure are transformed from FP32 to INT8), and the network structure after parameter transformation (i.e., the quantized network structure) is obtained. When determining the quantization parameters corresponding to the network structure in the target model, it is necessary to ensure that the output of the network structure after parameter transformation based on the determined quantization parameters is as close as possible to the output of the original network structure, that is, the output of the quantized network structure is as close as possible to the output of the network structure before quantization.

[0132] Based on this, in this embodiment, the quantization parameters corresponding to the network structure are constrained by comparing the output difference between the network structure after adjustment based on the quantization parameters and the network structure before quantization during each round of quantization, so that the output difference between the network structure after adjustment based on the finally determined quantization parameters and the network structure before quantization is as small as possible.

[0133] Exemplarily, for the first round of quantization in a multi-round quantization process, the multiple network structures that need to be quantized in the first round of quantization are first determined. Then, calibration data (e.g., part of the training data in the training set of the target model) is input into the multiple network structures that need to be quantized in the first round of quantization, and corresponding output data is obtained; and, after performing parameter transformation on the multiple network structures that need to be quantized in the first round of quantization based on the quantization parameters, the same calibration data is input into the multiple network structures after the parameter transformation to obtain output data. In this way, by obtaining the difference between the output data of the multiple network structures before the parameter transformation (i.e., the original multiple network structures) and the output data of the multiple network structures after the parameter transformation, the quantization parameters of the multiple network structures can be updated based on this difference (e.g., the quantization parameters are updated by gradient backpropagation), and ultimately, the outputs of the multiple network structures after the parameter transformation based on the quantization parameters can meet the requirements.

[0134] In addition, any round of quantization process except the first round of quantization process in the multi-round quantization process is regarded as a target round quantization process. Then, the target round quantization process specifically includes: first, inputting the first input data into at least two network structures to obtain the first output data. Among them, the at least two network structures input by the first input data are the network structures quantized by the target round quantization process, and the first input data are the outputs of the network structure before the at least two network structures. Since the current at least two network structures are not located at the very beginning in the target model, the inputs of the at least two network structures are actually the feature data (i.e., the first input data) output by the previous network structure. By inputting the first input data into the at least two network structures, the first output data output by the at least two network structures before quantization can be obtained.

[0135] Then, the second input data is input into at least two network structures that have been quantized based on the quantization parameter to obtain second output data, where the second input data is the output of the network structure before the at least two network structures after quantization. Since the input received by the network structure at the middle position is the output of the quantized network structure at the previous position when the entire target model is quantized, in this embodiment, the output data of the quantized network structure at the previous position (i.e., the second input data) is used as the input of the middle network structure (i.e., the at least two network structures mentioned above).

[0136] Secondly, based on the difference between the first output data and the second output data, the quantization parameters of the at least two network structures are updated. Specifically, the quantization parameters of the at least two network structures can be updated by gradient backpropagation, so that when the at least two network structures are quantized based on the updated quantization parameters, the outputs of the at least two network structures can be closer to the outputs before quantization. That is, the process of updating the quantization parameters based on the difference between the two output data is similar to the way of updating the weight parameters during the model training process, which can be understood as updating the quantization parameters by constructing a loss function indicating the difference in output data, and the goal of updating the quantization parameters is to make the loss function as small as possible (that is, the difference between the two output data is as small as possible).

[0137] In practical applications, by repeatedly executing the above-mentioned steps of updating the quantization parameters using different input data, the quantization parameters can be continuously updated until the output of the network structure after parameter transformation based on the quantization parameters meets the requirements, or the quantization parameter update is repeatedly executed a certain number of times, and the quantization of the network structure is finally completed.

[0138] In this scheme, in the process of determining the quantization parameters of the network structure, the output of the original network structure is used as a supervisory signal, and the difference is constructed with the output of the network structure after quantization based on the quantization parameters. Then, the quantization parameters are updated with the goal of minimizing the output difference, so that accurate quantization parameters can be obtained, ensuring the accuracy of the network structure after quantization based on the quantization parameters.

[0139] Optionally, the difference between the first output data and the second output data is, for example, a weighted average of the first difference value and the second difference value, the first difference value being the Euclidean distance (i.e., L2 distance) between the first output data and the second output data, and the second difference value being the relative entropy (i.e., KL divergence) between the first output data and the second output data. In this solution, by constructing the difference between the outputs of the network structure based on the Euclidean distance between the outputs and the weighted average of the relative entropy, it is possible to effectively deal with outliers in the feature data and improve the robustness of subsequent optimization of the network structure based on the difference.

[0140] For example, please refer to Figure 4, which is a schematic diagram of a process for performing quantization on a target model provided in an embodiment of the present application. As shown in Figure 4, in the process of performing quantization on the target model, the same calibration data needs to be input simultaneously into the original target model and the target model after parameter transformation using the quantization parameters (i.e., the target model in the quantization shown in Figure 4), and the difference between the outputs of the network structure at the same position of the original target model and the target model in the quantization is compared, and then the quantization parameters corresponding to the network structure are updated based on the difference in the output.

[0141] Specifically, during the first round of quantization, a parameter transformation is first performed on the target model based on a random quantization parameter to obtain the target model in quantization. Then, the same calibration data is input into the original target model and the target model in quantization, respectively, and the output of the network structure 2 in the original target model and the output of the network structure 2 after the parameter transformation based on the quantization parameter in the target model in quantization are obtained. Secondly, by comparing the output of the network structure 2 with the output of the network structure 2 after the parameter transformation, the difference 1 can be obtained. In this way, with the goal of minimizing the difference 1, the quantization parameters corresponding to the network structure 1 and the network structure 2 can be updated by gradient backpropagation until the final outputs of the network structure 1 and the network structure 2 after the parameter transformation based on the updated quantization parameter are close to the outputs of the network structure 1 and the network structure 2 after the parameter transformation is not performed, or the number of times the quantization parameter is updated reaches a preset number, thereby obtaining the quantized network structure 1 and the network structure 2.

[0142] During the second round of quantization, while the calibration data remains unchanged, the calibration data is input into the quantized network structure 1, and the output of the quantized network structure 1 is input into the network structures 2 and 3 in the target model being quantized, to obtain the output of the network structure 3 after parameter transformation. Then, by comparing the output of the original network structure 3 with the output of the network structure 3 after parameter transformation, the difference 2 can be obtained. In this way, with the goal of minimizing the difference 2, the quantization parameters corresponding to the network structures 2 and 3 can be updated by gradient backpropagation until the final outputs of the network structures 2 and 3 after parameter transformation based on the updated quantization parameters are close to the outputs of the network structures 2 and 3 after parameter transformation, or the number of times the quantization parameters are updated reaches a preset number, thereby obtaining the quantized network structures 2 and 3.

[0143] It should be noted that Network Structure 2 actually performs two rounds of quantization. The first round of quantization is performed by Network Structure 2 in conjunction with Network Structure 1, and the second round of quantization is performed by Network Structure 2 in conjunction with Network Structure 3. Only after Network Structure 2 has completed both rounds of quantization will the corresponding quantization parameters of Network Structure 2 be fixed, meaning that Network Structure 2 is considered to have completed the quantization process.

[0144] Similarly, based on a process similar to the second round of quantization, network structure 3 and network structure 4, network structure 4 and network structure 5, and subsequent network structures can be quantized until the quantization of the last network structure in the target model is completed, thereby obtaining the quantized target model.

[0145] Generally speaking, there are two types of quantization parameters for network structures: weight parameters and activation parameters. Weight-based quantization parameters can be used to transform the weight parameters in a network structure, thereby obtaining quantized weight parameters. Similarly, activation-based quantization parameters can be used to transform the activation values ​​in a network structure, thereby obtaining quantized activation values.

[0146] Exemplarily, the quantized weight parameter may be as shown in the following formula.

[0147] in, Represents the quantized weight parameter; Δ w Represents the quantization step size of the weight parameter; w represents the weight parameter; |·| means rounding the tensor down to the nearest integer.

[0148] In addition, the quantized activation value can be shown as the following formula.

[0149] in, represents the quantized activation value; Δ X Indicates the quantization step of the activation value; X represents the activation value.

[0150] Since each round of quantization actually performs quantization on at least two network structures, the process of optimizing the quantization parameters of the network structure in any round of quantization can be shown as the following formula.

[0151] Among them, argmin(f(x)) means finding the x that makes the function f(x) take the minimum value; Recon() means finding the difference value; B l,k (W l,k , X l,k ) represents the output of network structure l-network structure k; W l,k represents the weight parameter of network structure l-network structure k; X l,k Represents the activation value of network structure l-network structure k; Represents the output of network structure l-network structure k after parameter transformation based on quantization parameters; Represents the quantized weight parameters in network structure l-network structure k; Represents the quantized activation value in network structure l-network structure k.

[0152] Specifically, the definition of Recon() can be shown as the following formula.

[0153] Recon(f1,f2)=||f1-f2||+KLD(softmax(f1),softmax(f2))

[0154] Among them, KLD() means to obtain KL divergence; softmax() means normalization.

[0155] In some embodiments, the quantization parameters of the weight parameters may also include a weight compensation matrix in addition to the above-mentioned quantization step size. Among them, the quantization step size is used to update the weight parameters in the network structure, and the weight compensation matrix is ​​used to compensate for the updated weight parameters. Then, in the process of performing quantization on the target model, it is actually a process of learning the quantization step size and the weight compensation matrix. Since a certain error will be generated after the weight parameters are updated using the quantization step size, the error caused by rounding down during the fixed-point quantization of the weights can be compensated by adding the weight compensation matrix to the updated weight parameters. Generally, the values ​​of the elements in the weight compensation matrix are usually 0 and 1.

[0156] Optionally, the weight compensation matrix is ​​decomposed into multiple matrices using a low-rank decomposition method, wherein the sum of the parameters of the multiple matrices is less than the parameter of the weight compensation matrix, and the weight compensation matrix is ​​updated by updating the multiple matrices. Low-rank decomposition is a method of decomposing a high-dimensional matrix into a low-rank matrix, which can effectively reduce the dimensionality of the data and reduce redundant information.

[0157] For example, please refer to Figure 5, which is a schematic diagram of performing low-rank decomposition on a weight compensation matrix according to an embodiment of the present application. As shown in Figure 5, the dimension of the weight compensation matrix A is d×k. After performing low-rank decomposition on the weight compensation matrix A, the weight compensation matrix A can be decomposed into matrix A1 and matrix A2. Among them, the dimension of matrix A1 is d×r, and the dimension of matrix A2 is r×k, and r is much smaller than the minimum value of d and k.

[0158] It should be noted that the low-rank decomposition method shown in Figure 5 is to decompose a weight compensation matrix into two matrices. In practical applications, a weight compensation matrix can also be decomposed into two or more matrices. The specific method is determined by the low-rank decomposition method, and this embodiment does not specifically limit this.

[0159] That is to say, by performing low-rank decomposition on a weight compensation matrix with a large number of parameters, it is possible to use multiple matrices with smaller parameters to represent a weight compensation matrix with a large number of parameters, thereby avoiding the need to store and process a weight compensation matrix with a large number of parameters during the quantization process.

[0160] In this solution, by performing a low-rank decomposition of a weight compensation matrix with a large number of parameters into multiple matrices with smaller parameters, multiple matrices with smaller parameters can be learned during the quantization process, eliminating the need to learn a weight compensation matrix with a large number of parameters. This significantly reduces the number of parameters to be learned during the quantization process, as well as the number of iterations and resource usage. Especially for large language models with a large number of parameters, performing a low-rank decomposition can decompose the weight compensation matrix with a large number of parameters into multiple matrices with very small parameters, significantly reducing the cost of learning the weight compensation matrix.

[0161] Generally speaking, in a network structure, different neural network layers correspond to different quantization parameters. For example, the quantization parameter of the weight parameter actually converts the weight parameter in the current neural network layer from a floating-point number to a fixed-point number. The distribution range of the weight parameter in the current neural network layer will affect the actual value of the quantization parameter, and ultimately affect the specific value of each weight parameter after conversion to a fixed-point number. When the values ​​of some weight parameters in a neural network layer differ significantly from the values ​​of most other weight parameters, the weight parameters will have a large distribution range, making it difficult to determine an effective quantization parameter.

[0162] Based on this, in this embodiment, during each round of quantization, one or more outliers can be selected from the values ​​of multiple weight parameters in the same neural network layer in the network structure, and the values ​​of the multiple weight parameters include one or more outliers. That is, one or more outliers are determined from the values ​​of the multiple weight parameters. One or more outliers refer to one or more values ​​in the values ​​of the multiple weight parameters that are significantly different from other values. For example, assuming that there are currently 1000 weight parameter values, of which 998 weight parameter values ​​are distributed in the range of [0, 10], and the values ​​of the other two weight parameters are 150 and 160 respectively, it can be determined that the values ​​of these two weight parameters are outliers.

[0163] Then, the values ​​of the weight parameters corresponding to the one or more outliers are set to a preset threshold, that is, the values ​​of the weight parameters belonging to the outliers are modified to the preset threshold. The preset threshold can be a preset value determined based on the distribution range of the multiple weight parameters, such as 0 or any value within the distribution range of the other weight parameters excluding the outliers in the multiple weight parameters.

[0164] It should be noted that the above process of determining outliers for weight parameters is performed independently for each neural network layer in the network structure, that is, different neural network layers may determine different outliers. The above steps can be used to determine outliers for any neural network layer, and this embodiment will not be repeated here.

[0165] In this solution, by identifying outliers in the weight parameters and adjusting the values ​​of the weight parameters belonging to the outliers to a preset threshold, the distribution range of the weight parameters can be narrowed, thereby reducing the difficulty of learning the quantization parameters and facilitating the learning of better quantization parameters, thereby improving the accuracy of the quantized target model. In addition, since outliers in the activation values ​​are also caused to a certain extent by outliers in the weight parameters, adjusting the outliers in the weight parameters can also improve the outliers in the activation values, which is conducive to learning better quantization parameters for the activation values.

[0166] For example, please refer to Figures 6 and 7. Figure 6 is a comparative schematic diagram of the distribution of a weight parameter provided in an embodiment of the present application; Figure 7 is a comparative schematic diagram of the distribution of another weight parameter provided in an embodiment of the present application. As shown in Figure 6, the distribution range of the weight parameter is larger before the outliers are removed, and the distribution range of the weight parameter becomes smaller after the outliers are removed. As shown in Figure 7, before the outliers are removed, the distribution range of the weight parameter is [0,4], and after the outliers are removed, the distribution range of the weight parameter becomes [0,2.2], which effectively narrows the distribution range of the weight parameter and facilitates the subsequent determination of the quantization parameters of the weight parameter.

[0167] Optionally, the process of selecting outliers from the values ​​of multiple weight parameters of a network structure may specifically include: first determining a coarse-grained outlier interval based on the values ​​of the multiple weight parameters, wherein the values ​​of the weight parameters in the coarse-grained outlier interval are all greater than a first threshold or all less than a second threshold, and the first threshold or the second threshold is determined based on the distribution of the values ​​of the multiple weight parameters. That is, first determining a first threshold or a second threshold based on the distribution of the values ​​of the multiple weight parameters, and determining the range of weight parameter values ​​greater than the first threshold as the coarse-grained outlier interval, or determining the range of weight parameter values ​​less than the second threshold as the coarse-grained outlier interval. Specifically, the first threshold may be a positive value, and the second threshold may be a negative value. In general, the values ​​in the coarse-grained outlier interval are values ​​far from 0. The values ​​of the multiple weight parameters mentioned above are all positive or negative. For the weight parameters in the same neural network layer, this embodiment determines one coarse-grained outlier interval for multiple weight parameters with positive values, and another coarse-grained outlier interval for multiple weight parameters with negative values, thereby further determining outliers in each coarse-grained outlier interval.

[0168] Then, in the coarse-grained outlier interval, the fine-grained outlier interval is further determined. Among them, the coarse-grained outlier interval is relative to the fine-grained outlier interval, the fine-grained outlier interval is within the range of the coarse-grained outlier interval, and the distribution of values ​​in the fine-grained outlier interval meets the preset conditions. After determining the fine-grained outlier interval, it can be determined that the value of the weight parameter in the fine-grained outlier interval belongs to one or more of the above-mentioned outliers. That is, if the value of a certain weight parameter falls within the fine-grained outlier interval, it can be considered that the value of the weight parameter is an outlier. Generally, the coarse-grained outlier interval can be determined to be divided into a reserved interval and a fine-grained outlier interval, then the preset conditions met by the distribution of values ​​in the fine-grained outlier interval can be specifically: the distance between the fine-grained outlier interval and the reserved interval is as large as possible, and the numerical distribution in the reserved interval is as close as possible.

[0169] Among them, there are multiple ways to determine the first threshold and the second threshold for determining the coarse-grained outlier interval. For example, taking the first threshold as an example, the first threshold can be determined based on the third quartile and the interquartile range of multiple weight parameters. Alternatively, the first threshold can be a value at a specific position among multiple weight parameters, such as the first threshold being a value at the 90th position after the multiple weight parameters are arranged from small to large. In general, this embodiment does not limit the method for determining the first threshold and the second threshold.

[0170] For example, please refer to Figure 8, which is a schematic diagram of a process for preprocessing weight parameters provided by an embodiment of the present application. As shown in Figure 8, the process for preprocessing weight parameters in a network structure includes the following steps 801-805.

[0171] Step 801: Calculate the quartiles and interquartile ranges based on the distribution of the weight parameters.

[0172] Specifically, for multiple weight parameters in a neural network layer, the multiple weight parameters can be arranged in ascending order, and the first quartile Q1 and the third quartile Q3 can be determined. The first quartile Q1 is specifically the 25th percentile value of the multiple weight parameters after the multiple weight parameters are arranged in ascending order; the third quartile Q3 is specifically the 75th percentile value of the multiple weight parameters after the multiple weight parameters are arranged in ascending order. The interquartile range (IQP) is the difference between the third quartile Q3 and the first quartile Q1.

[0173] Step 802: Determine an interval threshold based on the quartiles and the interquartile range.

[0174] The interval threshold may specifically be: T=Q3+λ1IQR, where λ1 is a hyperparameter and the value of λ1 may specifically be 1.5.

[0175] Step 803: Determine the interval to which the weight parameter greater than the interval threshold belongs as a coarse-grained outlier interval.

[0176] Assuming that the coarse-grained outlier interval is O, the coarse-grained outlier interval O can be specifically expressed as: O = {x|x>T, x∈X}, where x is the value of the weight parameter and X is the distribution range of the value of the weight parameter.

[0177] Step 804 : searching for a threshold in the coarse-grained outlier interval to divide the coarse-grained outlier interval into a fine-grained outlier interval and a reserved interval, with the distance between the fine-grained outlier interval and the reserved interval being as large as possible and the values ​​in the reserved interval being as close as possible.

[0178] Specifically, the values ​​of the weight parameters contained in the coarse-grained outlier interval can be traversed one by one, and then the value of one of the weight parameters can be determined to be the searched threshold value, thereby dividing the coarse-grained outlier interval into a fine-grained outlier interval and a reserved interval based on the searched threshold value. For example, within the coarse-grained outlier interval, the range of weight parameter values ​​less than or equal to the searched threshold value is the reserved interval, and the range of weight parameter values ​​greater than the searched threshold value is the fine-grained outlier interval.

[0179] In order to make the distance between the fine-grained outlier interval and the reserved interval as large as possible and the numerical distribution within the reserved interval as close as possible, the following formula can be set to search for the corresponding threshold. Specifically, assuming that the fine-grained outlier interval is 0 outlier , the reserved interval is O reserved , then the fine-grained outlier interval O outlier With the reserved interval O reserved The distance between them, and the retention interval O reserved The internal numerical distribution satisfies the following formula.

[0180] M intra =var(O reserued )

[0181] M inter =(min(O outlier )-max(O reserved )) 2

[0182] M=M inter -λ2M intra

[0183] Among them, M intra Indicates the reserved interval O reserved Variance of internal values; M inter Represents the fine-grained outlier interval O outlier With the reserved interval O reservedThe distance between them; λ2 represents a hyperparameter, for example, 0.1. By traversing each value in the coarse-grained outlier interval, the value that can make M have the maximum value is determined to be the above threshold, and then the final fine-grained outlier interval O is determined. outlier With the reserved interval O reserved .

[0184] Step 805: Set the value of the weight parameter in the fine-grained outlier interval to 0.

[0185] After the fine-grained outlier interval is determined, the value of each weight parameter within the fine-grained outlier interval can be set to 0, thereby removing outlier values ​​of the weight parameters.

[0186] The above describes the process of identifying and handling outliers in weight parameters. Similarly, before quantizing activation values ​​in a network structure, outliers in activation values ​​can also be identified and handled.

[0187] For example, during each round of quantization, at least one outlier may be selected from a plurality of activation values ​​within a same neural network layer in the network structure. The at least one outlier refers to one or more values ​​in the plurality of activation values ​​that differ significantly from other values.

[0188] Then, the activation value corresponding to at least one outlier is scaled according to a target ratio, so that the scaled activation value is the same as the maximum activation value among the multiple activation values ​​excluding the outlier. It should be noted that the target scaling ratio is different for activation values ​​with different values. For example, suppose there are currently 1000 activation values, of which 998 activation values ​​are distributed in the range [0, 10], and the maximum activation value of these 998 activation values ​​is 10; and the values ​​of the other two activation values ​​that are outliers are 80 and 100 respectively. In this way, for the activation value of 80, this activation value can be scaled to 10 according to a ratio of 8:1; for the activation value of 100, this activation value can be scaled to 10 according to a ratio of 10:1.

[0189] For example, please refer to Figures 9 and 10. Figure 9 is a comparative schematic diagram of the distribution of activation values ​​provided in an embodiment of the present application; Figure 10 is a comparative schematic diagram of the distribution of another activation value provided in an embodiment of the present application. As shown in Figure 9, the distribution range of the activation values ​​is large before the outliers are removed, and the distribution range of the activation values ​​becomes smaller after the outliers are scaled. As shown in Figure 10, before the outliers are removed, the distribution range of the activation values ​​is [0,280], and after the outliers are removed, the distribution range of the activation values ​​becomes [0,40], which effectively narrows the distribution range of the activation values ​​and facilitates the subsequent determination of the quantization parameters of the activation values.

[0190] After determining the outlier value corresponding to the activation value, unlike setting the outlier value of the weight parameter to a preset threshold, in this embodiment, scaling is performed on the activation value. It is understandable that since the activation value is often the output of the neural network layer, if the activation value is directly set to a preset threshold (such as 0), it will directly have a greater impact on the output of the neural network layer, which is likely to affect the accuracy of the model. Therefore, in this embodiment, scaling is performed on the outlier value of the activation value. For the outlier value of the weight parameter, since the weight parameter indirectly affects the output of the neural network layer, setting the weight parameter to the preset threshold does not have much impact on the accuracy of the model, and can effectively reduce the difficulty of learning the quantization parameter.

[0191] In this scheme, by identifying outliers in the activation values ​​and performing scaling on the activation values ​​belonging to the outliers, the distribution range of the activation values ​​can be narrowed, thereby reducing the difficulty of learning quantization parameters and facilitating the learning of better quantization parameters, thereby improving the accuracy of the quantized target model.

[0192] It should be noted that the method of determining outliers based on the values ​​of multiple activation values ​​in this embodiment can be similar to the method of determining outliers based on the values ​​of multiple weight parameters mentioned above. Please refer to the above embodiment for details and will not be repeated here.

[0193] The above describes in detail the method provided by the embodiment of the present application. Next, the device provided by the embodiment of the present application for executing the above method will be introduced.

[0194] Please refer to Figure 11, which is a structural diagram of a model quantization device provided in an embodiment of the present application. As shown in Figure 11, the model quantization device provided in an embodiment of the present application includes: an acquisition module 1101, which is used to acquire a target model, and the target model includes multiple network structures connected in sequence; a processing module 1102, which is used to perform multiple rounds of quantization process to obtain a quantized target model, and the quantized target model includes multiple quantized network structures, and the multiple rounds of quantization process are performed in sequence based on the connection order of the multiple network structures; wherein, in the multiple rounds of quantization process, each round of quantization process is to quantize at least two consecutive network structures in the multiple network structures, and each round of quantization process is to quantize part of the network structures in the multiple network structures, and for any two consecutive rounds of quantization process in the multiple rounds of quantization process, there is at least one repeated network structure between the network structure quantized by the latter round of quantization process and the network structure quantized by the previous round of quantization process.

[0195] In a possible implementation, the first N network structures quantized in the next round of quantization process and the last N network structures quantized in the previous round of quantization process are the same N network structures, where N is an integer greater than or equal to 1.

[0196] In one possible implementation, the processing module 1102 is further used to: input first input data into at least two network structures to obtain first output data, where the at least two network structures are network structures quantized by a target round quantization process, the first input data are outputs of the network structure before at least two network structures, and the target round quantization process is one round of quantization in a multi-round quantization process; input second input data into at least two network structures after quantization based on quantization parameters to obtain second output data, where the second input data are outputs of the network structure before at least two network structures after quantization; and update the quantization parameters of the at least two network structures based on the difference between the first output data and the second output data.

[0197] In one possible implementation, the quantization parameters include a quantization step and a weight compensation matrix, the quantization step is used to update the weight parameters in the network structure, and the weight compensation matrix is ​​used to compensate for the updated weight parameters; wherein, the weight compensation matrix is ​​decomposed into multiple matrices by low rank, the sum of the parameter quantities of the multiple matrices is less than the parameter quantity of the weight compensation matrix, and the update of the weight compensation matrix is ​​achieved by updating multiple matrices.

[0198] In one possible implementation, the difference is a weighted average of the first difference value and the second difference value, the first difference value is the Euclidean distance between the first output data and the second output data, and the second difference value is the relative entropy between the first output data and the second output data.

[0199] In a possible implementation, during each round of quantization, the processing module 1102 is further used to: select one or more outliers from the values ​​of multiple weight parameters of the network structure; and set the values ​​of the weight parameters corresponding to the one or more outliers to a preset threshold.

[0200] In one possible implementation, the processing module 1102 is further used to: determine a coarse-grained outlier interval based on the values ​​of multiple weight parameters, wherein the values ​​of the weight parameters in the coarse-grained outlier interval are all greater than a first threshold or are all less than a second threshold, and the first threshold or the second threshold is determined based on the distribution of the values ​​of the multiple weight parameters; determine a fine-grained outlier interval based on the coarse-grained outlier interval, wherein the fine-grained outlier interval is within the range of the coarse-grained outlier interval, and the distribution of values ​​in the fine-grained outlier interval meets a preset condition, and the value of the weight parameter in the fine-grained outlier interval belongs to one or more outlier values.

[0201] In one possible implementation, during the execution of multiple rounds of quantization and before quantizing the activation values ​​in the network structure, the processing module 1102 is further used to: select at least one outlier from the values ​​of multiple activation values ​​of the network structure, the values ​​of the multiple activation values ​​include at least one outlier; and scale the value of the activation value corresponding to the at least one outlier according to a target ratio.

[0202] In one possible implementation, the target model is a large language model.

[0203] In a possible implementation, each of the multiple network structures includes one or more Transformer network blocks.

[0204] Please refer to Figure 12, which is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device 1200 can be specifically manifested as a mobile phone, a tablet, a laptop computer, a smart wearable device, a server, etc., which is not limited here. Specifically, the electronic device 1200 includes: a receiver 1201, a transmitter 1202, a processor 1203 and a memory 1204 (wherein the number of processors 1203 in the electronic device 1200 can be one or more, and Figure 12 takes one processor as an example), wherein the processor 1203 may include an application processor 12031 and a communication processor 12032. In some embodiments of the present application, the receiver 1201, the transmitter 1202, the processor 1203 and the memory 1204 may be connected via a bus or other means.

[0205] The memory 1204 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1203. A portion of the memory 1204 may also include non-volatile random access memory (NVRAM). The memory 1204 stores processor and operation instructions, executable modules, or data structures, or subsets or extended sets thereof. The operation instructions may include various operation instructions for implementing various operations.

[0206] Processor 1203 controls the operation of the electronic device. In specific applications, the various components of the electronic device are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all bus systems are referred to as a bus system in the figure.

[0207] The method disclosed in the above embodiment of the present application can be applied to the processor 1203, or implemented by the processor 1203. The processor 1203 can be an integrated circuit chip with signal processing capabilities. During the implementation process, each step of the above method can be completed by an integrated logic circuit of the hardware in the processor 1203 or an instruction in the form of software. The above-mentioned processor 1203 can be a general-purpose processor, a digital signal processor (digital signal processing, DSP), a microprocessor or a microcontroller, and can further include an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.

[0208] The processor 1203 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 1204, and the processor 1203 reads the information in the memory 1204 and completes the steps of the above method in combination with its hardware.

[0209] Receiver 1201 can be used to receive input digital or character information and generate signal input related to the relevant settings and function control of the electronic device. Transmitter 1202 can be used to output digital or character information through the first interface. Transmitter 1202 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group. Transmitter 1202 can also include a display device such as a display screen.

[0210] The electronic device provided in the embodiment of the present application may specifically be a chip, and the chip includes: a processing unit and a communication unit, the processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin or a circuit, etc. The processing unit may execute the computer execution instructions stored in the storage unit, so that the chip in the electronic device executes the classification method of multimedia data described in the above embodiment, or so that the chip in the training device executes the model quantization method described in the above embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit may also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0211] Specifically, see Figure 13 , which is a schematic diagram of the structure of a chip provided in an embodiment of the present application. The chip can be a neural network processor (NPU) 1300. NPU 1300 is mounted on the host CPU (host CPU) as a coprocessor, and the host CPU assigns tasks. The core of the NPU is arithmetic circuit 1303, which is controlled by controller 1304 to extract matrix data from memory and perform multiplication operations.

[0212] In some implementations, the arithmetic circuit 1303 includes multiple processing units (PEs). In some implementations, the arithmetic circuit 1303 is a two-dimensional systolic array. The arithmetic circuit 1303 can also be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1303 is a general-purpose matrix processor.

[0213] For example, assume there are input matrix A, weight matrix B, and output matrix C. The computation circuit retrieves the corresponding data of matrix B from weight memory 1302 and caches it on each PE in the computation circuit. The computation circuit then retrieves the data of matrix A from input memory 1301 and performs a matrix operation on it with matrix B. The partial or final matrix result is stored in accumulator 1308.

[0214] Unified memory 1306 is used to store input and output data. Weight data is directly transferred to weight memory 1302 through the Direct Memory Access Controller (DMAC) 1305. Input data is also transferred to unified memory 1306 through the DMAC.

[0215] BIU stands for Bus Interface Unit 1310 , which is used for interaction between the AXI bus, DMAC, and instruction fetch buffer (IFB) 1309 .

[0216] The bus interface unit 1310 (BIU) is used for the instruction fetch memory 1309 to obtain instructions from the external memory, and is also used for the storage unit access controller 1305 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0217] DMAC is mainly used to move input data in the external memory DDR to the unified memory 1306 or to move weight data to the weight memory 1302 or to move input data to the input memory 1301.

[0218] The vector calculation unit 1307 includes multiple operation processing units. When necessary, it further processes the output of the operation circuit 1303, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as batch normalization, pixel-level summation, and upsampling of feature planes.

[0219] In some implementations, the vector calculation unit 1307 can store the processed output vector in the unified memory 1306. For example, the vector calculation unit 1307 can apply a linear function or a nonlinear function to the output of the operation circuit 1303, such as linear interpolation of the feature plane extracted by the convolution layer, or accumulate a vector of values ​​to generate an activation value. In some implementations, the vector calculation unit 1307 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 1303, for example, for use in subsequent layers in a neural network.

[0220] An instruction fetch buffer 1309 connected to the controller 1304 is used to store instructions used by the controller 1304;

[0221] Unified memory 1306, input memory 1301, weight memory 1302, and instruction fetch memory 1309 are all on-chip memories. External memories are private to the NPU hardware architecture.

[0222] The processor mentioned in any of the above places can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above program.

[0223] Please refer to Figure 14, which is a schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of the present application. The present application also provides a computer-readable storage medium. In some embodiments, the method disclosed in Figure 2 above can be implemented as computer program instructions encoded in a machine-readable format on a computer-readable storage medium or on other non-transitory media or products.

[0224] 14 schematically illustrates a conceptual partial view of an example computer-readable storage medium including a computer program for executing a computer process on a computing device, arranged in accordance with at least some embodiments presented herein.

[0225] In one embodiment, computer readable storage medium 1400 is provided using signal bearing medium 1401. Signal bearing medium 1401 may include one or more program instructions 1402 that, when executed by one or more processors, may provide the functionality or portions of the functionality described above with respect to FIG.

[0226] In some examples, the signal bearing medium 1401 may include a computer readable medium 1403 such as, but not limited to, a hard drive, a compact disk (CD), a digital video disk (DVD), a digital tape, a memory, a ROM or RAM, or the like.

[0227] In some embodiments, the signal-bearing medium 1401 may include a computer-recordable medium 1404, such as, but not limited to, a memory, a read / write (R / W) CD, a R / W DVD, or the like. In some embodiments, the signal-bearing medium 1401 may include a communication medium 1405, such as, but not limited to, a digital and / or analog communication medium (e.g., a fiber optic cable, a waveguide, a wired communication link, a wireless communication link, or the like). Thus, for example, the signal-bearing medium 1401 may be communicated via a wireless form of the communication medium 1405 (e.g., a wireless communication medium conforming to the IEEE 802.X standard or other transmission protocol).

[0228] The one or more program instructions 1402 may be, for example, computer-executable instructions or logic-implemented instructions. In some examples, the computing device may be configured to provide various operations, functions, or actions in response to the program instructions 1402 communicated to the computing device via one or more of computer-readable media 1403, computer-recordable media 1404, and / or communication media 1405.

[0229] It should also be noted that the device embodiments described above are merely illustrative, in which the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.

[0230] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods of each embodiment of the present application.

[0231] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0232] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training equipment or data center to another website, computer, training equipment or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training equipment, data center, etc. that includes one or more available media integrations. Available media can be magnetic media, (e.g., floppy disk, hard disk, tape), optical media (e.g., DVD), or semiconductor media (e.g., solid-state drive (SSD)), etc.

Claims

1. A model quantization method, characterized in that: include: Acquire a target model, wherein the target model includes a plurality of network structures connected in sequence; Performing multiple rounds of quantization process to obtain a quantized target model, wherein the quantized target model includes the multiple quantized network structures, and the multiple rounds of quantization process are performed sequentially based on the connection order of the multiple network structures; Among them, in the multiple rounds of quantization process, each round of quantization process performs quantization on at least two consecutive network structures among the multiple network structures, and each round of quantization process performs quantization on part of the multiple network structures. For any two consecutive rounds of quantization process in the multiple rounds of quantization process, there is at least one repeated network structure between the network structure quantized by the latter round of quantization process and the network structure quantized by the previous round of quantization process.

2. The method according to claim 1, characterized in that The first N network structures quantized in the latter round of quantization process and the last N network structures quantized in the former round of quantization process are the same N network structures, where N is an integer greater than or equal to 1.

3. The method according to claim 1 or 2, characterized in that: The target round quantization process in the multi-round quantization process includes: Inputting first input data into at least two network structures to obtain first output data, wherein the at least two network structures are network structures quantized by a target round quantization process, the first input data are outputs of a network structure before the at least two network structures, and the target round quantization process is a round quantization process in the multiple rounds of quantization process; Inputting second input data into the at least two network structures after quantization based on the quantization parameter to obtain second output data, wherein the second input data is the output of the network structure before the at least two network structures after quantization; Based on the difference between the first output data and the second output data, the quantization parameters of the at least two network structures are updated.

4. The method according to claim 3, characterized in that The quantization parameters include a quantization step length and a weight compensation matrix, wherein the quantization step length is used to update the weight parameters in the network structure, and the weight compensation matrix is ​​used to compensate for the updated weight parameters; The weight compensation matrix is ​​decomposed into multiple matrices by low rank, the sum of the parameter quantities of the multiple matrices is less than the parameter quantity of the weight compensation matrix, and the updating of the weight compensation matrix is ​​achieved by updating the multiple matrices.

5. The method according to claim 3 or 4, characterized in that: The difference is a weighted average value between a first difference value and a second difference value, the first difference value is the Euclidean distance between the first output data and the second output data, and the second difference value is the relative entropy between the first output data and the second output data.

6. The method according to any one of claims 1 to 5, characterized in that: During each round of quantization, the method further includes: Select one or more outliers from the values ​​of multiple weight parameters of the network structure; The value of the weight parameter corresponding to the one or more outliers is set to a preset threshold.

7. The method according to claim 6, characterized in that The step of selecting one or more outliers from the values ​​of multiple weight parameters of the network structure comprises: Determining a coarse-grained outlier interval based on the values ​​of the multiple weight parameters, wherein the values ​​of the weight parameters in the coarse-grained outlier interval are all greater than a first threshold or are all less than a second threshold, and the first threshold or the second threshold is determined based on the distribution of the values ​​of the multiple weight parameters; Based on the coarse-grained outlier interval, a fine-grained outlier interval is determined, the fine-grained outlier interval is within the range of the coarse-grained outlier interval, and the distribution of values ​​in the fine-grained outlier interval meets a preset condition, and the value of the weight parameter in the fine-grained outlier interval belongs to the one or more outliers.

8. The method according to any one of claims 1 to 7, characterized in that: During each round of quantization, the method further includes: Select at least one outlier from the values ​​of multiple activation values ​​in the network structure; The activation value corresponding to the at least one outlier is scaled according to a target ratio.

9. The method according to any one of claims 1 to 8, characterized in that: The target model is a large language model.

10. The method according to claim 9, characterized in that Each of the multiple network structures includes one or more Transformer network blocks.

11. A model quantization device, characterized in that: include: An acquisition module, used for acquiring a target model, wherein the target model includes a plurality of network structures connected in sequence; A processing module, configured to execute a multi-round quantization process to obtain a quantized target model, wherein the quantized target model includes the quantized multiple network structures, and the multi-round quantization process is executed sequentially based on the connection order of the multiple network structures; Among them, in the multiple rounds of quantization process, each round of quantization process performs quantization on at least two consecutive network structures among the multiple network structures, and each round of quantization process performs quantization on part of the multiple network structures. For any two consecutive rounds of quantization process in the multiple rounds of quantization process, there is at least one repeated network structure between the network structure quantized by the latter round of quantization process and the network structure quantized by the previous round of quantization process.

12. The device according to claim 11, characterized in that The first N network structures quantized in the latter round of quantization process and the last N network structures quantized in the former round of quantization process are the same N network structures, where N is an integer greater than or equal to 1.

13. The device according to claim 11 or 12, characterized in that The processing module is further used for: Inputting first input data into at least two network structures to obtain first output data, wherein the at least two network structures are network structures quantized by a target round quantization process, the first input data are outputs of a network structure before the at least two network structures, and the target round quantization process is a round quantization process in the multiple rounds of quantization process; Inputting second input data into the at least two network structures after quantization based on the quantization parameter to obtain second output data, wherein the second input data is the output of the network structure before the at least two network structures after quantization; Based on the difference between the first output data and the second output data, the quantization parameters of the at least two network structures are updated.

14. The device according to claim 13, characterized in that The quantization parameters include a quantization step length and a weight compensation matrix, wherein the quantization step length is used to update the weight parameters in the network structure, and the weight compensation matrix is ​​used to compensate for the updated weight parameters; The weight compensation matrix is ​​decomposed into multiple matrices by low rank, the sum of the parameter quantities of the multiple matrices is less than the parameter quantity of the weight compensation matrix, and the updating of the weight compensation matrix is ​​achieved by updating the multiple matrices.

15. The device according to claim 13 or 14, characterized in that The difference is a weighted average value between a first difference value and a second difference value, the first difference value is a Euclidean distance between the first output data and the second output data, and the second difference value is a weighted average value of a relative entropy between the first output data and the second output data.

16. The device according to any one of claims 11 to 15, characterized in that: During each round of quantization, the processing module is further configured to: Select one or more outliers from the values ​​of multiple weight parameters of the network structure; The value of the weight parameter corresponding to the one or more outliers is set to a preset threshold.

17. The device according to claim 16, characterized in that The processing module is further used for: Determining a coarse-grained outlier interval based on the values ​​of the multiple weight parameters, wherein the values ​​of the weight parameters in the coarse-grained outlier interval are all greater than a first threshold or are all less than a second threshold, and the first threshold or the second threshold is determined based on the distribution of the values ​​of the multiple weight parameters; Based on the coarse-grained outlier interval, a fine-grained outlier interval is determined, the fine-grained outlier interval is within the range of the coarse-grained outlier interval, and the distribution of values ​​in the fine-grained outlier interval meets a preset condition, and the value of the weight parameter in the fine-grained outlier interval belongs to the one or more outliers.

18. The device according to any one of claims 11 to 17, characterized in that: During each round of quantization, the processing module is further configured to: Selecting at least one outlier from a plurality of activation values ​​of the network structure, wherein the plurality of activation values ​​include the at least one outlier; The activation value corresponding to the at least one outlier is scaled according to a target ratio.

19. The device according to any one of claims 11 to 18, characterized in that: The target model is a large language model.

20. The device according to claim 19, characterized in that Each of the multiple network structures includes one or more Transformer network blocks.

21. A model quantization device, characterized in that: The device comprises a memory and a processor; the memory stores codes, the processor is configured to execute the codes, and when the codes are executed, the device executes the method according to any one of claims 1 to 10.

22. A computer storage medium, characterized in that The computer storage medium stores instructions, which, when executed by a computer, cause the computer to implement the method of any one of claims 1 to 10.

Citation Information

Patent Citations

  • Model quantification method and related device

    CN120068975A

  • Target network quantification method and electronic equipment

    CN116541646A

  • Model quantification method and device, electronic equipment and storage medium

    CN117035006A

Cited By

  • Model weight quantification method, electronic device and program product

    CN121351913A