Quantitative model degradation correction method and device, electronic equipment and storage medium
By acquiring and registering the key operators of the quantization model and injecting low rank weights for fine-tuning, the error problem that the quantization model is uncontrollable during the end-side transformation process is solved, and the accuracy and output consistency of the model are improved.
Patent Information
- Application Number
- CN202510326222.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-08-19
AI Technical Summary
The uncontrollable error and accuracy reduction caused by the quantization model during the end-side transformation process, especially when the user input is not determined.
By obtaining the key operators of the original model, the model inference framework is used to implement the quantitative version of the model inference operator, and register it in the deep learning training framework, inject fine-tuning low rank weights, fine-tuning the quantitative model, obtain fine-tuning parameter results with higher accuracy, and loading it on the quantitative model to correct degeneration.
Effectively reduce the error of the quantization model, improve the accuracy of the model on specific tasks, and ensure the consistency and accuracy of the output after deployment.
Smart Images

Figure CN120509443A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of quantization model technology, and in particular to a quantization model degradation correction method, device, electronic device and storage medium. Background Art
[0002] The on-device deployment of large models is often accompanied by model quantization to achieve high-speed, low-memory generation. However, because model parameter precision is reduced, model quantization often leads to a decrease in inference accuracy. This decrease is data-specific, and without user input, it is often difficult to determine when and to what extent accuracy will decrease. Therefore, if a fine-tuned network (i.e., fine-tuned parameters) tailored to a specific task of the original model is loaded onto a quantized model, uncontrollable errors will occur. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a quantization model degradation correction method, device, electronic device, and storage medium to facilitate reducing errors generated by the quantization model.
[0004] In a first aspect, an embodiment of the present invention provides a method for correcting quantization model degradation, comprising: Get the key operators of the original model; Implement the key operators using a model inference framework to obtain a quantized version of the model inference operator; Registering the model inference operator in a deep learning training framework, and using the model inference operator to replace the corresponding operator in the deep learning training framework to obtain a quantized model; Injecting fine-tuned low-rank weights into each linear layer of the quantized model to obtain an injected quantized model; Fine-tuning the post-injection quantization model to obtain a fine-tuning parameter result; The fine-tuning parameter results are loaded onto the quantization model to correct its degradation.
[0005] According to a specific implementation of an embodiment of the present invention, implementing the key operator using a model inference framework to obtain a quantized version of the model inference operator includes: A deep learning compiler is used to implement the key operators to obtain a quantized version of the model inference operator.
[0006] According to a specific implementation of an embodiment of the present invention, registering the model inference operator into a deep learning training framework and using the model inference operator to replace a corresponding operator in the deep learning training framework includes: Register the model inference operator in the deep learning training framework, encapsulate the model inference operator, and use the encapsulated model inference operator to replace the corresponding operator in the deep learning training framework; wherein the interface of the encapsulated model inference operator is adapted to the corresponding interface in the deep learning training framework.
[0007] According to a specific implementation of an embodiment of the present invention, fine-tuning the post-injection quantization model to obtain a fine-tuning parameter result includes: The post-injection quantization model is fine-tuned on a dedicated task to obtain a fine-tuning parameter result for the dedicated task.
[0008] According to a specific implementation of the embodiment of the present invention, loading the fine-tuning parameter result onto the quantization model includes: A deep learning compiler is used to load the fine-tuning parameter results onto the quantization model.
[0009] According to a specific implementation method of an embodiment of the present invention, in the process of implementing the key operator using a model inference framework, the central processing unit is bound as the computing device of the key operator; wherein the computing device of the deep learning training framework is a graphics card.
[0010] In a second aspect, an embodiment of the present invention provides a quantization model degradation correction device, comprising: The operator implementation unit is used to implement the key operators of the original model using the model inference framework to obtain the quantized version of the model inference operator; A registration unit, configured to register the model inference operator into a deep learning training framework, and use the model inference operator to replace the corresponding operator in the deep learning training framework to obtain a quantized model; An injection unit, configured to inject fine-tuned low-rank weights into each linear layer of the quantization model to obtain an injected quantization model; A fine-tuning unit, configured to fine-tune the post-injection quantization model to obtain a fine-tuning parameter result; A loading unit is used to load the fine-tuning parameter result onto the quantization model to correct its degradation.
[0011] According to a specific implementation method of an embodiment of the present invention, the operator implementation unit is specifically used to implement the key operator using a deep learning compiler to obtain a quantized version of the model inference operator.
[0012] According to a specific implementation method of an embodiment of the present invention, the registration unit is specifically used to register the model inference operator into the deep learning training framework, encapsulate the model inference operator, and use the encapsulated model inference operator to replace the corresponding operator in the deep learning training framework; wherein, the interface of the encapsulated model inference operator is adapted to the corresponding interface in the deep learning training framework.
[0013] According to a specific implementation method of an embodiment of the present invention, the registration unit registers the model inference operator into a deep learning training framework, and uses the model inference operator to replace the corresponding operator in the deep learning training framework, including: registering the model inference operator into the deep learning training framework, encapsulating the model inference operator, and replacing the corresponding operator in the deep learning training framework with the encapsulated model inference operator; wherein, the interface of the encapsulated model inference operator is adapted to the corresponding interface in the deep learning training framework.
[0014] According to a specific implementation of an embodiment of the present invention, the fine-tuning unit is specifically configured to: fine-tune the post-injection quantization model on a dedicated task to obtain a fine-tuning parameter result for the dedicated task.
[0015] According to a specific implementation of an embodiment of the present invention, the loading unit is specifically configured to load the fine-tuning parameter result onto the quantization model using a deep learning compiler.
[0016] According to a specific implementation method of an embodiment of the present invention, in the process of the operator implementation unit using the model inference framework to implement the key operator, the central processing unit is bound as the computing device of the key operator; wherein the computing device of the deep learning training framework is a graphics card.
[0017] In a third aspect, an embodiment of the present invention provides an electronic device, comprising: a processor and a memory; the memory is used to store executable program code; the processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, so as to execute the method described in any of the aforementioned implementation methods.
[0018] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the method described in any of the aforementioned embodiments.
[0019] An embodiment of the present invention provides a method, device, electronic device and computer-readable storage medium for correcting degradation of a quantization model. After obtaining the key operator of the original model, a model inference framework is used to implement the key operator to obtain a quantized version of the model inference operator; the quantized version of the model inference operator is registered in a deep learning training framework, and the model inference operator is used to replace the corresponding operator in the deep learning training framework, so that a quantized model can be obtained; fine-tuning low-rank weights are injected into each linear layer of the quantization model to obtain an injected quantization model; then a dedicated task or a specific task is used to train and fine-tune the parameters of the injected quantization model to obtain a fine-tuning parameter result with higher accuracy. That is, in the process of training the quantization model, the feedback of the quantized inference model (including the quantized version of the model inference operator) must be used to make the fine-tuning parameters obtained by training more accurate; the fine-tuning parameter result with higher accuracy is loaded onto the quantization model to correct its degradation. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 A flowchart of a method for correcting quantization model degradation according to an embodiment of the present invention is shown; Figure 2 Schematic diagram of another flow chart of a method for correcting quantization model degradation according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a quantization model degradation correction device according to an embodiment of the present invention; Figure 4 The figure is a schematic structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0022] The following describes embodiments of the present invention in detail with reference to the accompanying drawings. It should be understood that the embodiments described are only some of the embodiments of the present invention, and not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are intended to fall within the scope of protection of the present invention.
[0023] This embodiment provides a method for correcting quantization model degradation to reduce errors generated by the quantization model.
[0024] Figure 1FIG. 1 is a flow chart of a method for correcting quantization model degradation according to the first embodiment of the present invention. Figure 1 As shown, the application scenario of this embodiment is that when the fine-tuning parameters for a specific task on the original model are loaded on the quantized model, uncontrollable degradation errors will be generated.
[0025] The method of this embodiment may include: Step 101: Obtain key operators of the original model.
[0026] Among them, the original model usually refers to an unmodified model obtained directly from the training process. The model contains all the parameters and structures learned from the training data. It is the starting point for model deployment and inference.
[0027] In deep learning models, layers and operators are the basic components of the model. They are responsible for performing a series of mathematical transformations on the input data to extract features, learn patterns, and ultimately make predictions. The following is a detailed introduction to layers and operators: A layer is a fundamental unit in deep learning models. It consists of a set of parameters (weights and biases) and one or more operators. Each layer is responsible for performing specific processing on the input data and then passing the results to the next layer. Common layer types include fully connected layers, convolutional layers, recurrent layers, pooling layers, and normalization layers.
[0028] Operators are functions or methods that perform specific mathematical operations. In deep learning frameworks, operators are components of layers, defining the forward and backward propagation logic of the layer. Common operators include matrix multiplication, convolution, activation, pooling, normalization, and loss functions.
[0029] Based on the above, key layers are layers that play an important role in the model. They are generally responsible for performing the most critical tasks in the model, such as feature extraction, information integration, or decision making. Similarly, key operators are functions or operations that perform core computations during the model's forward and backward propagation processes. These operators define the behavior of the layer and are crucial for model training and inference.
[0030] It should be noted here that the embodiment of the present invention does not limit the selection of specific key operators. During implementation, the implementer can select according to actual conditions.
[0031] Step 102: Implement the key operator using a model inference framework to obtain a quantized version of the model inference operator.
[0032] After obtaining the key operators of the original model, the key operators are re-implemented using a model inference framework to obtain quantized versions of the model inference operators.
[0033] Quantization is a model optimization technique that reduces the model's storage size and computational requirements by converting floating-point weights and activation values in the model into low-precision representations (such as int8 or int4). The purpose of quantization is to improve the model's running efficiency on hardware, especially in resource-constrained environments such as mobile devices or embedded systems.
[0034] The quantized version of the model inference operator can reduce storage size and computing requirements compared to the unquantized model inference operator.
[0035] The model inference framework can be TVM (Tensor Virtual Machine). TVM is an open-source deep learning compiler that can optimize deep learning models and deploy them on various hardware platforms, enabling efficient inference and training.
[0036] In addition to TVM, model inference frameworks can also use LLVM, CUDA, and others. LLVM is a modular and reusable compiler framework and tool chain, primarily used for developing compilers and related tools; CUDA is a parallel computing platform and programming model launched by NVIDIA.
[0037] The specific method of re-implementing the key operator using the model reasoning framework to obtain a quantized version of the model reasoning operator can be to recompile the operation logic of the key operator using the model reasoning framework method to form a key operator that conforms to the logic of the model reasoning framework.
[0038] Step 103: register the model inference operator in the deep learning training framework, and use the model inference operator to replace the corresponding operator in the deep learning training framework to obtain a quantized model.
[0039] The quantized model refers to a deep learning model that has been quantized. Since the model inference operator registered with the deep learning training framework is a quantized model inference operator, the quantized model is obtained by registering the model inference operator with the deep learning training framework and replacing the corresponding operator in the deep learning training framework with the model inference operator.
[0040] The deep learning training framework has an operator registration interface, and the model inference operator can be registered in the deep learning training framework through the operator registration interface, so that the deep learning training framework can apply the model inference operator to perform corresponding calculations.
[0041] In some implementations, the deep learning training framework can be PyTorch. PyTorch is an open source deep learning framework for machine learning and deep learning that primarily implements automatic differentiation and introduces dynamic computational graphs to make model building more flexible.
[0042] Registering the model inference operator with PyTorch means integrating custom functions or operations into PyTorch's computational graph, enabling it to take advantage of PyTorch's automatic differentiation, multi-device support (such as CPU and GPU), and other optimization features. This allows custom functions to work seamlessly within the PyTorch framework, just like PyTorch's built-in functions.
[0043] In addition to pytorch, deep learning training frameworks can also be TensorFlow, Scikit-learn, etc.
[0044] Step 104: inject fine-tuned low-rank weights into each linear layer of the quantization model to obtain an injected quantization model.
[0045] In deep learning, the linear layers of a model usually refer to the layers in the model that perform linear transformations. These layers play a core role in different types of neural networks, such as fully connected layers, embedding layers, self-attention layers, and normalization layers.
[0046] Fine-tuning low-rank weights is injected into each linear layer of the quantized model to fine-tune the quantized model to improve the performance of the model on a specific task.
[0047] Fine-tuning low-rank weights (also called low-rank matrices) is injected into each linear layer of the quantization model. Various methods such as lora, Qlora, adapter tuning, prefix tuning, prompt tuning, P-Tuning, and P-Tuning v2 can be used.
[0048] LoRa is a technology for fine-tuning large pre-trained models. Its core idea is to improve efficiency and effectiveness by updating only a small number of parameters during fine-tuning rather than updating the weights of the entire model. Specifically, LoRa significantly reduces the number of parameters that need to be adjusted by decomposing the weight update ΔW into the product of two low-rank matrices B and A. This approach allows the model to adapt to new tasks or datasets by adjusting a small number of parameters while keeping most of the pre-trained weights unchanged.
[0049] Low-rank weights refer to weight matrices represented by low-rank matrix approximation techniques in machine learning and deep learning. The core idea is to decompose a large, complex weight matrix into the product of two or more smaller matrices. The product of these smaller matrices can approximate the original weight matrix. The main advantage of low-rank weights is that they reduce the number of model parameters, thereby reducing the model's storage and computational costs while maintaining model performance.
[0050] Step 105: fine-tune the post-injection quantization model to obtain a fine-tuning parameter result.
[0051] Fine-tuning involves adjusting the parameters of a pre-trained model to adapt it to a new, specific task. This process typically occurs after a model has been trained to perform a task (such as image recognition or language translation) and then applied to a similar but slightly different task. The goal of fine-tuning is to leverage the knowledge of a pre-trained model to quickly adapt to a new task without training the entire model from scratch.
[0052] Fine-tuning parameter results refer to the final set of parameters obtained after adjusting the model parameters during the fine-tuning process. These parameters can make the model better suited to a specific task or dataset after fine-tuning.
[0053] Step 106: Load the fine-tuning parameter result into the quantization model to correct its degradation.
[0054] Degradation here refers to the performance degradation that may be caused by the quantization process. By loading fine-tuning parameters onto the quantized model, this performance degradation can be compensated or corrected, allowing the quantized model to maintain a small model size and fast inference while maintaining as high accuracy as possible.
[0055] The error in inference usually comes from the quantized inference model. This embodiment provides a method for correcting degradation of a quantized model. After obtaining the key operators of the original model, the model inference framework is used to implement the key operators to obtain a quantized version of the model inference operator; the quantized version of the model inference operator is registered in the deep learning training framework, and the model inference operator is used to replace the corresponding operator in the deep learning training framework, so that a quantized model can be obtained; the low-rank weights are injected into each linear layer of the quantized model to obtain an injected quantized model; then a dedicated task or a specific task is used to train and fine-tune the parameters of the injected quantized model to obtain a more accurate fine-tuning parameter result. That is, in the process of training the quantized model, the feedback of the quantized inference model (including the quantized version of the model inference operator) will inevitably be used to make the fine-tuning parameters obtained by training more accurate; the more accurate fine-tuning parameter result is loaded onto the quantized model to correct its degradation.
[0056] In some embodiments, registering the model inference operator into a deep learning training framework and using the model inference operator to replace the corresponding operator in the deep learning training framework includes: registering the model inference operator into a deep learning training framework, encapsulating the model inference operator, and replacing the corresponding operator in the deep learning training framework with the encapsulated model inference operator; wherein, the interface of the encapsulated model inference operator is compatible with the corresponding interface in the deep learning training framework.
[0057] In some embodiments, fine-tuning the post-injection quantization model to obtain a fine-tuning parameter result includes: fine-tuning the post-injection quantization model on a dedicated task to obtain a fine-tuning parameter result for the dedicated task.
[0058] Proprietary tasks are typically those specific to a particular domain or application. These tasks are often customized and require the model to understand and process specific types of data or problems. Proprietary tasks are designed for specific business needs or research objectives and are not universally applicable.
[0059] Fine-tuning the injected quantization model on different specialized tasks can yield multiple sets of fine-tuning parameter results corresponding to the multiple specialized tasks. When performing inference on a specialized task, applying the fine-tuning parameter results corresponding to the specialized task can improve the accuracy of the inference results.
[0060] In some embodiments, loading the fine-tuning parameter results into the quantization model includes: using a deep learning compiler to load the fine-tuning parameter results into the quantization model, thereby disabling the backpropagation function of the quantization model while retaining the forward propagation function of the quantization model, thereby facilitating the deployment of the quantization model on platforms with limited resources, such as mobile terminals such as mobile phones. The deep learning compiler may be the aforementioned TVM.
[0061] In some embodiments, in the process of implementing the key operator using a model inference framework, the central processing unit is bound as the computing device of the key operator; wherein the computing device of the deep learning training framework is a graphics card (also called a GPU).
[0062] Since the model inference framework tvm supports multiple devices, when using the model inference framework to implement the key operators, the computing device (such as CPU) for operating the key operators can be pre-specified. In the subsequent fine-tuning process, all key operators implemented using the model inference framework will be calculated using this computing device to implement the offload mechanism of additional devices and reduce the usage of video memory.
[0063] In addition, during training, the above method, also known as a memory offload strategy, can save video memory resources by storing frozen weight parameters in memory. At the same time, the inference framework can provide reliable performance support during the forward propagation process of training.
[0064] Figure 2 This is a flow chart of a method for correcting quantization model degradation according to the second embodiment of the present invention. Figure 2 As shown, the method of this embodiment may include: Step 201: Obtain key operators of the original model.
[0065] This embodiment can obtain key operators of the key layers of the original model.
[0066] The key layers of the original model selected are the Embedding layer and the Linear layer. The Embedding layer is a layer that converts discrete data (such as words, characters, or any category labels) into a continuous vector representation. This conversion enables the model to capture complex patterns and relationships in the data. The Linear layer, also known as the fully connected layer or dense layer, is a basic neural network layer whose main function is to perform linear transformations. The Linear layer maps the input data through a weight matrix, then adds a bias vector and outputs new data. This process can be expressed mathematically as:
[0067] in, represents the output vector, represents the input vector, represents the weight matrix, represents the weight matrix, Represents the bias vector.
[0068] The key operators selected in this embodiment may be the rope operator and the sdpa operator. The rope operator is a relative position encoding method that is integrated into the self-attention mechanism to enhance the performance of the architecture. The core of the rope operator is to encode the relative position information of words in a sequence through rotation transformation, allowing the model to capture the relative distance relationship between words. The sdpa operator is an optimized self-attention mechanism that uses a specific kernel implementation to improve computational efficiency, especially on GPUs.
[0069] Step 202: Use TVM to implement the key operator and obtain a quantized version of the TVM inference operator.
[0070] TVM is an open-source machine learning compiler framework that allows machine learning engineers to efficiently optimize and run computations on a variety of hardware backends. TVM aims to provide optimizations throughout the entire process, from model definition to execution on hardware. It supports importing models from front-end frameworks and converting them into TVM's high-level model language, Relax. Dynamic shape refers to situations where the input shape of a model is determined only at runtime. TVM supports dynamic shape through the Relax IRModule format, allowing models to run with varying input sizes. TVM's dynamic shape support is particularly important for models in fields such as natural language processing (NLP), which often need to process sequences of varying lengths. TVM supports a variety of devices, including CPUs, GPUs, and machine learning accelerators. Users can specify the device on which operators should execute at compile time, allowing TVM to optimize operator performance to suit specific hardware characteristics. Furthermore, TVM can offload graphics cards by compiling modules to other hardware, migrating computational tasks from the host computer to other devices, reducing the host's computational burden and memory usage.
[0071] Step 203: Register the TVM operator in PyTorch, and use the TVM operator to replace the corresponding operator in PyTorch to obtain a quantized model.
[0072] The key layers in the quantized model are the quantized version key layers, which include the mbedding layer and the Linear layer.
[0073] The quantized version of the Embedding layer and the quantized version of the Linear layer both contain quantization functions, dequantization functions, and forward propagation functions. The following will introduce the quantization function, forward propagation function, and dequantization function in detail one by one: Quantization functions refer to a series of functions and processes used to convert the weights and activations of a model from floating point numbers to low-precision representations (such as int8). Generally, quantization functions are used in the model quantization process, which converts high-precision models (fp32, fp16) to low-precision (int8, int4).
[0074] The forward propagation function is a core concept in deep learning models, defining how data flows through network layers and produces output. In neural networks, the forward propagation function describes the data transformation process from the input layer to the output layer. The forward propagation function is a mixed-precision matrix multiplication. The input is typically FP16 or FP32, and the weights are usually in INT4 or INT8. The forward propagation requires a dequantization function to restore the precision.
[0075] The dequantization function converts quantized values back to their original floating-point state before quantization. In deep learning and model optimization, quantization is a technique for reducing model size and accelerating inference. It is achieved by converting the model's weights and activation values from floating-point numbers to integers. However, the quantization process may introduce some loss of accuracy, so the dequantization function is used to convert these quantized values back to floating-point numbers when needed (such as when analyzing or visualizing the model).
[0076] When using the TVM operators to replace the corresponding operators in PyTorch, for Embedding and Linear layers, you can directly assign the registered TVM operators to the Linear and Embedding layers that come with PyTorch, overwriting the original implementation. For the attention mechanism of large models, the position encoding and attention mechanism calculation parts are replaced with the rope operator and the sdpa operator.
[0077] When using the TVM operator to replace the corresponding operator in PyTorch, it is necessary to adapt the TVM operator to the PyTorch interface. This means ensuring that the custom operator (TVM operator) can be seamlessly integrated into the PyTorch framework and work with other parts of PyTorch. Specifically, the following steps are required to adapt the TVM operator to the PyTorch interface: (1) Understand the interface of PyTorch operators, including the input and output tensor formats, data types, and device types (such as CPU or CUDA).
[0078] (2) If the pytorch operator is not built-in to pytorch, you need to use pytorch's C++ extension or Python wrapper to implement the custom operator.
[0079] (3) Custom operators need to be registered in PyTorch so that they can be recognized and called. This can be achieved by inheriting torch.autograd.Function and defining the forward and backward static methods in Python. For the specific case of this invention, only the forward method needs to be implemented.
[0080] (4) The interface (input and output parameters) of the TVM operator needs to be consistent with the interface in PyTorch to ensure that the TVM operator can replace the corresponding part in PyTorch without causing errors.
[0081] Step 204: Inject the low-rank weights of LoRa into each linear layer of the quantization model to obtain the injected quantization model.
[0082] Step 2041: Package the linear layer in the replaced model into a linear layer with lora low-rank weights.
[0083] Specifically, first define a new class, which will wrap the original linear layer and add the functionality of Lora low-rank weights. This class will contain all the properties of the original linear layer, such as weights and biases, as well as additional parameters required by Lora; secondly, in this new class, two new parameters need to be initialized, namely Lora's low-rank matrices A and B. The ranks of these two matrices are much smaller than the dimensions of the original weight matrix, and they will be used to adjust the original weights; again, define a forward propagation method for this new class, which will be called when the data passes through this wrapper layer. In this method, we first calculate the Lora low-rank weights, that is, multiply the matrices A and B, and then add the result to the original weights to form new weights; finally, in the forward propagation method, we use the updated weights (that is, the original weights plus the weights adjusted by Lora) to perform a linear transformation and then output the result.
[0084] Step 2042: traverse all linear layers in the replaced model and replace each linear layer with a linear layer with lora weights.
[0085] Specifically, first write a function that will traverse all layers in the model and look for linear layers, usually by recursively checking each submodule of the model; secondly, when a linear layer is found, use the wrapper class defined in the above steps to create a new wrapper layer instance. This new wrapper layer will contain the adjustment of lora low-rank weights; again, when the new wrapper layer is created, it can be used to replace the original linear layer; finally, during the replacement process, make sure to retain other states of the model, such as the weights and biases of other layers, and the overall architecture of the model.
[0086] Step 205: Fine-tune the post-injection quantization model on the dedicated task to obtain fine-tuning parameter results.
[0087] Specifically, first determine the specific task that the injected quantized model needs to perform and prepare the corresponding dataset, which is closely related to the task. For example, if it is a text classification task, a dataset containing texts of different categories is required; secondly, ensure that the injected model has been injected with lora low-rank weights and is ready for fine-tuning, that is, the injected model has been loaded into an appropriate deep learning framework, such as PyTorch or TensorFlow; again, determine the parameters in the fine-tuning process, including learning rate, batch size, training cycle (epoch), etc.; finally, use the dataset of the proprietary task to fine-tune the injected model and obtain the fine-tuning parameter results.
[0088] Step 206: Load the fine-tuning parameter results into the quantization model to correct its degradation.
[0089] Here, the quantized model can be obtained by quantizing the model, that is, converting the model's weights and activations from floating-point numbers to low-precision representations, such as int8 or int4, which can reduce the size of the model, speed up inference, and reduce energy consumption on specific hardware.
[0090] Post-quantization model degradation refers to the phenomenon in which model performance may degrade due to the conversion of weights and activation values from floating-point numbers to lower-precision representations (such as integers) during the model quantization process. This performance degradation may manifest as a decrease in accuracy, an increase in inference errors, or degradation of other evaluation metrics.
[0091] This embodiment provides a method for correcting degradation of a quantization model. After obtaining the key operators of the original model, TVM is used to implement the key operators to obtain a quantized version of the TVM operator; the TVM operator is registered in pytorch, and the TVM operator is used to replace the corresponding operator in pytorch to obtain a quantized model; fine-tuning low-rank weights are injected into each linear layer of the quantization model to obtain an injected quantization model; then a dedicated task or a specific task is used to train and fine-tune the parameters of the injected quantization model to obtain a fine-tuning parameter result with higher accuracy. That is to say, in the process of training the quantization model, the feedback of the quantized inference model (including the quantized version of the TVM operator) must be used to make the fine-tuning parameters obtained by training more accurate; the fine-tuning parameter result with higher accuracy is loaded onto the quantization model to correct its degradation.
[0092] The following is an example of an embodiment of the present invention: For intent recognition, we used the PEFT library to fine-tune LoRa or QLoRa on the original model. We then used PEFT to load the LoRa weights onto the original model, which we call Model 1. We then loaded the LoRa weights onto the model quantized using MLCLLM, which we call Model 2. When MLCLLM uses TVM as the underlying runtime, the recognized intent for the same user input is shown in the following table:
[0093] As shown in Table 1, Model 2 has incorrect intent predictions, indicating that the mlcllm quantization model causes degradation in the intent recognition task. Quantization can lead to accuracy differences compared to the original model, as well as to accuracy differences caused by differences in the peft quantization implementation.
[0094] The solution described in this embodiment of the present invention enables the network (fine-tuned parameters) fine-tuned by LoRa to be loaded onto the quantized model without loss of quality, ensuring that the prediction results after training do not degrade due to deployment. That is, for the same user input, consistent output is produced after fine-tuning and after deployment via MLCLLM.
[0095] This embodiment provides a method for correcting degradation of a quantization model. After obtaining the key operators of the original model, a model inference framework is used to implement the key operators to obtain a quantized version of the model inference operator; the quantized version of the model inference operator is registered in a deep learning training framework, and the model inference operator is used to replace the corresponding operator in the deep learning training framework, so that a quantized model can be obtained; fine-tuning low-rank weights are injected into each linear layer of the quantization model to obtain an injected quantization model; then a dedicated task or a specific task is used to train and fine-tune the parameters of the injected quantization model to obtain a fine-tuning parameter result with higher accuracy. That is to say, in the process of training the quantization model, the feedback of the quantized inference model (including the quantized version of the model inference operator) must be used to make the fine-tuning parameters obtained by training more accurate; the fine-tuning parameter result with higher accuracy is loaded onto the quantization model to correct its degradation.
[0096] Since the model inference framework supports multiple devices, in the process of implementing the key operators using the model inference framework, the computing device (such as CPU) for operating the key operators can be pre-specified. In the subsequent fine-tuning process, all key operators implemented using the model inference framework will be calculated using this computing device to implement the offload mechanism of additional devices and reduce the usage of video memory.
[0097] Figure 3 FIG. 1 is a schematic diagram showing the structure of a quantization model degradation correction device according to the third embodiment of the present invention. Figure 3 As shown, the device of this embodiment may include: An operator implementation unit 31 is used to implement the key operators of the original model using a model inference framework to obtain a quantized version of the model inference operator; A registration unit 32 is used to register the model inference operator in the deep learning training framework and use the model inference operator to replace the corresponding operator in the deep learning training framework to obtain a quantized model; An injection unit 33 is configured to inject a fine-tuned low-rank weight into each linear layer of the quantization model to obtain an injected quantization model; A fine-tuning unit 34 is used to fine-tune the post-injection quantization model to obtain a fine-tuning parameter result; The loading unit 35 is used to load the fine-tuning parameter result into the quantization model to correct its degradation.
[0098] The quantization model degradation correction device of this embodiment, after obtaining the key operators of the original model, adopts the model inference framework to implement the key operators to obtain a quantized version of the model inference operator; registers the quantized version of the model inference operator into the deep learning training framework, and uses the model inference operator to replace the corresponding operator in the deep learning training framework, so as to obtain a quantized model; injects fine-tuned low-rank weights into each linear layer of the quantization model to obtain an injected quantization model; then uses a dedicated task or a specific task to train and fine-tune the parameters of the injected quantization model to obtain a fine-tuning parameter result with higher accuracy. That is to say, in the process of training the quantization model, the feedback of the quantized inference model (including the quantized version of the model inference operator) must be used to make the fine-tuning parameters obtained by training more accurate; the fine-tuning parameter result with higher accuracy is loaded onto the quantization model to correct its degradation.
[0099] In some embodiments, the operator implementation unit is specifically used to implement the key operator using a deep learning compiler to obtain a quantized version of the model inference operator.
[0100] In some embodiments, the registration unit is specifically used to register the model inference operator into the deep learning training framework, encapsulate the model inference operator, and use the encapsulated model inference operator to replace the corresponding operator in the deep learning training framework; wherein, the interface of the encapsulated model inference operator is compatible with the corresponding interface in the deep learning training framework.
[0101] In some embodiments, the registration unit registers the model inference operator into a deep learning training framework, and uses the model inference operator to replace the corresponding operator in the deep learning training framework, including: registering the model inference operator into the deep learning training framework, encapsulating the model inference operator, and replacing the corresponding operator in the deep learning training framework with the encapsulated model inference operator; wherein, the interface of the encapsulated model inference operator is compatible with the corresponding interface in the deep learning training framework.
[0102] In some embodiments, the fine-tuning unit is specifically configured to fine-tune the post-injection quantization model on a dedicated task to obtain a fine-tuning parameter result for the dedicated task.
[0103] In some embodiments, the loading unit is specifically configured to load the fine-tuning parameter results onto the quantization model using a deep learning compiler.
[0104] In some embodiments, in the process of the operator implementation unit implementing the key operator using the model inference framework, the central processing unit is bound as the computing device of the key operator; wherein the computing device of the deep learning training framework is a graphics card.
[0105] The above-mentioned device embodiment can implement the aforementioned quantization model degradation correction method embodiment. Its implementation process and technical effects are basically the same and will not be repeated here.
[0106] Figure 4 This is a schematic diagram of the structure of an embodiment of the electronic device of the present invention, which can realize the present invention. Figure 1 and Figure 2 The process of the embodiment shown is as follows: Figure 4 As shown, the above-mentioned electronic device may include: a processor 52 and a memory 53, wherein the memory 53 is used to store executable program code; the processor 52 runs the program corresponding to the executable program code by reading the executable program code stored in the memory 53, so as to execute the quantitative model degradation correction method described in any of the above-mentioned embodiments.
[0107] For details on the specific execution process of the above steps by the processor 52 and the steps further executed by the processor 52 by running the executable program code, please refer to the present invention. Figure 1 and Figure 2 The description of the illustrated embodiment will not be repeated here.
[0108] This electronic device exists in many forms, including but not limited to: (1) Mobile communication devices: These devices are characterized by their mobile communication capabilities and are primarily designed to provide voice and data communications. These terminals include smartphones (e.g., iPhones), multimedia phones, feature phones, and low-end phones.
[0109] (2) Personal computer devices: These devices fall under the category of personal computers, have computing and processing capabilities, and generally also have mobile Internet access features.
[0110] (3) Server: A device that provides computing services. The server consists of a processor, hard disk, memory, system bus, etc. The server is similar to a general computer architecture, but because it needs to provide highly reliable services, it has higher requirements in terms of processing power, stability, reliability, security, scalability, and manageability.
[0111] (4) Other electronic devices with data interaction functions.
[0112] This embodiment provides an electronic device that can implement the aforementioned embodiment of the quantization model degradation correction method. Its implementation process and technical effects are basically the same and will not be repeated here.
[0113] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by one or more processors, the above-mentioned fine-tuning network degradation correction method on the quantization model is implemented.
[0114] It should be noted that in some embodiments of the present invention, the computer-readable medium described above may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable storage media may include, but are not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In some embodiments of the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or component. Furthermore, in some embodiments of the present invention, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. This propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wire, optical cable, RF (radio frequency), or any suitable combination thereof.
[0115] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0116] The computer-readable medium may be included in the apparatus, or may exist independently and not incorporated into the electronic device. The computer-readable medium carries one or more programs. When executed by the electronic device, the electronic device: obtains the calling functions of the key layers and key operators of the original model; registers the calling functions with PyTorch to obtain the PyTorch operators; replaces the corresponding parts of the original model with the PyTorch operators to obtain the replaced model; injects LoRa's low-rank weights into each linear layer of the replaced model to obtain the injected model; fine-tunes the injected model on a dedicated task to obtain the fine-tuned parameter results; and loads the fine-tuned parameter results into the quantized model to correct its degradation.
[0117] Computer program code for performing the operations of some embodiments of the present invention may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0118] This embodiment provides a computer-readable storage medium for correcting degradation of a fine-tuned network on a quantized model. This medium can eliminate errors caused by fine-tuning weights on a quantized model loaded with different runtimes, different quantized versions, and a non-quantized version. That is, the model weights generated after fine-tuning can be loaded onto the quantized model without loss of precision, thereby correcting the degradation of the fine-tuned network on the quantized model under different runtimes. At the same time, this embodiment of the present invention supports decoupling video memory resources during training, allowing partial memory to be used for offloading, allowing large models to be efficiently fine-tuned on machines with small video memory.
[0119] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0120] Each embodiment in this specification is described in a related manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
[0121] In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0122] For the convenience of description, the above device is described as being divided into various units / modules based on their functions. Of course, when implementing the present invention, the functions of each unit / module can be implemented in the same or multiple software and / or hardware.
[0123] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0124] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A method for correcting quantization model degradation, characterized in that: include: Get the key operators of the original model; Implement the key operators using a model inference framework to obtain a quantized version of the model inference operator; Registering the model inference operator in a deep learning training framework, and using the model inference operator to replace the corresponding operator in the deep learning training framework to obtain a quantized model; Injecting fine-tuned low-rank weights into each linear layer of the quantized model to obtain an injected quantized model; Fine-tuning the post-injection quantization model to obtain a fine-tuning parameter result; The fine-tuning parameter results are loaded onto the quantization model to correct its degradation.
2. The quantization model degradation correction method according to claim 1, characterized in that: The key operator is implemented using the model reasoning framework to obtain a quantized version of the model reasoning operator, including: A deep learning compiler is used to implement the key operators to obtain a quantized version of the model inference operator.
3. The quantization model degradation correction method according to claim 1, characterized in that: Registering the model inference operator in the deep learning training framework and replacing the corresponding operator in the deep learning training framework with the model inference operator includes: Register the model inference operator in the deep learning training framework, encapsulate the model inference operator, and use the encapsulated model inference operator to replace the corresponding operator in the deep learning training framework; wherein the interface of the encapsulated model inference operator is adapted to the corresponding interface in the deep learning training framework.
4. The quantization model degradation correction method according to claim 1, characterized in that: The fine-tuning of the post-injection quantization model to obtain a fine-tuning parameter result includes: The post-injection quantization model is fine-tuned on a dedicated task to obtain a fine-tuning parameter result for the dedicated task.
5. The quantization model degradation correction method according to claim 1, characterized in that: The step of loading the fine-tuning parameter result onto the quantization model includes: A deep learning compiler is used to load the fine-tuning parameter results onto the quantization model.
6. The quantization model degradation correction method according to claim 1, characterized in that: In the process of implementing the key operator using the model inference framework, the central processing unit is bound as the computing device of the key operator; wherein, the computing device of the deep learning training framework is a graphics card.
7. A quantization model degradation correction device, characterized in that: include: The operator implementation unit is used to implement the key operators of the original model using the model inference framework to obtain the quantized version of the model inference operator; A registration unit, configured to register the model inference operator into a deep learning training framework, and use the model inference operator to replace the corresponding operator in the deep learning training framework to obtain a quantized model; An injection unit, configured to inject fine-tuned low-rank weights into each linear layer of the quantization model to obtain an injected quantization model; A fine-tuning unit, configured to fine-tune the post-injection quantization model to obtain a fine-tuning parameter result; A loading unit is used to load the fine-tuning parameter result onto the quantization model to correct its degradation.
8. The quantization model degradation correction device according to claim 7, characterized in that: The registration unit is specifically used to register the model inference operator into the deep learning training framework, encapsulate the model inference operator, and use the encapsulated model inference operator to replace the corresponding operator in the deep learning training framework; wherein, the interface of the encapsulated model inference operator is adapted to the corresponding interface in the deep learning training framework.
9. An electronic device, characterized in that: The electronic device includes: a processor and a memory; the memory is used to store executable program code; the processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, so as to execute the method described in any of the above claims.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the method according to any one of the preceding claims.
Citation Information
Cited By
Depth learning model output visualization migration consistency analysis method, system and device, medium and product
CN121119141A
A method, system, device, medium and product for analyzing migration consistency of deep learning model output visualization
CN121119141B