Quantification method, system and device of end side large model based on distillation and medium

By adopting a distillation-based quantization method in the end-to-end quantization parameter optimization training of layer-by-layer distillation and self-distillation, the problems of model and chip adaptation and performance losses are solved, and efficient and easy-to-use quantization effect is achieved.

CN120216991APending Publication Date: 2025-06-27BEIJING KNOWLEDGE ATLAS TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510294818.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When existing large language models run on end-side devices, they need to solve the adaptation problem between the model and the chip, and the quantization process leads to a degradation of model performance, making it difficult to find a balance between reducing memory footprint and minimizing performance losses.

Method used

The distillation-based end-side large-model quantization method is adopted, and the end-to-end quantization parameter optimization training is optimized for layer-by-layer distillation and self-distillation, which reduces computing power demand and improves the quantization effect.

Benefits of technology

It effectively avoids the problem of a large amount of computing power demand, improves the practical application, scalability and ease of use of the quantization effect, and significantly improves the operating performance on Qualcomm chips, reducing the quantization loss by more than 90%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216991A_ABST
    Figure CN120216991A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and relates to a distillation-based end side large model quantification method, system and device and a medium, and the quantification method comprises the steps: 1) carrying out simulation quantification on a target large model M composed of N transformer structure layers through weight quantification and activation quantification, 2) based on the target large model M and the simulated and quantified large model # imgabs 1 #, carrying out layer-by-layer distillation on each transformer structure layer of the simulated and quantified large model # imgabs 2 # to obtain a simulated and quantified large model # imgabs 0 # 2; and end-to-end quantization parameter optimization training is carried out on the preliminarily quantized large model # imgabs5 # based on the target large model M and the preliminarily quantized large model # imgabs4 #, so that a finally quantized large model # imgabs6 # can be obtained, and through layer-by-layer distillation and end-to-end quantization parameter optimization training based on self-distillation, the final quantized large model # imgabs6 # can be obtained. The problem that large model quantification needs a large amount of calculation power is avoided, and the method has reliability, expansibility and usability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence, and relates to a quantization method, system, device and medium for large models on the edge side, and particularly relates to a quantization method, system, device and medium for large models on the edge side based on distillation. Background Art

[0002] In recent years, large language models have demonstrated excellent capabilities in various tasks, greatly improving people's work efficiency and life convenience. However, the huge number of parameters and computational requirements of these large models limit their widespread popularity. To address this issue, quantization technology is regarded as an effective strategy, which can reduce the size of the running model and accelerate the inference speed. In particular, to enable large language models to run on edge devices, the adaptation problem between large models and edge chips must be solved. A direct benefit of model quantization is the reduction of video memory occupancy and the alleviation of the burden of running large models on edge devices. However, the quantization process inevitably leads to a decline in model performance. How to minimize the performance loss caused by quantization while reducing video memory occupancy has become a key challenge in the popularization and application of large language models.

[0003] The current mainstream quantization technologies for large language models are mainly divided into two categories: PTQ (Post-Training Quantization, a quantization method performed after model training, where the weights and activations of the model are converted from floating-point numbers to fixed-point numbers (usually int4 or int8) without retraining the model) and QAT (Quantization-Aware Training, a method that simulates the quantization effect during model training, where the weights and activations of the model are in the form of simulated quantization during training, i.e., using floating-point numbers to simulate the behavior after quantization). The PTQ method has attracted much attention in the field of large language models due to its simple process and no need for an additional training process. This method can directly perform quantization after model training, quickly achieving model compression and acceleration. However, the performance loss that PTQ may cause cannot be ignored because it does not fully consider the impact of quantization on model inference accuracy. In contrast, the QAT method has considered the loss caused by quantization during the model training stage, thus training a model with high quantization performance. The advantage of QAT is that it can more accurately simulate the impact of quantization on the model, ensuring that the quantized model has high accuracy. However, its disadvantage is that QAT requires additional training time, and the training process is more complex than traditional floating-point training.

[0004] According to different quantization schemes, the quantization methods of large language models can be further divided into weight-only quantization and weight-activation quantization, where weight-activation quantization can be further divided into static quantization and dynamic quantization, etc. Quantization-related settings include per-channel quantization, per-token quantization, per-tensor quantization, as well as 4-bit and 8-bit quantization, etc. In the PTQ scheme, representative methods include AWQ and GPTQ, and these methods are all weight-only quantization methods, mainly applicable to cloud large model inference. However, in the edge scenario, especially when both weights and activations are int quantized, these methods are difficult to ensure the performance of the quantized model. Representative works of the QAT scheme include LLM-QAT and EfficientQAT, where LLM-QAT has too high computing power requirements and poor generality, while although EfficientQAT has lower computing power requirements, it still requires appropriate fine-tuning data for model training and has limited scalability.

[0005] Therefore, in view of the above-mentioned defects existing in the prior art, it is necessary to develop a new quantization method, system, device and medium for edge large models. Summary of the Invention

[0006] To overcome the defects of the prior art, the present invention proposes a quantization method, system, device and medium for edge large models based on distillation, which avoids the problem of high computing power required for large model quantization through layer-by-layer distillation and end-to-end quantization parameter optimization training based on self-distillation, and has reliability, scalability and ease of use.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] A quantization method for edge large models based on distillation, characterized by comprising the following steps:

[0009] 1) Perform simulated quantization on the target large model M composed of N layers of transformer structural layers through weight optimization and activation quantization to obtain a simulated quantized large model

[0010] 2) Based on the target large model M and the simulated quantized large model Perform layer-by-layer distillation on each layer of the transformer structural layer of the simulated quantized large model To obtain a preliminarily quantized large model

[0011] 3) Based on the target large model M and the preliminarily quantized large model Perform end-to-end quantization parameter optimization training on the preliminarily quantized large model To obtain a finally quantized large model

[0012] Preferably, step 1) specifically includes:

[0013] 11) Quantize the linear transformation weights W and activations X of each layer of the transformer structure layer of the target large model M to obtain quantized weights and quantized activations where

[0014]

[0015] In the formula, represents the rounding operation, n represents the target number of bits for weight quantization input by the user, s represents the scaling factor for weight quantization, z represents the zero point for weight quantization, n' represents the target number of bits for activation quantization input by the user, s' represents the scaling factor for activation quantization, and z' represents the zero point for activation quantization;

[0016] 12) Based on the quantized weights and quantized activations obtain dequantized weights and dequantized activations where

[0017]

[0018] 13) Take s' = abs(X) / (2 n′ -1), z' = 0, calculate to obtain the quantized activation and dequantized activation

[0019] 14) Based on the dequantized weights and dequantized activations obtain the dequantized large model Input the same input into the target large model M and the dequantized large model respectively and compare the error between their outputs. Optimize the linear transformation weight W, the scaling factor s for weight optimization, and the zero point z for weight optimization based on the error between their outputs; and obtain the optimized quantized weights

[0020] 15) Based on the optimized quantized weights and quantized activations obtain the simulated quantized large model

[0021] Preferably, step 2) specifically includes:

[0022] 21) Start a layer-by-layer distillation with N iterations. Before the i-th iteration, check if i is greater than N. If it is, end the iteration. If not, proceed to step 22) and perform distillation on the i-th layer Transformer structure layer of the simulated quantized large model ;

[0023] 22) Input the input vector of the i-th layer into the i-th layer Transformer structure layer of the target large model M and the i-th layer Transformer structure layer of the simulated quantized large model respectively, and obtain the output vector Y of the i-th layer Transformer structure layer of the target large model M and the output vector of the i-th layer Transformer structure layer of the simulated quantized large model

[0024] 23) Calculate the error loss between the output vector and Y through the following formula:

[0025]

[0026] where the error loss calculation function L calculates the error loss using the mean squared error loss;

[0027] 24) Utilize the error loss to optimize the parameters of the i-th layer Transformer structure layer of the simulated quantized large model through backpropagation to obtain the preliminarily quantized large model

[0028] Preferably, step 3) specifically includes:

[0029] 31) Provide the same input for the preliminarily quantized large model and the target large model M, and collect their final outputs and P;

[0030] 32) Use the KL divergence as the loss function to calculate the divergence between the final outputs and P as the loss;

[0031] 33) Utilize the loss to optimize the parameters of the preliminarily quantized large model through backpropagation to obtain the finally quantized large model

[0032] In addition, the present invention also provides a quantization system for an edge-side large model based on distillation, which is characterized in that it includes:

[0033] A simulation quantization module, which is used to perform simulation quantization on a target large model M composed of N layers of transformer structure layers to obtain a simulated quantized large model

[0034] A layer-by-layer distillation module, which is used to perform layer-by-layer distillation on each layer of the transformer structure layer of the simulated quantized large model based on the target large model M and the simulated quantized large model on the simulated quantized large model to obtain a preliminarily quantized large model

[0035] An end-to-end quantization parameter optimization training module, which is used to perform end-to-end quantization parameter optimization training on the preliminarily quantized large model based on the target large model M and the preliminarily quantized large model on the preliminarily quantized large model to obtain a finally quantized large model

[0036] Moreover, the present invention also provides a quantization device for an edge large model based on distillation, which is characterized by including:

[0037] One or more processors;

[0038] A memory for storing one or more programs;

[0039] When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the quantization method for the edge large model based on distillation as described above.

[0040] Finally, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and is characterized in that when the program is executed by a processor, the steps of the quantization method for the edge large model based on distillation as described above are implemented.

[0041] Compared with the prior art, the quantization method, system, device and medium for the edge large model based on distillation of the present invention have one or more of the following beneficial technical effects:

[0042] 1. The present invention innovatively incorporates activation quantization into the self-distillation quantization training process, ensuring a high degree of matching between the quantization effect of the model and the actual chip operation, thereby improving the practical applicability of the quantization effect.

[0043] 2. Different from the traditional QAT method, the present invention does not need to rely on a specific fine-tuning data set. It only needs to use data in a specific business scenario and an appropriate amount of publicly available fine-tuning data sets to learn quantization parameters, greatly enhancing the scalability and applicability of the method.

[0044] 3. The present invention fully considers usability and significantly reduces the computing power requirements during the quantization process through the method of layer-by-layer distillation. In addition, the final supplementary training stage (end-to-end quantization parameter optimization training based on self-distillation) can also achieve the expected effect with only a small amount of data. Generally speaking, the computing power requirements of the present invention are relatively low.

[0045] 4. In terms of the implementation effect, the present invention significantly improves the running performance of the quantized model on Qualcomm chips. Compared with Qualcomm's official quantization scheme, on relevant evaluation datasets, the quantization loss of the present invention is reduced by more than 90%, fully demonstrating its superiority in terms of effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 is a flowchart of the quantization method for the distillation-based edge large model of the present invention.

[0047] Figure 2 is a schematic diagram of the composition of the quantization system for the distillation-based edge large model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] Before detailing any embodiment of the present invention, it should be understood that in its application, the present invention is not limited to the construction and arrangement details of the components described in the following description or illustrated in the following drawings. The present invention can have other embodiments and can be practiced or carried out in various ways. Additionally, it should be understood that the wording and terms used herein are for the purpose of description and should not be considered restrictive. As used herein, the terms "including" or "having" and their variants are intended to cover the items listed hereinafter and their equivalents as well as additional items. Unless otherwise specified or limited, the terms "mounted", "connected", "supported", and "coupled" and their variants are used widely and cover direct mounting and indirect mounting, connection, support, and coupling. In addition, "connection" and "coupling" are not limited to physical or mechanical connection or coupling.

[0049] And, on the one hand, in the disclosure of the present invention, the orientation or positional relationship indicated by the terms "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the above terms should not be construed as limiting the present invention; on the other hand, the term "one" should be understood as "at least one" or "one or more". That is, in one embodiment, the number of an element can be one, while in other embodiments, the number of the element can be multiple. The term "one" should not be construed as limiting the quantity.

[0050] In view of the problems existing in the prior art, the present invention proposes a quantization method for end-side large models based on distillation, which is completely based on self-distillation technology. First, preliminary quantization parameter training is carried out through layer-by-layer distillation of the simulated quantization model and the floating-point model (i.e., the target model); then, end-to-end quantization parameter optimization training based on self-distillation is adopted, and finally ideal quantization parameters are obtained. This method based on distillation throughout effectively avoids the problem of relying on fine-tuning data, and the methods of layer-by-layer distillation and end-to-end quantization parameter optimization training based on self-distillation avoid the problem of requiring a large amount of computing power, and have reliability, scalability and ease of use.

[0051] Figure 1 The flowchart of the quantization method for end-side large models based on distillation according to the present invention is shown. As Figure 1 shown, the quantization method for end-side large models based on distillation according to the present invention includes the following steps:

[0052] I. Simulated quantization.

[0053] Considering that the method of the present invention needs to be applicable to specific end-side computing chips, such as MTK, intel, Snapdragon and other chips, and these chips only support the method of activation quantization. Therefore, the activation quantization method is also introduced in the forward propagation process of the present invention. So, in the present invention, both weight quantization and activation quantization are adopted to perform simulated quantization on the target large model M composed of N layers of transformer structure layers to obtain the simulated quantization large model Specifically, it includes:

[0054] 1. Quantize the linear transformation weights W and activations X of each layer of the transformer structure layer of the target large model M to obtain quantized weights and quantized activations where

[0055]

[0056] In the formula, represents the rounding operation, n represents the target number of bits of weight quantization input by the user, s represents the scaling factor of weight quantization, z represents the zero point of weight quantization, n′ represents the target number of bits of activation quantization input by the user, s′ represents the scaling factor of activation quantization, and z′ represents the zero point of activation quantization.

[0057] 2. Based on the quantized weights and quantized activations respectively obtain dequantized weights and dequantized activations where

[0058]

[0059] 3. Take \(s'=\frac{|X|}{2 - 1}\), \(z' = 0\), and calculate the quantized activation n′ and the dequantized activation

[0060] In the present invention, since it is a large model on the edge side, when performing activation quantization, a relatively widely used and easy-to-use dynamic quantization method for activation is considered. Therefore, each time the quantized activation and the dequantized activation are calculated, \(s'=\frac{|X|}{2 - 1}\) and \(z' = 0\) can be taken. In this way, it is not necessary to train and learn the scaling factor \(s'\) of activation quantization and the zero point \(z'\) of activation quantization. n′

[0061] 4. Based on the dequantized weight and the dequantized activation obtain the dequantized large model That is, the dequantized large model is a large model whose linear transformation weight is the dequantized weight and whose activation is the dequantized activation .

[0062] Then, input the same input into the target large model \(M\) and the dequantized large model respectively and compare the error between their outputs. Optimize the linear transformation weight \(W\), the scaling factor \(s\) of weight optimization, and the zero point \(z\) of weight optimization based on the error between their outputs.

[0063] In the present invention, when optimizing the linear transformation weight \(W\), the scaling factor \(s\) of weight optimization, and the zero point \(z\) of weight optimization, the same as the traditional QAT method, the updated parameters can be obtained through forward propagation and backward propagation, that is, the optimized linear transformation weight \(W\), the scaling factor \(s\) of weight optimization, and the zero point \(z\) of weight optimization. Since the specific method of obtaining the updated parameters is the same as the traditional QAT method and is not the focus of the present invention, for the sake of simplicity, it will not be described in detail here.

[0064] Finally, based on the optimized linear transformation weight \(W\), the scaling factor \(s\) of weight optimization, and the zero point \(z\) of weight optimization, obtain the optimized quantized weight That is, after obtaining the optimized linear transformation weight \(W\), the scaling factor \(s\) of weight optimization, and the zero point \(z\) of weight optimization, the optimized quantized weight can be obtained based on the formula in step 1

[0065] 5. Based on the optimized quantized weight and the quantized activation ​​Obtain the large model of the simulated quantization

[0066] With the optimized quantization weights And the quantization activation obtained in step 3 Based on the optimized quantization weights And the quantization activation Obtain the large model of the simulated quantization That is, the large model of the simulated quantization Has a linear transformation weight of the optimized quantization weight And the activation is the quantization activation Of the large model

[0067] II. Layer-by-layer distillation

[0068] After obtaining the large model of the simulated quantization Based on the target large model M and the large model of the simulated quantization Perform layer-by-layer distillation on each layer of the transformer structure layer of the large model of the simulated quantization To obtain the preliminarily quantized large model

[0069] The layer-by-layer distillation method adopted in the present invention is mainly to solve the problem of excessive computational overhead of the traditional QAT method. By gradually distilling each layer of the transformer structure layer of the large model, the effect obtained by the traditional QAT method can be finally approximated

[0070] In the present invention, the layer-by-layer distillation specifically includes the following steps

[0071] 1. Start a layer-by-layer distillation including N iterations, where N is the number of layers of the transformer structure layer included in the target large model M. Iterate from the first time to the Nth layer, and distill one layer of the transformer structure layer of the large model of the simulated quantization each time. Through N-layer iteration, the distillation of all transformer structure layers of the large model of the simulated quantization can be realized Of the large model of the simulated quantization Of all transformer structure layers

[0072] Among them, before the i-th iteration, check whether i is greater than N times. If it is greater than N, it means that the distillation of all transformer structure layers of the large model of the simulated quantization Has ended, and the iteration ends. If it is not greater than N, enter step 2 to distill the i-th layer of the transformer structure layer of the large model of the simulated quantization Of the large model of the simulated quantization

[0073] 2. Input the input vectors of the i-th layer into the i-th layer Transformer structure layer of the target large model M and the i-th layer Transformer structure layer of the simulated quantization large model respectively, and obtain the output vector Y of the i-th layer Transformer structure layer of the target large model M and the output vector of the i-th layer Transformer structure layer of the simulated quantization large model respectively. In the present invention, since the linear transformation weight W and the activation are quantized in the simulation quantization stage, therefore, the output vector Y of the i-th layer Transformer structure layer of the target large model M = F(W, X), and the output vector of the i-th layer Transformer structure layer of the simulated quantization large model where F can abstractly represent the mapping function of a layer of Transformer structure layer.

[0074] In the present invention, since the linear transformation weight W is quantized and the activation is also quantized in the simulation quantization stage, therefore, the output vector Y of the i-th layer Transformer structure layer of the target large model M = F(W, X), and the output vector of the i-th layer Transformer structure layer of the simulated quantization large model where F can abstractly represent the mapping function of a layer of Transformer structure layer. where F can abstractly represent the mapping function of a layer of Transformer structure layer.

[0075] 3. Calculate the error loss between the output vector and Y through the following formula:

[0076]

[0077] where the error loss calculation function L calculates the error loss using the mean square error loss.

[0078] 4. Use the error loss to optimize the parameters of the i-th layer Transformer structure layer of the simulated quantization large model through backpropagation. After distilling and iterating layer by layer through all N layers of Transformer structure layers, the preliminarily quantized large model is obtained.

[0079] After distilling and iterating layer by layer through all N layers of Transformer structure layers, the preliminarily quantized large model is obtained.

[0080] III. End-to-end quantization parameter optimization training.

[0081] Since the quantization effect of the obtained preliminarily quantized large model still has room for improvement, therefore, the present invention considers further applying end-to-end global training based on self-distillation. That is, based on the target large model M and the preliminarily quantized large model perform end-to-end quantization parameter optimization training on the preliminarily quantized large model to obtain the finally quantized large model Specifically, it includes:

[0082] 1. For the preliminarily quantized large model Provide the same input to the target large model M and collect their final outputs and P.

[0083] In the present invention, due to the introduction of activation quantization, and P = H(W, X). The function H here represents the abstract mathematical mapping function corresponding to the target model M itself (i.e., the N-layer transformer structure layer).

[0084] 2. Use the KL divergence as the loss function Calculate the final output and the divergence between P as the loss.

[0085] In the present invention, the specific calculation method of the loss is as follows:

[0086]

[0087] 3. Use the loss to optimize the parameters of the preliminarily quantized large model through backpropagation of to obtain the final quantized large model The final quantized large model contains the finally learned weights and parameters and

[0088] So far, the quantization of the target large model M is completed, and the final quantized large model is obtained

[0089] Figure 2 shows a schematic diagram of the composition of the quantization system of the end-side large model based on distillation of the present invention. As Figure 2 shown, the quantization system of the end-side large model based on distillation of the present invention includes:

[0090] 1. Analog quantization module.

[0091] The analog quantization module is used to perform analog quantization on the target large model M composed of N-layer transformer structure layers through weight quantization and activation quantization to obtain an analog quantized large model

[0092] 2. Layer-by-layer distillation module.

[0093] The layer-by-layer distillation module is used to perform layer-by-layer distillation on each layer of the transformer structure layer of the analog quantized large model based on the target large model M and the analog quantized large model to obtain a preliminarily quantized large model ​

[0094] 3. End-to-end quantization parameter optimization training module.

[0095] The end-to-end quantization parameter optimization training module is used to perform end-to-end quantization parameter optimization training on the preliminarily quantized large model based on the target large model M for the preliminarily quantized large model to obtain the finally quantized large model

[0096] In addition, the present invention also relates to a quantization device for a large model at the edge based on distillation, which includes: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the quantization method for the large model at the edge based on distillation as described above.

[0097] Finally, the present invention also relates to a computer-readable storage medium, on which a computer program is stored, and characterized in that when the program is executed by a processor, the steps of the quantization method for the large model at the edge based on distillation as described above are implemented.

[0098] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting the protection scope of the present invention. Those skilled in the art can modify or equivalently replace the technical solutions of the present invention according to the idea of the present invention, without departing from the essence and scope of the technical solutions of the present invention.

Claims

1. A quantization method for a large end-side model based on distillation, characterized in that: The following steps are involved: 1) The target large model M consisting of N layers of transformer structure layers is simulated and quantized through weight quantization and activation quantization to obtain a simulated quantized large model 2) Based on the target large model M and the simulated quantized large model A large model for quantifying the simulation Each layer of the transformer structure is distilled layer by layer to obtain a large model with preliminary quantization 3) Based on the target large model M and the preliminary quantified large model A large model for the initial quantification Perform end-to-end quantization parameter optimization training to obtain the final quantized large model 2. The method for quantizing a large end-to-end model based on distillation according to claim 1, characterized in that: The step 1) specifically includes: 11) Quantize the linear transformation weight W and activation X of each transformer structure layer of the target large model M to obtain the quantized weight and quantized activation in, In the formula, represents the rounding operation, n represents the target number of bits of weight quantization input by the user, s represents the scaling factor of weight quantization, z represents the zero point of weight quantization, n′ represents the target number of bits of activation quantization input by the user, s′ represents the scaling factor of activation quantization, and z′ represents the zero point of activation quantization; 12) Based on the quantized weight and quantized activation Get the dequantized weights and dequantized activation in, 13) Take s′=abs(X) / (2 n′ -1), z′=0, and the quantized activation is calculated and dequantized activation 14) Based on the dequantized weight and dequantized activation Get a dequantized large model The same input is input into the target large model M and the dequantized large model respectively. and compare the errors between their outputs, optimize the linear transformation weight W, the weight-optimized scaling factor s, and the weight-optimized zero point z based on the errors between their outputs; and obtain the optimized quantization weight based on the optimized linear transformation weight W, the weight-optimized scaling factor s, and the weight-optimized zero point z 15) Based on the optimized quantization weights and the quantized activation Get the large model of simulation quantization 3. The method for quantizing a large end-to-end model based on distillation according to claim 1, characterized in that: The step 2) specifically includes: 21) Start a layer-by-layer distillation including N iterations, check whether i is greater than N before the i-th iteration, if it is greater, end the iteration, if not, go to step 22), and simulate the large model of quantization The i-th transformer structure layer is distilled; 22) Input the input vector of the i-th layer into the i-th layer transformer structure layer of the target large model M and the large model of simulation quantization respectively In the i-th transformer structure layer of the target large model M, the output vector Y of the i-th transformer structure layer and the simulated quantized large model are obtained respectively. The output vector of the i-th Transformer structure layer 23) Calculate the output vector using the following formula And the error loss of Y: Among them, the error loss calculation function L is to calculate the error loss using the mean square error loss; 24) Utilizing the error loss, the large model of simulation quantization is optimized by back propagation The parameters of the i-th layer of the Transformer structure layer are used to obtain the preliminary quantized large model 4. The method for quantizing a large end-to-end model based on distillation according to claim 1, characterized in that: The step 3) specifically includes: 31) is the large model for the preliminary quantification Provide the same input as the target large model M and collect their final output and P; 32) Use KL divergence as the loss function to calculate the final output The divergence with P is used as loss; 33) Using the loss, optimize the preliminary quantized large model through back propagation The parameters of the final quantized large model are obtained 5. A quantization system for a large end-side model based on distillation, characterized in that: include: The simulation quantization module is used to simulate the quantization of the target large model M composed of N layers of transformer structure layers through weight quantization and activation quantization to obtain a simulated quantized large model A layer-by-layer distillation module is used to generate a large model based on the target large model M and the simulated quantized large model. A large model for quantifying the simulation Each layer of the transformer structure is distilled layer by layer to obtain a large model with preliminary quantization An end-to-end quantization parameter optimization training module is used for the target large model M and the preliminary quantized large model. A large model for the initial quantification Perform end-to-end quantization parameter optimization training to obtain the final quantized large model 6. A quantization device for a large end-side model based on distillation, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the distillation-based end-side large model quantization method as described in any one of claims 1 to 4.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the distillation-based end-side large model quantization method as described in any one of claims 1 to 4 are implemented.