Model compression method fusing pruning quantization joint optimization
By combining pruning and quantitative optimization in large language model compression, the problem of misdeletion of important parameters in the existing technology is solved, and the efficient operation of the model in resource-constrained environment is achieved.
Patent Information
- Application Number
- CN202510164883.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-06
AI Technical Summary
When compressing large language models, the existing technology only relies on model pruning technology, resulting in that when a certain pruning rate is reached, important parameters may be accidentally deleted, affecting the performance of the model.
The model compression method of fusion pruning quantization joint optimization is adopted. By obtaining the pre-trained natural language processing model and quantization module, the training corpus is used for joint training, and fusion pruning and quantization optimization is used, so that the impact of quantization on model performance is considered during the training process.
By perceiving the performance of pruning and quantization in joint training in advance, the trainingable parameters are optimized, which significantly reduces the number of parameters and calculation requirements of the model, reduces the storage space usage, and makes it easy to deploy on resource-constrained hardware.
Smart Images

Figure CN120106167A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of neural network technology, specifically to the field of deep learning model compression, and more specifically to a model compression method that integrates pruning and quantization joint optimization. Background Art
[0002] In recent years, with the development of deep learning technology, large language models such as GPT (Generative Pre-trained Transformer), GLM (General Language Model Pretraining with Autoregressive BlankInfilling), LLaMA (Large Language Model Meta AI), and QWen have been increasingly used in various natural language processing tasks such as text generation, machine translation, and question-answering systems due to their powerful expressive power and wide applicability. However, the scale and complexity of these large language models have greatly increased, with the number of parameters reaching billions or even hundreds of billions. Therefore, in resource-constrained scenarios such as edge computing, the deployment of large language models faces severe challenges. To address this issue, many studies have begun to focus on how to effectively compress large language models so that they can run efficiently on resource-constrained devices.
[0003] Existing model compression technology mainly uses model pruning to remove unimportant parameters (such as weights, neurons, or attention heads) to reduce computing and storage requirements. After the model training is completed, the amplitude of the parameters is analyzed to remove weights with less contribution. Or during the training process, the weights or structures to be pruned are dynamically selected according to the gradient or sensitivity. However, for the current natural language processing models with a large number of parameters, only using model pruning technology to remove a large number of parameters, in order to achieve a certain pruning rate, more important parameters may be removed based on the sorting and pruning rate, thereby affecting the performance of the model.
[0004] It should be noted that this background technology is only used to introduce the relevant information of the present invention to help understand the technical solution of the present invention, but it does not mean that the relevant information is necessarily the prior art. The relevant information is submitted and disclosed together with the present invention solution. If there is no evidence that the relevant information has been disclosed before the application date of the present invention, the relevant information shall not be regarded as the prior art. Summary of the invention
[0005] Therefore, the purpose of the present invention is to overcome the defects of the above-mentioned prior art and provide a model compression method that integrates pruning and quantization joint optimization.
[0006] The objective of the present invention is achieved through the following technical solutions:
[0007] According to a first aspect of the present invention, a model compression method is provided, the method comprising: obtaining a model to be pruned, which is a pre-trained natural language processing model and includes multiple trainable parameters; obtaining a quantization module constructed for the model, which quantizes the trainable parameters of the model through learnable quantization parameters to obtain corresponding quantization values, wherein the quantization value corresponding to each trainable parameter is smaller than the data volume of the trainable parameter itself; using training corpus to perform joint training of pruning and quantization optimization on the model and the quantization module to obtain a jointly trained model and quantization module, wherein in forward propagation during training, the quantization values corresponding to the trainable parameters are used to temporarily replace the trainable parameters for calculation, and in back propagation, the trainable parameters and quantization parameters are updated with the goal of minimizing the value of a preset total loss function; pruning the jointly trained model to obtain a pruned model; using the jointly trained quantization module to quantize the trainable parameters in the pruned model to obtain a quantized model. The technical solution of this embodiment can at least achieve the following beneficial technical effects: during the training process, the present invention uses a quantization module to replace the processed data of the trainable parameters with their quantized values, and the quantization parameters in the quantization module are themselves trainable. Through the joint training of the model and the quantization module, the impact of quantization on the prediction performance of the model can be considered during the joint training, so that the performance of the quantized model actually obtained later is guaranteed.
[0008] Optionally, the total loss function of the joint training is a weighted sum of multiple sub-loss functions, and the multiple sub-loss functions include: a first sub-loss function, which is configured to calculate the prediction loss of the model; a second sub-loss function, which is configured to be positively correlated with the mean of the absolute values of the trainable parameters; and a third sub-loss function, which is configured to be positively correlated with the mean of the deviations between the trainable parameters and the dequantization results of their quantized values.
[0009] Optionally, the total loss function is:
[0010]
[0011] in, represents the first sub-loss function, represents the second sub-loss function, represents the third sub-loss function, express The weighting coefficient of express The weighting coefficient of express The weighting coefficient of .
[0012] Optionally, the second sub-loss function is:
[0013]
[0014] in, Represents the total number of trainable parameters, Indicates the number Trainable parameters.
[0015] Optionally, the third sub-loss function is:
[0016]
[0017] in, Represents the total number of trainable parameters, Indicates the number Trainable parameters, express The corresponding quantized value, Represents a dequantization function that performs dequantization using a quantization parameter.
[0018] Optionally, the quantization module uses quantization parameters to truncate, bias and scale the trainable parameters, where:
[0019] The quantization parameters include the bias parameter and the scaling parameter, or
[0020] The quantization parameters include a truncation lower limit parameter, a truncation upper limit parameter, a bias parameter, and a scaling parameter.
[0021] Optionally, the quantization parameter includes a bias parameter and a scaling parameter, wherein the quantization module performs quantization in the following manner:
[0022]
[0023] in, represents the rounding function, represents the truncation function, represents the bias parameter, Represents the scaling parameter.
[0024] Optionally, the process of pruning the jointly trained model includes: obtaining a pruning sensitivity index of each trainable parameter and a set of prunable parameters; from the set of prunable parameters, removing at least part of the trainable parameters of the model in the order of the pruning sensitivity index from small to large, to obtain a pruned model.
[0025] Optionally, an electronic device comprises: one or more processors; and a memory, wherein the memory is used to store executable instructions; the one or more processors are configured to implement the steps of the method described in the first aspect by executing the executable instructions or to perform inference using a quantized model obtained by the method described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The embodiments of the present invention are further described below with reference to the accompanying drawings, in which:
[0027] Figure 1 4 is a flow chart of a model compression method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0028] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below through specific embodiments in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0029] As mentioned in the background technology section, for the current natural language processing models with a huge number of parameters, only using model pruning technology to remove a large number of parameters, in order to achieve a certain pruning rate, more important parameters may be removed based on the sorting and pruning rate, thereby affecting the performance of the model. In this regard, before pruning and quantization, the present invention first uses training corpus to perform joint training of pruning and quantization optimization on the model and quantization module to obtain a jointly trained model and quantization module, thereby integrating pruning optimization and quantization optimization in the process of joint training, thereby allowing the model to perceive the performance after pruning and quantization in advance, thereby optimizing the trainable parameters to adapt to the state after pruning and quantization, and better guaranteeing the performance of the model after pruning and quantization.
[0030] Some specific implementations are given below to illustrate the model compression method of the present invention.
[0031] Implementation Method 1
[0032] According to one embodiment of the present invention, see Figure 1 , a model compression method is provided, which includes steps A1, A2, A3, A4 and / or A5. In order to better understand the present invention, each step is described in detail below in conjunction with specific embodiments.
[0033] Step A1: Obtain a model to be pruned, which is a pre-trained natural language processing model and includes multiple trainable parameters.
[0034] According to one embodiment of the present invention, the pre-trained natural language processing model (hereinafter referred to as the model) can be a full-precision large language model. Usually, the model, as a benchmark for compression, needs to have high initial performance (such as accuracy, perplexity, etc.), for example: QWen or LLaMA (Large Language Model Meta AI) can be used, or it can be a natural language processing model with a structure customized by the implementer and pre-trained.
[0035] Step A2: Obtain a quantization module constructed for the model, which quantizes the trainable parameters of the model through learnable quantization parameters to obtain corresponding quantization values, wherein the quantization value corresponding to each trainable parameter is smaller than the data amount of the trainable parameter itself.
[0036] During the training phase, the quantization module is used to continuously pseudo-quantize the trainable parameters and optimize the self-parameters. The purpose is to learn the quantization parameters used to quantize the trainable parameters so that the model can still maintain good performance after its trainable parameters are truly quantized.
[0037] According to one embodiment of the present invention, the quantization module uses quantization parameters to truncate, bias and scale the trainable parameters, and the quantization parameters include bias parameters and scaling parameters. For each processing layer (or module) in the model, the quantization module sets a set of bias parameters and scaling parameters, and the bias parameters and scaling parameters of different processing layers can be updated independently. In this way, the space occupied by the quantization module and the optimization overhead are reduced. Among them, the quantization module performs quantization in the following manner:
[0038]
[0039] in, represents the rounding function, represents the truncation function, represents the bias parameter, In addition, the quantization bit width of the quantization module can be set as a hyperparameter by the implementer to limit the quantization range of the quantization module.
[0040] Step A3: Use the training corpus to perform joint training of pruning and quantization optimization on the model and quantization module to obtain the jointly trained model and quantization module, wherein in the forward propagation during training, the quantization values corresponding to the trainable parameters are used to temporarily replace the trainable parameters for calculation, and in the back propagation, the trainable parameters and quantization parameters are updated with the goal of minimizing the value of the preset total loss function.
[0041] The training corpus is the corpus used to train the model to perform natural language processing to complete the preset prediction task. It may usually be a natural language text. Depending on the task, it may also be divided into input text and preset answer text. In this field, the training corpus is relatively abundant, and the present invention does not make any limitation to this.
[0042] Conventionally, model training is done in batches. Each batch of corpus is input into the model. The model extracts semantic features through trainable parameters and makes predictions based on the semantic features to obtain prediction results. Based on the prediction results and the preset answer text, the prediction loss can be calculated, and the gradient is calculated through back propagation, and the trainable parameters of the model are updated through the gradient descent method. However, in the method of the present invention, improvements are made to improve the adaptability of the model to pruning and quantization.
[0043] According to one embodiment of the present invention, the joint training is, on the one hand, the joint training of the model and the quantization module, and on the other hand, it is also the joint training of pruning optimization and quantization optimization of the model.
[0044] Joint training involves improvements in the model's data processing and the overall loss function.
[0045] During joint training, when the model processes the input text, the quantization module will be called to pseudo-quantize each trainable parameter based on its latest quantization parameter to obtain the quantized value to replace the trainable parameter to process the data (feature). Because joint training will update the quantization parameters and trainable parameters multiple times, the quantization value is used to temporarily replace (pseudo-quantize) the trainable parameter to process the data during joint training.
[0046] During joint training, the total loss function adds losses related to pruning optimization and quantization optimization in addition to the existing prediction loss, so as to jointly optimize the parameters of the model to adapt to the actual pruned and quantized state in the later stage.
[0047] According to one embodiment of the present invention, the total loss function of the joint training is a weighted sum of multiple sub-loss functions, including: a first sub-loss function configured to calculate the prediction loss of the model; a second sub-loss function configured to be positively correlated with the mean of the absolute value of the trainable parameter; a third sub-loss function configured to be positively correlated with the mean of the deviation between the trainable parameter and the dequantization result of its quantized value. Schematically, the total loss function is:
[0048]
[0049] in, represents the first sub-loss function, represents the second sub-loss function, represents the third sub-loss function, express The weighting coefficient of express The weighting coefficient of express The weighting coefficient of . , and The sizes can be set according to the needs of the implementer, and this embodiment does not limit this. For example, they can be set to 1, 0.2 and 0.3, or 0.7, 0.1, 0.2 respectively.
[0050] The first sub-loss function is a loss function used to guide the optimization of the trainable parameters of the model to improve the preset training task. For example, it is used to guide the model to improve the accuracy and / or perplexity of the prediction. The first sub-loss function can adopt the prediction loss function designed by each existing model (such as Tongyi Qianwen or LLaMA model), and the present invention does not impose any restrictions on this.
[0051] The second sub-loss function is a loss function used to guide the pruning optimization of the trainable parameters of the model so that the trainable parameters of the model tend to zero (0). In other words, the second sub-loss function can be called a pruning regularization term, which is used to control the sparsity of the trainable parameters of the model and guide the pruning process by adding sparse constraints.
[0052] According to one embodiment of the present invention, the second sub-loss function is:
[0053]
[0054] in, Represents the total number of trainable parameters, Indicates the number Trainable parameters.
[0055] The third sub-loss function is a loss function used to guide the pruning optimization of the trainable parameters of the model so that the trainable parameters of the model are closer to the quantization values within the quantization range of the quantization module. In other words, the third sub-loss function is a quantization error term, which is used to minimize the error introduced by the quantization process and optimize the distribution of trainable parameters (including weights and / or activation values) to adapt to the quantization values of low bit width.
[0056] According to one embodiment of the present invention, the third sub-loss function is:
[0057]
[0058] in, Represents the total number of trainable parameters, Indicates the number Trainable parameters, express The corresponding quantized value, represents a dequantization function for dequantization using a quantization parameter, where , represents the bias parameter, Represents the scaling parameter.
[0059] In this embodiment, the training corpus can also be divided into multiple batches. Each batch iteratively trains the model and quantization module based on the above model, quantization module, total loss function and the corpus of the current batch until the training converges (for example: the set total number of training times is reached or the change in the value of the total loss function is stable), thereby obtaining a jointly trained model and a jointly trained quantization module.
[0060] Step A4: Prune the jointly trained model to obtain a pruned model.
[0061] According to one embodiment of the present invention, in this embodiment, the implementer can set a preset pruning rate and use a preset pruning sensitivity index to prune the jointly trained model, wherein a set of trainable parameters that can be pruned is first determined, and based on the sorting of the pruning sensitivity indexes of each trainable parameter, the trainable parameters in the set that have relatively small impact on the gradient are removed until the preset pruning rate is reached to obtain the pruned model.
[0062] According to an embodiment of the present invention, the pruning sensitivity index may be calculated in the following manner:
[0063]
[0064] in, Represents trainable parameters The pruning sensitivity index is represents partial derivative, represents the first sub-loss function, Indicates the number Trainable parameters.
[0065] Alternatively, according to another embodiment of the present invention, the pruning sensitivity index may also be calculated in the following manner:
[0066]
[0067] in, Represents trainable parameters The pruning sensitivity index is represents partial derivative, represents the first sub-loss function, Indicates the number Trainable parameters. In the pruning sensitivity index, the method of multiplying the gradient by the trainable parameter and then taking the absolute value can simultaneously consider the influence of the gradient size and the size of the trainable parameter itself, thereby reducing the pruning of some trainable parameters with large values but small corresponding gradients, reducing the impact on the quantization process, and improving model performance.
[0068] Step A5: quantize the trainable parameters in the pruned model using the jointly trained quantization module to obtain a quantized model.
[0069] According to an embodiment of the present invention, this step is a process of actually quantizing the trainable parameters in the pruned model. At this time, the quantization module is used to perform quantization in the following manner:
[0070]
[0071] in, represents the rounding function, represents the truncation function, represents the bias parameters of the quantization module after joint training, Represents the scaling parameters of the jointly trained quantization module.
[0072] This step uses the jointly trained quantization module to calculate each trainable parameter The corresponding quantized value (low bit integer), using the quantized value Replace the trainable parameters of the original model , thereby avoiding floating-point calculations, further reducing the number of parameters after pruning, and improving reasoning efficiency.
[0073] Implementation Method 2
[0074] In order to better perform quantization and further reduce the impact of the quantization process on the model prediction performance, it is also possible to consider improving implementation mode 1.
[0075] The difference between this embodiment and embodiment 1 is that, in addition to the bias parameter and the scaling parameter, the quantization parameters of this embodiment also have a learnable truncation lower limit parameter and a truncation upper limit parameter in the truncation function, wherein the truncation upper limit parameter is used to limit the maximum quantization value in the quantization range, and the truncation lower limit parameter is used to limit the minimum quantization value in the quantization range.
[0076] During joint training, the trainable parameters of the model and the quantization parameters of this implementation, namely, the bias parameter, the scaling parameter, the truncation upper limit parameter and the truncation lower limit parameter, are updated.
[0077] Implementation 3
[0078] In order to perform more refined quantization of the model and avoid over-quantization of some important parameters, the aforementioned implementation method may also be improved to enhance the performance of the final quantized model.
[0079] The difference between this embodiment and all the above embodiments is that this embodiment sets independent quantization parameters for each trainable parameter, and the quantization parameters of different trainable parameters are updated independently during joint training. In this way, each trainable parameter can be personalized to find its most suitable quantization value, so as to better guarantee the performance of the final quantized model in the process of compressing the model.
[0080] Implementation 4
[0081] The sub-loss function of pruning optimization, that is, the second sub-loss function, may also adopt other forms. According to one embodiment of the present invention, the second sub-loss function may be configured to first calculate the layer mean of the trainable parameters of each processing layer, and then sum the layer means and then take the average. Thus, the entire processing layer (convolutional layer, fully connected layer or attention head, etc.) is removed.
[0082] Implementation method 5
[0083] In the quantization module, bias or scaling parameters may also be used in other forms to achieve the purpose of quantization.
[0084] Schematically, the quantization module performs quantization in the following manner:
[0085]
[0086] in, represents the rounding function, represents the truncation function, represents the bias parameter, Represents the scaling parameter; correspondingly, the inverse quantization function , represents the bias parameter, Represents the scaling parameter
[0087] Alternatively, the quantization module performs quantization in the following manner:
[0088]
[0089] in, represents the rounding function, represents the truncation function, represents the bias parameter, Represents the scaling parameter; correspondingly, the inverse quantization function , represents the bias parameter, Represents the scaling parameter.
[0090] Implementation 6
[0091] In the above implementation, pruning is performed with trainable parameters as the minimum pruning object. Since there are many repeatedly stacked processing layers (convolutional layers, fully connected layers, or attention heads, etc.) in the natural language processing model, it is also possible to consider removing the entire processing layer.
[0092] According to one embodiment of the present invention, in step A4, a layer set consisting of processing layers that can be pruned is first determined, and the average of each trainable parameter contained in each processing layer in the layer set is calculated to obtain a pruning sensitivity index for each processing layer. Based on the sorting of the pruning sensitivity indexes of each processing layer, the processing layers in the layer set that have relatively small impact on the gradient are removed until a preset pruning rate is reached to obtain a pruned model.
[0093] Of course, when pruning, you can also set up staged pruning. In the first stage, first remove the processing layers in the layer set that have a relatively small impact on the gradient to achieve a preset pruning ratio; in the second stage, determine the parameter set of trainable parameters that can be pruned after the first stage of pruning, and remove the trainable parameters in the parameter set that have a relatively small impact on the gradient based on the sorting of the pruning sensitivity index of each trainable parameter until the preset pruning rate is reached to obtain the pruned model. Among them, the pruning ratio is less than the pruning rate to further improve the pruning effect.
[0094] Implementation 7
[0095] The above implementation method can obtain a quantized model, which the implementer can deploy to the target hardware platform (i.e., electronic device, such as edge server or embedded device) to use the quantized model for reasoning, thereby improving the speed of the target hardware platform using the model for reasoning (compared to before compression).
[0096] In order to verify the effect of the method of the present invention, the inventors also conducted tests:
[0097] Models such as Llama and Qwen were tested on multiple data sets, and the test results are shown in the following table:
[0098]
[0099] Among them, ACC (Accuracy) is an evaluation indicator, which is usually used to measure the accuracy of the model on a specific task. The higher the value of the ACC indicator, the higher the accuracy of the model prediction.
[0100] PPL is the abbreviation of perplexity, which is an important indicator to measure the performance of language models. The lower the perplexity, the better the model fits the data.
[0101] Throughput and latency are performance indicators. The higher the throughput, the better, and the lower the latency, the better.
[0102] @arc-c indicates that the test was conducted on a specific dataset or task (such as ARC-C). ARC-C is a benchmark test set based on complex reasoning, which is used to evaluate the performance of the model in logical reasoning and problem solving. Therefore, ACC@arc-c can represent the accuracy score of the model on the ARC-C dataset.
[0103] @mmlu indicates that the test was conducted on the MMLU (Massive Multitask Language Understanding) dataset. MMLU is a large-scale multi-task language understanding dataset that contains knowledge questions from multiple fields and is designed to test the model's extensive knowledge and reasoning capabilities.
[0104] @wikitext2 indicates that the test was conducted on the WikiText-2 dataset. WikiText-2 is a widely used language modeling dataset containing consecutive text paragraphs from Wikipedia articles.
[0105] It can be seen that after using the method of the present invention to compress the existing large model, the size of the video memory occupied by the model and the inference delay can be reduced. At the same time, under certain conditions, the accuracy (the ACC@arc-c index of llama2-7b and llama3.1-8b after compression using the method of the present invention) is improved compared with the original model.
[0106] In general, the model compression method provided by the present invention has at least one of the following beneficial effects:
[0107] (1) Before pruning and quantization, the present invention first uses training corpus to jointly train the model and quantization module for pruning and quantization optimization, thereby obtaining the jointly trained model and quantization module. In the process of joint training, pruning optimization and quantization optimization are integrated, thereby allowing the model to perceive the performance after pruning and quantization in advance, thereby optimizing the trainable parameters to adapt to the state after pruning and quantization, and better ensuring the performance of the model after pruning and quantization.
[0108] (2) By combining pruning and quantization-aware training, it is equivalent to providing a joint optimization framework for pruning and quantization-aware training. By integrating the pruning and quantization-aware training processes, the model is optimized collaboratively between sparse structure and low-precision quantization. As a result, the number of parameters and computing requirements of large language models are significantly reduced, and the storage space occupied by the model is reduced, making it easy to deploy on resource-constrained hardware.
[0109] (3) The synergy between pruning and quantization is considered in the compression process of large language models. The pruning optimization process considers the impact of quantization on the distribution of trainable parameters, thereby reducing the precision loss caused by quantization. At the same time, the quantization optimization process improves the quantization accuracy by adapting to the sparse structure and reduces the precision loss in the quantization process by dynamically evaluating the importance of trainable parameters. While reducing the model size, the overall performance of the model is maintained. Compared with the traditional single pruning method, the model performance of the present invention is more stable, and the accuracy of the compressed model is significantly improved.
[0110] It should be noted that although the above describes the various steps in a specific order, it does not mean that the various steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently or even in a different order as long as the required functions can be achieved.
[0111] The present invention may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present invention.
[0112] A computer-readable storage medium may be a tangible device that holds and stores instructions used by an instruction execution device. Computer-readable storage media may include, for example, but are not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a protruding structure in a groove on which instructions are stored, and any suitable combination thereof.
[0113] The embodiments of the present invention have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A model compression method, the method comprising: Obtaining a model to be pruned, which is a pre-trained natural language processing model and includes a plurality of trainable parameters; Obtain a quantization module constructed for the model, which quantizes the trainable parameters of the model through the learnable quantization parameters to obtain corresponding quantized values, wherein the quantized value corresponding to each trainable parameter is smaller than the data amount of the trainable parameter itself; The model and quantization module are jointly trained by pruning and quantization optimization using the training corpus to obtain the jointly trained model and quantization module, wherein the quantization values corresponding to the trainable parameters are used to temporarily replace the trainable parameters for calculation in the forward propagation during training, and the trainable parameters and quantization parameters are updated in the back propagation with the goal of minimizing the value of the preset total loss function; Pruning the jointly trained model to obtain a pruned model; The trainable parameters in the pruned model are quantized using the jointly trained quantization module to obtain a quantized model.
2. The method according to claim 1, characterized in that The total loss function of joint training is the weighted sum of multiple sub-loss functions, which include: A first sub-loss function, which is configured to calculate the prediction loss of the model; The second sub-loss function is configured to be positively correlated with the mean of the absolute values of the trainable parameters; The third sub-loss function is configured to be positively correlated with the mean of the deviations of the trainable parameters from the dequantized results of their quantized values.
3. The method according to claim 2, characterized in that The total loss function is: in, represents the first sub-loss function, represents the second sub-loss function, represents the third sub-loss function, express The weighting coefficient of express The weighting coefficient of express The weighting coefficient of .
4. The method according to claim 3, characterized in that The second sub-loss function is: in, Represents the total number of trainable parameters, Indicates the number Trainable parameters.
5. The method according to claim 3, characterized in that: The third sub-loss function is: in, Represents the total number of trainable parameters, Indicates the number Trainable parameters, express The corresponding quantized value, Represents a dequantization function that performs dequantization using a quantization parameter.
6. The method according to any one of claims 1 to 5, characterized in that: The quantization module uses quantization parameters to truncate, bias, and scale the trainable parameters, where: The quantization parameters include the bias parameter and the scaling parameter, or The quantization parameters include a truncation lower limit parameter, a truncation upper limit parameter, a bias parameter, and a scaling parameter.
7. The method according to any one of claims 1 to 5, characterized in that: The quantization parameters include bias parameters and scaling parameters, where the quantization module performs quantization in the following manner: in, represents the rounding function, represents the truncation function, represents the bias parameter, Represents the scaling parameter.
8. The method according to any one of claims 1 to 5, characterized in that: The process of pruning the jointly trained model includes: Get the pruning sensitivity index of each trainable parameter and the set of parameters that can be pruned; From the set of prunable parameters, at least part of the trainable parameters of the model are removed in ascending order of the pruning sensitivity index to obtain a pruned model.
9. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 8.
10. An electronic device, characterized in that: include: one or more processors; as well as A memory, wherein the memory is used to store executable instructions; The one or more processors are configured to implement the steps of the method according to any one of claims 1 to 8 by executing the executable instructions or to perform inference using the quantized model obtained by the method according to any one of claims 1 to 8.