A model compression method, an image recognition method, a product, a device, and a medium

By using the predicted pruning mask vector in the pruning stage to remove the redundant channels and adjust the quantization accuracy in the quantization stage, the accuracy loss problem caused by step-by-step execution of pruning and quantization in the prior art is solved, and the efficiency and accuracy retention of model compression are achieved.

CN119721168BActive Publication Date: 2025-07-11LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510229142.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-11
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

The existing pruning and quantization scheme is implemented step by step, ignoring the mutual influence between pruning and quantization, resulting in a large loss of accuracy after model compression, which cannot effectively reduce the storage and computing costs of computing devices.

Method used

The redundant output channel is removed in the pruning stage by predicting the pruning mask vector, and the quantization accuracy is adjusted through pseudo-quantization bitstep length in the quantization stage, and the collaborative optimization is carried out in combination with pruning and quantization, reducing the number of model parameters and calculation complexity while maintaining the model accuracy.

Benefits of technology

While reducing the storage and computing costs of computing devices, model accuracy should be preserved as much as possible and adapted to the image recognition tasks of computing devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119721168B_ABST
    Figure CN119721168B_ABST
Patent Text Reader

Abstract

The present invention discloses a model compression method, an image recognition method, a product, a device and a medium, which relate to the field of data processing. To solve the problem that the model compression effect is poor and cannot be well adapted to the computing device, the method comprises using the predicted pruning mask vectors of each convolutional layer to perform pruning operations on the output channels of each convolutional layer; performing quantization operations on each convolutional layer layer by layer in the current iteration; in response to the current iteration satisfying the end condition, the compressed model after all convolutional layers have completed the quantization operation is deployed as the target model on the computing device. The present invention can retain the accuracy of the model as much as possible while reducing the storage and computing costs of the computing device, thereby better adapting to the image recognition task of the computing device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to a model compression method, an image recognition method, a product, a device and a medium. Background Art

[0002] In the field of image recognition, model compression of deep neural networks is one of the key technologies to achieve efficient deployment. Pruning and quantization are two important optimization techniques. Pruning reduces the number of model parameters by removing redundant network structures, while quantization reduces the storage and computing costs of the model on the computing device by reducing the accuracy of model parameters. However, most existing pruning and quantization schemes adopt a step-by-step execution method, that is, pruning operations are first performed to remove redundant structures, and then the compressed model is quantized. This step-by-step processing method ignores the mutual influence between pruning and quantization, resulting in a large loss of model accuracy after compression, poor compression effect, and failure to effectively reduce the storage and computing costs of the computing device. It may even be impossible to complete the image recognition task assigned to the computing device.

[0003] Therefore, how to provide a solution to the above technical problems is a problem that those skilled in the art need to solve at present. Summary of the invention

[0004] The purpose of the present invention is to provide a model compression method, image recognition method, product, device and medium, which can reduce the storage and computing costs of the computing device while retaining the accuracy of the model as much as possible, so as to better adapt to the image recognition tasks of the computing device.

[0005] In order to solve the above technical problems, the present invention provides a model compression method, comprising:

[0006] Determine the predicted pruning mask vectors of each convolutional layer of the preset model;

[0007] Performing a pruning operation on the output channels of each of the convolutional layers using the predicted pruning mask vectors of each of the convolutional layers;

[0008] In the current iteration, a quantization operation is performed on each of the convolutional layers layer by layer, wherein the quantization operation includes: determining a first highest quantization bit, a first lowest quantization bit, and an initial weight value of the current convolutional layer in the current iteration, obtaining a pseudo quantization bit step length using the first highest quantization bit, and quantizing the initial weight value based on the pseudo quantization bit step length, the first highest quantization bit, and the first lowest quantization bit to obtain a quantized weight value;

[0009] In response to the current iteration satisfying the end condition, the compressed model after all the convolutional layers have performed the quantization operation is deployed on the computing device as the target model.

[0010] Optionally, the process of determining the first highest quantization bit and the first lowest quantization bit of the current convolutional layer in the current iteration includes:

[0011] Determine the quantization bit selection function of the current convolutional layer in the current iteration; the quantization bit selection functions of different convolutional layers in the current iteration are different;

[0012] Predict the first highest quantization bit and the first lowest quantization bit of the current convolutional layer in the current iteration based on the quantization bit selection function.

[0013] Optionally, the quantization bit selection function is b1 = Mish(2t) + 2, b2 = b1 × [Mish(-2t) + 2], where -1 ≤ t ≤ 1, b1 is the first lowest quantization bit, b2 is the first highest quantization bit, Mish( ) is the activation function, and t is the input gating variable of the quantization bit selection function.

[0014] Optionally, the process of obtaining the pseudo quantization bit step size using the first highest quantization bit and quantizing the initial weight value based on the pseudo quantization bit step size, the first highest quantization bit, and the first lowest quantization bit to obtain the quantized weight value includes:

[0015] Quantize the initial weight value using the first highest quantization bit to obtain the first weight value;

[0016] Quantize the initial weight value using the first lowest quantization bit to obtain the second weight value;

[0017] Quantize the initial weight value using the pseudo quantization bit step size to obtain the third weight value;

[0018] Obtain the quantized weight value of the current convolutional layer in the current iteration according to the first weight value, the second weight value, and the third weight value.

[0019] Optionally, the process of obtaining the pseudo quantization bit step size using the first highest quantization bit includes:

[0020] Obtain the pseudo quantization bit step size using the first relational expression, and the first relational expression is ;

[0021] where is the pseudo quantization bit step size, b2 is the first highest quantization bit, is the quantization step size.

[0022] Optionally, the process of obtaining the quantized weight value of the current convolutional layer in the current iteration according to the first weight value, the second weight value, and the third weight value includes:

[0023] Obtaining the quantized weight value of the current convolutional layer in the current iteration according to a second relational expression, where the second relational expression is ;

[0024] where is the quantized weight value, is the first weight value, is the second weight value, is the third weight value.

[0025] Optionally, the quantization operation further includes:

[0026] Determining the second highest quantization bit of the quantization layer of the input data of the current convolutional layer in the current iteration, and quantizing the input data by using the second highest quantization bit to obtain a quantized activation value.

[0027] Optionally, before pruning the output channels of each convolutional layer by using the predicted pruning mask vector of each convolutional layer, the model compression method further includes:

[0028] Determining the predicted highest quantization bit and the predicted lowest quantization bit of each convolutional layer according to the network structure order of the preset model, and performing the quantization operation on each convolutional layer based on the predicted highest quantization bit and the predicted lowest quantization bit of each convolutional layer to obtain an initially compressed model;

[0029] After pruning the output channels of each convolutional layer by using the predicted pruning mask vector of each convolutional layer, the model compression method further includes:

[0030] Retraining the initially compressed model to obtain an intermediate model;

[0031] The process of performing the quantization operation layer by layer on each convolutional layer in the current iteration includes:

[0032] Performing the quantization operation layer by layer on each convolutional layer of the intermediate model in the current iteration.

[0033] Optionally, the model compression method further includes:

[0034] If the current iteration is the first iteration, obtain the current pruning mask vector of the current convolutional layer, determine the model accuracy loss corresponding to the current convolutional layer based on the quantized weight values, the quantized activation values, and the current pruning mask vector, and determine the quantization order of all the convolutional layers in the next iteration based on the model accuracy losses of all the convolutional layers in the current iteration;

[0035] If the current iteration is not the first iteration, determine the model accuracy loss corresponding to the current convolutional layer based on the quantized weight values and the quantized activation values, and determine the quantization order of all the convolutional layers in the next iteration based on the model accuracy losses of all the convolutional layers in the current iteration;

[0036] The process of performing quantization operations layer by layer on each of the convolutional layers in the current iteration includes:

[0037] If the current iteration is not the first iteration, perform quantization operations layer by layer on each of the convolutional layers according to the quantization order determined in the previous iteration;

[0038] If the current iteration is the first iteration, sort the model accuracy losses of all the convolutional layers in the current iteration in descending order to obtain the quantization order of each of the convolutional layers in the current iteration, and perform quantization operations layer by layer on each of the convolutional layers according to the quantization order of the current iteration.

[0039] Optionally, the process of determining the quantization order of all the convolutional layers in the next iteration based on the model accuracy losses of all the convolutional layers in the current iteration includes:

[0040] Sort the model accuracy losses of all the convolutional layers in the current iteration in descending order to obtain the quantization order of each of the convolutional layers in the next iteration.

[0041] Optionally, the model compression method further includes:

[0042] Determine the first constraint term and the second constraint term corresponding to the current convolutional layer in the current iteration;

[0043] Obtain the loss function value of the compressed model in the current iteration based on the first constraint term and the second constraint term;

[0044] Update the gating variable and the initial weight value of the current convolutional layer in the next iteration by using the loss function value, where the gating variable includes the gating variable used to calculate the current pruning mask vector.

[0045] To solve the above technical problems, the present invention further provides an image recognition method, including:

[0046] Receiving an image to be recognized;

[0047] Obtaining input image data based on the image to be recognized;

[0048] The input image data is input into a target model to obtain a recognition result of the image to be recognized; the target model is a model obtained based on any of the model compression methods described above.

[0049] To solve the above technical problems, the present invention also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the model compression method described in any one of the above items and / or the steps of the image recognition method described in the above items.

[0050] In order to solve the above technical problems, the present invention further provides an electronic device, comprising:

[0051] Memory for storing computer programs;

[0052] A processor is used to implement the steps of any one of the above-mentioned model compression methods and / or the steps of the above-mentioned image recognition method when executing the computer program.

[0053] To solve the above technical problems, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the model compression method described in any one of the above items and / or the steps of the image recognition method described above are implemented.

[0054] The present invention provides a model compression method. For each convolutional layer of a preset model, a pruning mask vector is predicted in the pruning stage, and the influence of quantization on the model accuracy is considered in advance. According to the predicted pruning mask vector, redundant output channels in the convolutional layer are removed, thereby reducing the number of parameters and the computational complexity of the model. After the pruning operation is performed on the convolutional layer, in the quantization stage, the quantization accuracy is adjusted by a pseudo-quantization bit step, thereby achieving coordinated optimization between pruning and quantization. While reducing the storage and computing costs of a computing device, the accuracy of the model is retained as much as possible, thereby better adapting to the image recognition task of the computing device.

[0055] The present invention also provides an image recognition method, a computer program product, an electronic device and a computer-readable storage medium, which have the same beneficial effects as the above-mentioned model compression method. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0057] Figure 1 A flow chart of the steps of a model compression method provided by the present invention;

[0058] Figure 2 A flowchart of the steps of an image recognition method provided by the present invention;

[0059] Figure 3 A schematic diagram of the structure of an electronic device provided by the present invention;

[0060] Figure 4 This is a schematic diagram of the structure of a computer-readable storage medium provided by the present invention. DETAILED DESCRIPTION

[0061] The core of the present invention is to provide a model compression method, image recognition method, product, device and medium, which can reduce the storage and computing costs of the computing device while retaining the accuracy of the model as much as possible, thereby better adapting to the image recognition tasks of the computing device.

[0062] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0063] First, please refer to Figure 1 The present invention provides a model compression method, comprising:

[0064] S101: Determine the predicted pruning mask vector of each convolutional layer of the preset model;

[0065] In this embodiment, the network structure of the preset model includes multiple convolutional layers, and the pruning mask vector of each convolutional layer is predicted to obtain the predicted pruning mask vector of each convolutional layer. It can be understood that each convolutional layer includes Cout output channels and Cin input channels, and the convolution kernel width of the convolutional layer is k. w , the convolution kernel height is k h, where the predicted pruning mask vector is used to indicate which output channels in the convolutional layer are redundant, providing a basis for subsequent pruning operations. This step can be achieved by analyzing the weight distribution, channel importance, or other means of the convolutional layer. Specifically, the predicted pruning mask vector can be dynamically generated through the constraint values during the training process.

[0066] S102: Prune the output channels of each convolutional layer using the predicted pruning mask vector of each convolutional layer;

[0067] In this embodiment, according to the predicted pruning mask vector of each convolutional layer, the redundant output channels in the convolutional layer are removed. Since the number of output channels of the pruned convolutional layer is reduced, the number of model parameters and the computational complexity are reduced, and at the same time, a foundation is laid for subsequent quantization operations.

[0068] It can be understood that the predicted pruning mask vector is usually generated based on the importance of the weights. In the pruning stage, through the predicted pruning mask vector, the weights that are more important to the output can be preferentially retained. These retained weights are more likely to be quantized into more accurate values during subsequent quantization due to their high importance, thereby reducing the accuracy loss caused by quantization. That is, in this embodiment, the impact of quantization on accuracy is considered in advance during the pruning stage, making pruning and quantization no longer independent steps, but a process of mutual cooperation.

[0069] S103: Perform quantization operations layer by layer on each convolutional layer in the current iteration. The quantization operations include: determining the first highest quantization bit, the first lowest quantization bit, and the initial weight value of the current convolutional layer in the current iteration, obtaining the pseudo quantization bit step size using the first highest quantization bit, and quantizing the initial weight value based on the pseudo quantization bit step size, the first highest quantization bit, and the first lowest quantization bit to obtain the quantized weight value;

[0070] The model compression process in this embodiment is achieved through multiple iterations. In each iteration, quantization operations need to be performed layer by layer on each convolutional layer. Taking a convolutional layer in one iteration process as an example to illustrate the quantization process, the quantization processes of other iterations and other convolutional layers are the same. In the current iteration, obtain the first highest quantization bit, the first lowest quantization bit, and the initial weight value of the current convolutional layer. Among them, the first lowest quantization bit and the first highest quantization bit jointly define the data range after quantization. The quantization step size in the quantization process can be calculated through the first lowest quantization bit and the first highest quantization bit to optimize the quantization accuracy, reduce the quantization error, and at the same time control the storage and computational costs. By reasonably selecting the highest and lowest quantization bits, a balance can be achieved between model compression and accuracy retention.

[0071] This embodiment uses the first highest quantization bit to calculate the pseudo quantization bit step size, which is used for precision adjustment in the quantization process, reducing the order of quantization bit decomposition, and actually applying quantization bit decomposition to the mixed precision quantization stage, rather than just the quantization bit search stage. Then, the initial weight value is quantized based on the pseudo quantization bit step size, the first highest quantization bit, and the first lowest quantization bit to obtain the quantized weight value, completing the convolution layer weight value quantization step in the model compression process. Through this dynamic quantization method, it is possible to reduce storage and computing costs while minimizing precision loss.

[0072] S104: In response to the current iteration satisfying the end condition, the compressed model after all convolutional layers have completed the quantization operation is deployed on the computing device as the target model.

[0073] In this embodiment, if the current iteration does not meet the end condition, the next iteration is entered. If the current iteration meets the end condition, the compressed model after all convolutional layers have completed the quantization operation is used as the target model and deployed on the computing device. While maintaining high accuracy, the target model significantly reduces storage and computing costs and can better adapt to scenarios with limited hardware resources. Among them, the end condition includes but is not limited to the model accuracy reaching a preset threshold, the compression rate meeting the requirements, or the number of iterations reaching the upper limit threshold, etc., which can be selected according to the actual engineering needs, and this embodiment does not make specific limitations here.

[0074] This embodiment avoids the problem of pruning and quantization being independent of each other in the step-by-step execution method in the related art by jointly optimizing pruning and quantization operations. This joint optimization method can better balance the sparsity and quantization accuracy of the model, thereby retaining the accuracy of the model as much as possible while reducing the number of model parameters and computational costs. By dynamically calculating the pseudo-quantization bit step size based on the first highest quantization bit, the quantization accuracy can be flexibly adjusted according to the characteristics of the convolutional layer. This dynamic adjustment mechanism helps to better balance the accuracy loss and compression rate during the quantization process. It is particularly suitable for hardware devices that need to be efficiently deployed, and can effectively meet the requirements of real-time and resource constraints.

[0075] The present invention provides a model compression method. For each convolutional layer of a preset model, a pruning mask vector is predicted in the pruning stage, and the influence of quantization on the model accuracy is considered in advance. According to the predicted pruning mask vector, redundant output channels in the convolutional layer are removed, thereby reducing the number of parameters and the computational complexity of the model. After the pruning operation is performed on the convolutional layer, in the quantization stage, the quantization accuracy is adjusted by a pseudo-quantization bit step, thereby achieving coordinated optimization between pruning and quantization. While reducing the storage and computing costs of a computing device, the accuracy of the model is retained as much as possible, thereby better adapting to the image recognition task of the computing device.

[0076] Based on the above embodiments:

[0077] In an exemplary embodiment, the process of determining the first highest quantization bit and the first lowest quantization bit of the current convolutional layer in the current iteration includes:

[0078] Determine the quantization bit selection function of the current convolutional layer in the current iteration; the quantization bit selection functions of different convolutional layers in the current iteration are different;

[0079] Predict the first highest quantization bit and the first lowest quantization bit of the current convolutional layer in the current iteration based on the quantization bit selection function.

[0080] In this embodiment, the quantization bit selection function is constructed based on the Mish activation function. The Mish activation function is Mish(t) = t × tanh(ln(1 + e t )). On this basis, the specific relationship of the quantization bit selection function is as follows:

[0081] b1 = Mish(2t) + 2, b2 = b1 × [Mish(-2t) + 2];

[0082] where -1 ≤ t ≤ 1, t is the input gating variable of the quantization bit selection function and is the input of the Mish function, b1 is the first lowest quantization bit, and b2 is the first highest quantization bit. The quantization bit selection function constructed by the Mish activation function can dynamically adjust the quantization bits in different convolutional layers and iteration processes, thereby reducing the computational complexity while maintaining the model accuracy.

[0083] In each iteration, determine the quantization bit selection function for each output channel of the convolutional layer, and assign a random initial value to the input gating variable t required for each quantization bit selection function. The value range of t is [-1, 1]. Please refer to Table 1, which shows the corresponding relationship between the values of b1 and b2 when t takes different values.

[0084] Table 1 Value Table of Quantization Bit Selection Function under Different Parameter Settings

[0085]

[0086] The quantization bit selection function constructed in this embodiment has the characteristic of tending to select low quantization bits. It can be understood that when the value of t is small, the output of the Mish function will make the values of b1 and b2 small. Low quantization bits can significantly reduce the storage space and computational complexity of the model, thereby improving the running efficiency of the model on hardware, especially suitable for resource-constrained devices such as mobile devices and embedded systems. At the same time, there are also various possibilities of selecting high quantization bits, which can compress the model more while ensuring the accuracy of the quantized model. It can be understood that although the quantization bit selection function tends to select low quantization bits, due to the non-linear characteristics of the Mish function and the value range of t, the quantization bit selection function can still output higher quantization bits. For example, when t is close to 1 or -1, the output of the Mish function may cause the values of b1 and b2 to be large, thus selecting higher quantization bits.

[0087] As an alternative embodiment, for the Cout output channels of each convolutional layer, the Cout output channels can be pre-divided into G output channel groups. Each output channel group includes at least one output channel, and each output channel group has its own quantization bit selection function. Random initial values are assigned to the vector composed of the input gating variables required for these G quantization bit selection functions, and the value range is [-1, 1]; then, according to the quantization bit selection function in the above relationship, the initial values of the first lowest quantization bit b1 and the first highest quantization bit b2 are calculated and generated within the c-th output channel group, where c is the label of the output channel group.

[0088] In an exemplary embodiment, the process of obtaining the quantized weight value by quantizing the initial weight value based on the pseudo quantization bit step, the first highest quantization bit, and the first lowest quantization bit includes:

[0089] Quantize the initial weight value using the first highest quantization bit to obtain the first weight value;

[0090] Quantize the initial weight value using the first lowest quantization bit to obtain the second weight value;

[0091] Quantize the initial weight value using the pseudo quantization bit step to obtain the third weight value;

[0092] Obtain the quantized weight value of the current convolutional layer in the current iteration according to the first weight value, the second weight value, and the third weight value.

[0093] In this embodiment, the quantization bit decomposition based on the pseudo quantization bit is described:

[0094] Linear asymmetric quantization is usually adopted when compressing the activation feature map, such as , linear symmetric quantization is adopted when compressing the weight parameters, such as .

[0095] Among them, z x is the quantized activation value, z w is the quantized weight parameter, x is the original activation value, w is the original weight parameter, x0 is the zero point, v x , v w are the data maximum values, which are used to normalize the activation value and the weight parameter, s is the quantization step, b is the quantization bit, n is used to determine the quantization range, clip is the limiting function, and round is the rounding function.

[0096] Obtain the linearly symmetric quantized weight value within the c-th output channel group according to the first lowest quantization bit b1 , that is ; obtain the linearly symmetric quantized weight value within the c-th output channel group according to the first highest quantization bit b2 , that is ; adaptively estimate the pseudo quantization bit b3 based on the first highest quantization bit b2, and the quantization step of the pseudo quantization bit ;

[0097] Among them, is the pseudo quantization bit step, b2 is the first highest quantization bit, is the quantization step. Use the pseudo quantization bit step to obtain the linearly symmetric quantized weight value within the c-th output channel group , if the first highest quantization bit b2 is 4 bits, then , b3 is a bit value close to 3 bits, if the first highest quantization bit b2 is 8 bits, then , b3 is a bit value close to 6 bits, w c is the weight parameter within the c-th output channel group.

[0098] Obtain the quantized weight value of the current convolutional layer in the current iteration according to the second relational expression, and the second relational expression is ; among them, is the quantized weight value, is the first weight value, is the second weight value, is the third weight value.

[0099] In this embodiment, a pseudo quantization bit is introduced in the quantization bit decomposition, the value of the first lowest quantization bit b1 is not preset, and the order of the quantization bit decomposition is reduced from the K-1 order to the first order.

[0100] Among them, , b3 is the pseudo - quantization bit constructed by the method of the present invention, and the quantization step of b3 is k times that of b2. To ensure that the quantization error of the approximate quantization value is smaller than the original quantization error , the upper limit of the value of k is as shown in to . . To make the formula expression more concise, hereinafter, let the highest quantization bit b k = N, b k-1 = N / 2. When N = 4, k ≤ 2.5, k can take 2; when N = 8, k ≤ 8.5, k can take 2 or 4. z is the original data, z N is the data after quantization using the quantization bit N, z N / 2 is the data after quantization using the quantization bit N / 2, is the quantization residual, s N is the quantization step of the quantization bit N, s N / 2 is the quantization step of the quantization bit N / 2.

[0101] If the data dimension is d, that is, assuming the original data z is a vector of dimension d, the difference between the quantization error of the high - quantization bit and the quantization error of the low - quantization bit has an upper bound.

[0102] It shows that the difference between the quantization error based on the pseudo - quantization bit and the quantization error of the high - quantization bit in the method of the present invention also has an upper bound. It shows that the difference between the quantization error of the method of the present invention and the quantization error of the high - quantization bit , and the difference between the quantization error of the method of the present invention and the quantization error of the low - quantization bit , and the difference between the two does not exceed the error upper bound either.

[0103] In an exemplary embodiment, the quantization operation further includes:

[0104] Determine the second - highest quantization bit of the quantization layer of the input data of the current convolutional layer in the current iteration, and use the second - highest quantization bit to quantize the input data to obtain the quantized activation value.

[0105] In this embodiment, the quantization process further includes predicting the lowest / highest quantization bit of the quantization layer of the input data of the current convolutional layer. Assign a random initial value to the input gating variable t required by the quantization - bit selection function for the quantization layer of the input data, - 1 ≤ t ≤ 1. Calculate the predicted values of the second - lowest quantization bit and the second - highest quantization bit of the quantization layer of the current input data according to the quantization - bit selection function, and obtain the quantized activation value according to the second - highest quantization bit , that is , where It refers to performing linear asymmetric quantization on the input data x using a quantization step size. Perform linear asymmetric quantization on the input data x.

[0106] In an exemplary embodiment, before performing pruning operations on the output channels of each convolutional layer using the predicted pruning mask vectors of each convolutional layer, the model compression method further includes:

[0107] Determine the predicted maximum quantization bit and the predicted minimum quantization bit of each convolutional layer according to the network structure order of the preset model, and perform quantization operations on each convolutional layer based on the predicted maximum quantization bit and the predicted minimum quantization bit of each convolutional layer to obtain an initially compressed model;

[0108] After performing pruning operations on the output channels of each convolutional layer using the predicted pruning mask vectors of each convolutional layer, the model compression method further includes:

[0109] Retrain the initially compressed model to obtain an intermediate model;

[0110] The process of performing quantization operations layer by layer on each convolutional layer in the current iteration includes:

[0111] Perform quantization operations layer by layer on each convolutional layer of the intermediate model in the current iteration.

[0112] In this embodiment, after completing the initial quantization and pruning operations, the initially compressed model is retrained to obtain an intermediate model. The purpose of retraining is to restore the accuracy loss caused by pruning and quantization by fine-tuning the weights of the model, and at the same time further optimize the performance of the model.

[0113] In an exemplary embodiment, the model compression method further includes:

[0114] If the current iteration is the first iteration, obtain the current pruning mask vector of the current convolutional layer, determine the model accuracy loss corresponding to the current convolutional layer based on the quantized weight value, the quantized activation value, and the current pruning mask vector, and determine the quantization order of all the convolutional layers in the next iteration based on the model accuracy losses of all the convolutional layers in the current iteration;

[0115] If the current iteration is not the first iteration, determine the model accuracy loss corresponding to the current convolutional layer based on the quantized weight value and the quantized activation value, and determine the quantization order of all the convolutional layers in the next iteration based on the model accuracy losses of all the convolutional layers in the current iteration;

[0116] The process of performing quantization operations layer by layer on each convolutional layer in the current iteration includes:

[0117] If the current iteration is not the first iteration, perform quantization operations layer by layer on each convolutional layer according to the quantization order determined in the previous iteration;

[0118] If the current iteration is the first iteration, sort the model accuracy losses of all convolutional layers in the current iteration in descending order to obtain the quantization order of each convolutional layer in the current iteration, and perform quantization operations layer by layer on each convolutional layer according to the quantization order of the current iteration.

[0119] In this embodiment, in the prediction stage, that is, in the first iteration process, obtain the current pruning mask vector of the current convolutional layer. After calculating the quantized weight value and quantized activation value of the current convolutional layer, perform a convolution calculation on the two, and multiply the calculation result by the current pruning mask vector. The result of the multiplication is the output result of the current convolutional layer. Based on the output result, pass it down to complete the forward inference of the model, and the model accuracy loss corresponding to the current convolutional layer in the current iteration can be obtained. In the quantization stage, that is, in the subsequent iteration process, after calculating the quantized weight value and quantized activation value of the current convolutional layer, perform a convolution calculation on the two, and the result of the convolution calculation is the output result of the current convolutional layer. Based on the output result, pass it down to complete the forward inference of the model, and the model accuracy loss corresponding to the current convolutional layer in the current iteration can be obtained. Sort the model accuracy losses of each convolutional layer in descending order to determine the quantization order of each convolutional layer in the next iteration process. It can be understood that the greater the model accuracy loss of a certain convolutional layer in the current iteration, the earlier the quantization operation order of this convolutional layer in the next iteration process. It can be understood that the convolutional layer with a larger accuracy loss has a more significant impact on the model accuracy. Prioritizing the quantization operation on it can more effectively optimize the compression effect of the model. By dynamically evaluating the impact of each convolutional layer and dynamically adjusting the quantization order, the model can better balance the compression rate and accuracy in each iteration.

[0120] In an exemplary embodiment, the process of determining the quantization order of all convolutional layers in the next iteration based on the model accuracy losses of all convolutional layers in the current iteration includes:

[0121] Sort the model accuracy losses of all convolutional layers in the current iteration in descending order to obtain the quantization order of each convolutional layer in the next iteration.

[0122] In an exemplary embodiment, the model compression method further includes:

[0123] Determine the first constraint term and the second constraint term corresponding to the current convolutional layer in the current iteration;

[0124] Obtain the loss function value of the compressed model in the current iteration based on the first constraint term and the second constraint term;

[0125] Update the gating variables and initial weight values of the current convolutional layer in the next iteration using the loss function value. The gating variables include the gating variable ρ for calculating the current pruning mask vector and the gating variable t for calculating the quantization bit selection function.

[0126] In this embodiment, when implementing hybrid compression of pruning and quantization, Accumulate the quantization losses of each output channel group as the constraint term of the l-th convolutional layer, R l is the constraint term of the l-th convolutional layer. Where g c is the mask value of the pruning mask vector of the c-th output channel group, is the quantization loss of the c-th output channel group quantized by quantization bit b1, is the quantization loss of the c-th output channel group quantized by quantization bit b2, is the quantization loss of the c-th output channel group quantized by quantization bit b K There is , where , C in , k w , k h correspond to the number of output channels, number of input channels, width of the convolutional kernel, and height of the convolutional kernel of the convolutional kernel respectively. G is the number of output channel groups, is the first variable, is the second variable, and take values of 0 or 1, representing whether to use quantization bits b2 and b k in the quantization decomposition. Specifically, when takes a value of 0, it means not to use quantization bit b2 in the quantization decomposition. When takes a value of 1, it means to use quantization bit b2 in the quantization decomposition. When takes a value of 0, it means not to use quantization bit b k in the quantization decomposition. When takes a value of 1, it means to use quantization bit b k .

[0127] Since this embodiment reduces the order of the quantization decomposition, the first constraint term Rl1 of the loss function can be modified to the following recursive calculation method:

[0128] ;

[0129] Correspondingly, the second constraint term Rl2 = m w m h b2, where m w is the width of the input data, m hare the high of the input data, is the sigmoid function, is the L1 norm, , w c is the weight parameter within the c-th output channel group, and ρ is the learnable gating variable.

[0130] Based on the above first constraint value and second constraint value, obtain the loss function value of the compressed model in the current iteration. Update the initial weight value of the current convolutional layer, the pruning gating variable ρ, and the input gating variable t required for calculating the quantization bit selection function in the next iteration through the loss function value, further realizing the joint optimization of quantization and pruning.

[0131] In summary, the hybrid compression method based on quantization bit decomposition proposed by the present invention designs a new quantization bit selection function, which can obtain both the lowest quantization bit and the highest quantization bit at the same time, and designs an adaptive pseudo quantization bit estimation method. By introducing pseudo quantization bits into the existing quantization bit decomposition formula, the decomposition order is reduced and the quantization loss is reduced, achieving the joint optimization of quantization and pruning, and can automatically complete channel-level pruning and mixed-precision quantization. In addition, a quantization bit selection function is designed based on the Mish function and first-order quantization bit decomposition is used to lightweight the image recognition network and accelerate image recognition, which is beneficial to the deployment and implementation of deep learning-based image recognition applications on edge devices, such as face recognition, liveness detection, etc.

[0132] In a second aspect, please refer to Figure 2 , the present invention also provides an image recognition method, including:

[0133] S201: Receive the image to be recognized;

[0134] S202: Obtain input image data based on the image to be recognized;

[0135] S203: Input the input image data into the target model to obtain the recognition result of the image to be recognized; the target model is the model obtained based on any one of the above model compression methods.

[0136] In a third aspect, the present invention also provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, they implement the steps of the model compression method described in any one of the above embodiments and / or the steps of the above image recognition method.

[0137] In a fourth aspect, please refer to Figure 3 , the present invention also provides an electronic device, including:

[0138] A memory 21 for storing computer programs;

[0139] A processor 22 for implementing the steps of the model compression method described in any of the above embodiments and / or the steps of the above image recognition method when executing a computer program.

[0140] The electronic device further includes:

[0141] An input interface 23, connected to the processor 22 via a communication bus 26, for obtaining externally imported computer programs, parameters, and instructions, and storing them in the memory 21 under the control of the processor 22. The input interface can be connected to an input device to receive parameters or instructions manually input by the user. The input device can be a touch layer covering the display screen, or a button, trackball, or touchpad provided on the terminal housing.

[0142] A display unit 24, connected to the processor 22 via a communication bus 26, for displaying data sent by the processor 22. The display unit can be a liquid crystal display screen or an electronic ink display screen, etc.

[0143] A network port 25, connected to the processor 22 via a communication bus 26, for communicating with external terminal devices. The communication technology used for this communication connection can be a wired communication technology or a wireless communication technology, such as Mobile High-Definition Link technology, Universal Serial Bus, High-Definition Multimedia Interface, Wi-Fi technology, Bluetooth communication technology, Bluetooth Low Energy communication technology, communication technology based on IEEE802.11s, etc.

[0144] In a fifth aspect, please refer to Figure 4 , the present invention further provides a computer-readable storage medium 30, on which a computer program 31 is stored. When the computer program 31 is executed by a processor, it implements the steps of the model compression method described in any of the above embodiments and / or the steps of the image recognition method described above.

[0145] The computer-readable storage medium 30 may include: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.

[0146] In a sixth aspect, the present invention further provides a model compression system, including:

[0147] A determination module for determining a predicted pruning mask vector for each convolutional layer of a preset model;

[0148] A pruning module for performing a pruning operation on the output channels of each convolutional layer using the predicted pruning mask vector of each convolutional layer;

[0149] A quantization module for performing layer-by-layer quantization operations on each convolutional layer in the current iteration. The quantization operations include: determining a first highest quantization bit, a first lowest quantization bit, and an initial weight value of the current convolutional layer in the current iteration, obtaining a pseudo quantization bit step size using the first highest quantization bit, and quantizing the initial weight value based on the pseudo quantization bit step size, the first highest quantization bit, and the first lowest quantization bit to obtain a quantized weight value.

[0150] An acquisition module for deploying the compressed model after all convolutional layers have completed the quantization operation as the target model on a computing device in response to the current iteration satisfying the end condition.

[0151] In an exemplary embodiment, the process of determining the first highest quantization bit and the first lowest quantization bit of the current convolutional layer in the current iteration includes:

[0152] Determining a quantization bit selection function for the current convolutional layer in the current iteration; the quantization bit selection functions of different convolutional layers in the current iteration are different;

[0153] Predicting the first highest quantization bit and the first lowest quantization bit of the current convolutional layer in the current iteration based on the quantization bit selection function.

[0154] In an exemplary embodiment, the quantization bit selection function is b1 = Mish(2t) + 2, b2 = b1 × [Mish(-2t) + 2], where -1 ≤ t ≤ 1, b1 is the first lowest quantization bit, b2 is the first highest quantization bit, and Mish( ) is an activation function.

[0155] In an exemplary embodiment, the process of obtaining a pseudo quantization bit step size using the first highest quantization bit and quantizing the initial weight value based on the pseudo quantization bit step size, the first highest quantization bit, and the first lowest quantization bit to obtain a quantized weight value includes:

[0156] Quantizing the initial weight value using the first highest quantization bit to obtain a first weight value;

[0157] Quantizing the initial weight value using the first lowest quantization bit to obtain a second weight value;

[0158] Quantizing the initial weight value using the pseudo quantization bit step size to obtain a third weight value;

[0159] Obtaining the quantized weight value of the current convolutional layer in the current iteration according to the first weight value, the second weight value, and the third weight value.

[0160] In an exemplary embodiment, the process of obtaining a pseudo quantization bit step size using the first highest quantization bit includes:

[0161] Obtain the pseudo - quantization bit step size using the first relational expression, where the first relational expression is ;

[0162] Among them, is the pseudo - quantization bit step size, b2 is the first highest quantization bit, is the quantization step size.

[0163] In an exemplary embodiment, the process of obtaining the quantized weight value of the current convolutional layer in the current iteration according to the first weight value, the second weight value, and the third weight value includes:

[0164] Obtain the quantized weight value of the current convolutional layer in the current iteration according to the second relational expression, where the second relational expression is ;

[0165] Among them, is the quantized weight value, is the first weight value, is the second weight value, is the third weight value.

[0166] In an exemplary embodiment, the quantization operation further includes:

[0167] Determine the second highest quantization bit of the quantization layer of the input data of the current convolutional layer in the current iteration, and use the second highest quantization bit to quantize the input data to obtain the quantized activation value.

[0168] In an exemplary embodiment, before using the prediction pruning mask vectors of each convolutional layer to perform pruning operations on the output channels of each convolutional layer, the model compression system is further configured to:

[0169] Determine the predicted highest quantization bit and the predicted lowest quantization bit of each convolutional layer according to the network structure order of the preset model, and perform quantization operations on each convolutional layer based on the predicted highest quantization bit and the predicted lowest quantization bit of each convolutional layer to obtain the initial compressed model;

[0170] After using the prediction pruning mask vectors of each convolutional layer to perform pruning operations on the output channels of each convolutional layer, the model compression system is further configured to:

[0171] Retrain the initial compressed model to obtain an intermediate model;

[0172] The process of performing quantization operations layer by layer on each convolutional layer in the current iteration includes:

[0173] Perform quantization operations layer by layer on each convolutional layer of the intermediate model in the current iteration.

[0174] In an exemplary embodiment, the model compression system is further configured to:

[0175] If the current iteration is the first iteration, obtain the current pruning mask vector of the current convolutional layer, determine the model accuracy loss corresponding to the current convolutional layer based on the quantized weight values, the quantized activation values, and the current pruning mask vector, and determine the quantization order of all the convolutional layers in the next iteration based on the model accuracy losses of all the convolutional layers in the current iteration;

[0176] If the current iteration is not the first iteration, determine the model accuracy loss corresponding to the current convolutional layer based on the quantized weight values and the quantized activation values, and determine the quantization order of all the convolutional layers in the next iteration based on the model accuracy losses of all the convolutional layers in the current iteration;

[0177] The process of performing quantization operations layer by layer on each convolutional layer in the current iteration includes:

[0178] If the current iteration is not the first iteration, perform quantization operations layer by layer on each convolutional layer according to the quantization order determined in the previous iteration;

[0179] If the current iteration is the first iteration, sort the model accuracy losses of all the convolutional layers in the current iteration in descending order to obtain the quantization order of each convolutional layer in the current iteration, and perform quantization operations layer by layer on each convolutional layer according to the quantization order of the current iteration.

[0180] In an exemplary embodiment, the process of determining the quantization order of all the convolutional layers in the next iteration based on the model accuracy losses of all the convolutional layers in the current iteration includes:

[0181] Sort the model accuracy losses of all the convolutional layers in the current iteration in descending order to obtain the quantization order of each convolutional layer in the next iteration.

[0182] In an exemplary embodiment, the model compression system is further configured to:

[0183] Determine a first constraint term and a second constraint term corresponding to the current convolutional layer in the current iteration;

[0184] Obtain the loss function value of the compressed model in the current iteration based on the first constraint term and the second constraint term;

[0185] Update the gating variable and the initial weight value of the current convolutional layer in the next iteration by using the loss function value, where the gating variable includes the gating variable used to calculate the current pruning mask vector.

[0186] In a seventh aspect, the present invention further provides an image recognition system, including:

[0187] A receiving module, configured to receive an image to be recognized;

[0188] A processing module, configured to obtain input image data based on the image to be recognized;

[0189] A result acquisition module, configured to input the input image data into a target model to obtain a recognition result of the image to be recognized; the target model is a model obtained based on any one of the above model compression methods.

[0190] For the introduction of a model compression system, an image recognition method, a system, a computer program product, an electronic device, and a computer-readable storage medium provided by the present invention, please refer to the above embodiments, and the present invention will not be elaborated herein.

[0191] A model compression system, an image recognition method, a system, a computer program product, an electronic device, and a computer-readable storage medium provided by the present invention have the same beneficial effects as the above model compression method.

[0192] It should also be noted that in this specification, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0193] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A model compression method, characterized in that, Including: Determine the predicted pruning mask vectors of each convolutional layer of the preset model; Perform pruning operations on the output channels of each convolutional layer using the predicted pruning mask vectors of each convolutional layer; In the current iteration, perform quantization operations layer by layer on each convolutional layer. The quantization operations include: determining the first highest quantization bit, the first lowest quantization bit, and the initial weight value of the current convolutional layer in the current iteration, obtaining the pseudo quantization bit step size using the first highest quantization bit, and quantizing the initial weight value based on the pseudo quantization bit step size, the first highest quantization bit, and the first lowest quantization bit to obtain the quantized weight value; In response to the current iteration satisfying the end condition, deploy the compressed model after all convolutional layers have completed the quantization operations as the target model on the computing device, so that the computing device performs the image recognition task assigned to itself through the target model; The process of obtaining the pseudo quantization bit step size using the first highest quantization bit and quantizing the initial weight value based on the pseudo quantization bit step size, the first highest quantization bit, and the first lowest quantization bit to obtain the quantized weight value includes: Quantize the initial weight value using the first highest quantization bit to obtain the first weight value; Quantize the initial weight value using the first lowest quantization bit to obtain the second weight value; Quantize the initial weight value using the pseudo quantization bit step size to obtain the third weight value; Obtain the quantized weight value of the current convolutional layer in the current iteration according to the first weight value, the second weight value, and the third weight value; The process of obtaining the pseudo quantization bit step size using the first highest quantization bit includes: The pseudo-quantization bit step size is obtained by using the first relational expression, and the first relational expression is ; Wherein, is the pseudo quantization bit step size, and b2 is the first highest quantization bit, is the quantization step size.

2. The model compression method according to claim 1, wherein The process of determining the first highest quantization bit and the first lowest quantization bit of the current convolutional layer in the current iteration includes: Determine the quantization bit selection function of the current convolutional layer in the current iteration; the quantization bit selection functions of different convolutional layers in the current iteration are different; Predict the first highest quantization bit and the first lowest quantization bit of the current convolutional layer in the current iteration based on the quantization bit selection function; 3. The model compression method according to claim 2, wherein The quantization bit selection function is b1 = Mish(2t)+2, b2 = b1×[Mish(-2t)+2], where -1≤t≤1, b1 is the first lowest quantization bit, b2 is the first highest quantization bit, Mish( ) is the activation function, and t is the input gating variable of the quantization bit selection function.

4. The model compression method according to claim 1, wherein The process of obtaining the quantized weight value of the current convolutional layer in the current iteration according to the first weight value, the second weight value, and the third weight value includes: Obtain the quantized weight value of the current convolutional layer in the current iteration according to the second relational expression, where the second relational expression is ; Among them, is the quantized weight value, is the first weight value, is the second weight value, is the third weight value.

5. The model compression method according to any one of claims 1-4, characterized in that, The quantization operations further include: Determine the second highest quantization bit of the quantization layer of the input data of the current convolutional layer in the current iteration, and quantize the input data using the second highest quantization bit to obtain the quantized activation value; 6. The model compression method according to claim 5, wherein Before performing pruning operations on the output channels of each convolutional layer using the predicted pruning mask vectors of each convolutional layer, the model compression method further includes: Determine the predicted highest quantization bit and the predicted lowest quantization bit of each convolutional layer according to the network structure order of the preset model, and perform the quantization operations on each convolutional layer based on the predicted highest quantization bit and the predicted lowest quantization bit of each convolutional layer to obtain the initial compressed model; After pruning the output channels of each of the convolutional layers using the predicted pruning mask vectors of each of the convolutional layers, the model compression method further includes: Retraining the initially compressed model to obtain an intermediate model; The process of performing quantization operations layer by layer on each of the convolutional layers in the current iteration includes: Performing quantization operations layer by layer on each of the convolutional layers of the intermediate model in the current iteration.

7. The model compression method according to claim 5, wherein The model compression method further includes: If the current iteration is the first iteration, obtaining the current pruning mask vector of the current convolutional layer, determining the model accuracy loss corresponding to the current convolutional layer based on the quantized weight values, the quantized activation values, and the current pruning mask vector, and determining the quantization order of all the convolutional layers in the next iteration based on the model accuracy losses of all the convolutional layers in the current iteration; If the current iteration is not the first iteration, determining the model accuracy loss corresponding to the current convolutional layer based on the quantized weight values and the quantized activation values, and determining the quantization order of all the convolutional layers in the next iteration based on the model accuracy losses of all the convolutional layers in the current iteration; The process of performing quantization operations layer by layer on each of the convolutional layers in the current iteration includes: If the current iteration is not the first iteration, performing quantization operations layer by layer on each of the convolutional layers in accordance with the quantization order determined in the previous iteration; If the current iteration is the first iteration, arranging the model accuracy losses of all the convolutional layers in the current iteration in descending order to obtain the quantization order of each of the convolutional layers in the current iteration, and performing quantization operations layer by layer on each of the convolutional layers in accordance with the quantization order of the current iteration.

8. The model compression method according to claim 7, wherein The process of determining the quantization order of all the convolutional layers in the next iteration based on the model accuracy losses of all the convolutional layers in the current iteration includes: Arranging the model accuracy losses of all the convolutional layers in the current iteration in descending order to obtain the quantization order of each of the convolutional layers in the next iteration.

9. The model compression method according to claim 7, wherein The model compression method further includes: Determining a first constraint term and a second constraint term corresponding to the current convolutional layer in the current iteration; Obtaining the loss function value of the compressed model in the current iteration based on the first constraint term and the second constraint term; Updating the gating variable and the initial weight value of the current convolutional layer in the next iteration using the loss function value, where the gating variable includes the gating variable for calculating the current pruning mask vector.

10. An image recognition method, characterized in that, Includes: Receiving an image to be recognized; Obtaining input image data based on the image to be recognized; Inputting the input image data into a target model to obtain a recognition result of the image to be recognized; The target model is a model obtained by the model compression method according to any one of claims 1-9.

11. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the steps of the model compression method according to any one of claims 1 to 9 and / or the steps of the image recognition method according to claim 10 are implemented.

12. An electronic device, characterized in that, Includes: A memory for storing a computer program; A processor for implementing the steps of the model compression method according to any one of claims 1 to 9 and / or the steps of the image recognition method according to claim 10 when executing the computer program.

13. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the steps of the model compression method according to any one of claims 1 to 9 and / or the steps of the image recognition method according to claim 10.

Citation Information

Patent Citations

  • Neural network compression method and device and storage medium

    CN114764614A

  • Automatic driving model compression method and device based on pruning and quantitative training

    CN116126815A