A model compression method, device and readable storage medium

CN118690796BActive Publication Date: 2026-08-21SHENZHEN MICROBT ELECTRONICS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310319361.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-22
Publication Date
2026-08-21
Estimated Expiration
2043-03-22

AI Technical Summary

Technical Problem

[0003]为了提高深度神经网络模型的性能,模型的参数量和计算量也随之急剧增加,从而对模型的训练和部署带来了巨大挑战

Benefits of technology

[0020] The model compression method provided in this invention involves quantizing a first network model to be compressed to obtain a second network model. Then, target network layers are determined within the second network model, and corresponding cross-layer connection modules are added to each target network layer to obtain a third network model. The quantization loss of the target network layer meets a preset condition; that is, the quantization loss of the target network layer exceeds an acceptable range, causing the overall model's accuracy to fail to meet the preset accuracy. This invention adds corresponding cross-layer connection modules to each target network layer. These modules can add a small amount of computation to the second network model, thereby increasing the model's expressive power and compensating for the errors caused by quantization in the target network layer. Finally, training the third network model improves the model's accuracy, resulting in the target network model. The target network model is lightweight and features high accuracy and speed, allowing deployment on resource-constrained hardware devices and enabling rapid inference to obtain correct prediction results with low power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118690796B_ABST
    Figure CN118690796B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a model compression method, device and readable storage medium. The method comprises: obtaining a first network model, the first network model being a floating point model trained based on a training set; performing model quantization on the first network model to obtain a second network model; determining a target network layer in the second network model; the quantization loss of the target network layer satisfying a preset condition; adding a corresponding cross-layer connection module at each target network layer to obtain a third network model; the input of the cross-layer connection module being connected to the input of the target network layer, and the output of the cross-layer connection module being connected to the output of the target network layer; and training the third network model to obtain a target network model. Embodiments of the present application can improve the accuracy of the model while compressing the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a model compression method, apparatus, and readable storage medium. Background Technology

[0002] With the development of deep learning technology, deep neural network models are widely used in various application scenarios, such as image processing, speech recognition, reasoning / prediction, knowledge representation, and operation control.

[0003] To improve the performance of deep neural network models, the number of parameters and computational cost have increased dramatically, posing significant challenges to model training and deployment. Particularly in deployment, the sheer volume of parameters and computational demands makes it difficult to deploy deep neural network models on resource-constrained hardware devices such as mobile devices.

[0004] To reduce the hardware consumption of deep neural network models and enable their deployment on resource-constrained hardware, model compression techniques are commonly used to improve model efficiency and reduce the computational and storage resources required. However, the accuracy of compressed neural network models is currently difficult to guarantee. Summary of the Invention

[0005] This invention provides a model compression method, apparatus, and readable storage medium, which can improve the accuracy of the model while compressing it, and reduce the computing and storage resources occupied by the model.

[0006] In a first aspect, embodiments of the present invention disclose a model compression method, the method comprising:

[0007] Obtain the first network model, which is a floating-point model trained based on the training set;

[0008] The first network model is quantized to obtain the second network model;

[0009] In the second network model, a target network layer is determined; the quantization loss of the target network layer satisfies a preset condition.

[0010] A third network model is obtained by adding a corresponding cross-layer connection module at each target network layer; the input of the cross-layer connection module is connected to the input of the target network layer, and the output of the cross-layer connection module is connected to the output of the target network layer.

[0011] The third network model is trained to obtain the target network model.

[0012] Secondly, embodiments of the present invention disclose a model compression apparatus, the apparatus comprising:

[0013] The model acquisition module is used to acquire the first network model, which is a floating-point model trained based on the training set.

[0014] The model quantization module is used to quantize the first network model to obtain the second network model.

[0015] The target determination module is used to determine the target network layer in the second network model; the quantization loss of the target network layer satisfies a preset condition.

[0016] The model processing module is used to add a corresponding cross-layer connection module at each of the target network layers to obtain a third network model; the input of the cross-layer connection module is connected to the input of the target network layer, and the output of the cross-layer connection module is connected to the output of the target network layer.

[0017] The model training module is used to train the third network model to obtain the target network model.

[0018] Thirdly, embodiments of the present invention disclose a machine-readable medium having instructions stored thereon that, when executed by one or more processors of a device, cause the device to perform the model compression method as described above.

[0019] The embodiments of the present invention have the following advantages:

[0020] The model compression method provided in this invention involves quantizing a first network model to be compressed to obtain a second network model. Then, target network layers are determined within the second network model, and corresponding cross-layer connection modules are added to each target network layer to obtain a third network model. The quantization loss of the target network layer meets a preset condition; that is, the quantization loss of the target network layer exceeds an acceptable range, causing the overall model's accuracy to fail to meet the preset accuracy. This invention adds corresponding cross-layer connection modules to each target network layer. These modules can add a small amount of computation to the second network model, thereby increasing the model's expressive power and compensating for the errors caused by quantization in the target network layer. Finally, training the third network model improves the model's accuracy, resulting in the target network model. The target network model is lightweight and features high accuracy and speed, allowing deployment on resource-constrained hardware devices and enabling rapid inference to obtain correct prediction results with low power consumption. Attached Figure Description

[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a flowchart illustrating the steps of an embodiment of the model compression method of the present invention;

[0023] Figure 2 This is a schematic diagram of a model structure for adding a cross-layer connection module according to the present invention;

[0024] Figure 3 This is a structural block diagram of an embodiment of a model compression device according to the present invention. Detailed Implementation

[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0026] The terms "first," "second," etc., used in the specification and claims of this invention are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, the first object can be one or more. Furthermore, the term "and / or" in the specification and claims is used to describe the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. In embodiments of this invention, the term "multiple" refers to two or more, and other quantifiers are similar.

[0027] Reference Figure 1 The diagram illustrates a flowchart of an embodiment of a model compression method according to the present invention. The method may include the following steps:

[0028] Step 101: Obtain the first network model, which is a floating-point model trained based on the training set;

[0029] Step 102: Perform model quantization on the first network model to obtain the second network model;

[0030] Step 103: Determine the target network layer in the second network model; the quantization loss of the target network layer satisfies a preset condition;

[0031] Step 104: Add a corresponding cross-layer connection module at each target network layer to obtain a third network model; the input of the cross-layer connection module is connected to the input of the target network layer, and the output of the cross-layer connection module is connected to the output of the target network layer.

[0032] Step 105: Train the third network model to obtain the target network model.

[0033] The model compression method provided in this embodiment of the invention can reduce the loss of model accuracy while compressing the model.

[0034] First, a first network model is obtained, which is a floating-point model trained on a training set. A floating-point model refers to a model where the weights and activations of each layer are floating-point numbers. This first network model is a deep neural network model to be compressed. In practical applications, this first network model can be used to perform a target task. For example, for input X, the first network model can output Y, which, after being mapped by Softmax (a normalized exponential function), outputs the probability value of the corresponding category.

[0035] The embodiments of the present invention do not limit the application scenarios of the first network model. For example, the target task includes, but is not limited to, face recognition, image classification, object detection, semantic segmentation, speech recognition, machine translation, natural language processing, and recommendation systems.

[0036] In this embodiment of the invention, the first network model can be obtained by pre-training an existing neural network in a supervised or unsupervised manner using a training set and machine learning methods. It should be noted that this embodiment of the invention does not limit the network structure or training method of the first network model. The first network model can integrate multiple neural networks. The neural networks include, but are not limited to, at least one or a combination, superposition, or nesting of at least two of the following: CNN (Convolutional Neural Network), LSTM (Long Short-Term Memory) network, RNN (Recurrent Neural Network), attention neural network, etc.

[0037] In practical implementation, for different application scenarios, training data required for the corresponding target task can be collected to obtain a training set. For example, for image classification applications, a large number of images can be collected and manually labeled to obtain a training set. The first network model trained using this training set can perform the target task of image classification. Similarly, for speech recognition applications, a large number of recordings can be collected and manually labeled to obtain a training set. The first network model trained using this training set can perform the target task of speech recognition.

[0038] Taking image classification as an example, a training set is obtained, which includes training data, specifically multiple collected images. Each training data point is labeled with its corresponding category, yielding the true label for each image. Furthermore, before training the first network model using the training set, data augmentation can be performed on the training data. The augmented training set is then used for training to increase the diversity of the training data and improve the accuracy of the first network model. Data augmentation methods include, but are not limited to, basic methods such as scaling, rotation, segmentation, image matting, color enhancement, and noise addition, as well as advanced methods such as mixup and random erasure.

[0039] This invention does not limit the model structure of the first network model. Exemplarily, the model structure may include convolutional neural networks of various sizes, such as one or more of 3×3 convolutions, 5×5 convolutions, dilated convolutions, and grouped convolutions; the model structure may also include batch normalization layers; the activation function of the model structure may use any one of ReLU, Swish, and sigmoid; the model structure may also include residual structures, such as the residual structure of ResNet or the inverse residual structure of Mobilenent; the final output layer of the model structure may be a global pooling layer or a fully connected layer.

[0040] The first network model is trained using the collected training set, and the network parameters are updated using backpropagation until the model converges, resulting in the trained first network model. Model convergence means that the output of the first network model meets the preset accuracy.

[0041] Next, the first network model is quantized to obtain the second network model. The second network model is a compressed model.

[0042] Model quantization refers to the process of approximating floating-point activations or weights (usually represented as 32-bit floating-point numbers) to low-bit integers (such as 16-bit or 8-bit), thereby completing calculations in this low-bit representation. Generally, model quantization can compress model parameters, thereby reducing model storage overhead; and by reducing memory accesses and effectively utilizing low-bit computation instructions, it can improve inference speed. In this embodiment of the invention, model quantization is performed on the first network model to obtain a second network model, achieving model compression; the resulting second network model is a lightweight model.

[0043] The method used for model quantization of the student network model in this embodiment of the invention is not limited. In an optional embodiment of the invention, the step of model quantization of the first network model to obtain a second network model may include:

[0044] Step S11: Insert pseudo-quantization nodes at the target nodes of the first network model;

[0045] Step S12: Obtain a calibration set, which is a subset of the training set;

[0046] Step S13: Input the calibration data in the calibration set into the first network model that has been inserted with pseudo-quantization nodes in sequence, and obtain the quantization information of each pseudo-quantization node.

[0047] Step S14: Based on the quantization information of each pseudo-quantization node, initialize the quantization parameters of each pseudo-quantization node to obtain the second network model.

[0048] The pseudo-quantization node includes a quantization node and a dequantization node. The quantization node quantizes the floating-point input according to the quantization parameters to obtain the quantized output, which is fixed-point data. The dequantization node dequantizes the quantized output according to the quantization parameters to obtain the floating-point output.

[0049] For a quantization node, the quantization output can be represented as follows:

[0050]

[0051] Where Q represents quantized output, R represents floating-point input, and S and Z represent quantization parameters.

[0052] For the dequantization node, the floating-point output can be represented as follows:

[0053] R=(QZ)*S (2)

[0054] Where R represents floating-point output, Q represents quantization input, and S and Z represent quantization parameters.

[0055] Pseudo-quantization nodes are used to perform quantization and dequantization on data. Inserted into the student network model, a pseudo-quantization node first quantizes the input high-precision floating-point data to map it to low-precision fixed-point data, and then dequantizes the low-precision fixed-point data to obtain the output high-precision floating-point data. The pseudo-quantization node keeps both the input and output data as floating-point numbers, but the difference is that the input data is a continuously variable floating-point number, while the output data is a discretized floating-point number. A continuously variable floating-point number refers to any floating-point number within a preset continuous range. For example, if the preset continuous range is 0 to 1, then the continuously variable floating-point number can include any decimal between 0 and 1. A discretized floating-point number refers to any floating-point number within a preset discrete range. For example, if the preset discrete range includes the following three floating-point numbers: 0.33, 0.66, and 0.99, then the discretized floating-point number can include any of these three floating-point numbers.

[0056] In this embodiment of the invention, pseudo-quantized nodes are inserted at the target nodes of the first network model. The target nodes refer to network layer nodes in the first network model that support quantization operations. For example, the target nodes include, but are not limited to, nodes for weights, activations, and the input and output of operators. In specific implementations, deep learning algorithms consist of computational units, also called operators. In the network model, operators correspond to the computational logic in the network layers. For example, a convolutional layer implementing convolution is an operator; a pooling layer implementing pooling is an operator; an activation layer implementing activation is an operator; a fully connected layer implementing fully connected operations is an operator; and so on. For an operator that requires quantization, this embodiment of the invention can insert a pseudo-quantized node before (input node) and after (output node) the operator.

[0057] In one example, assuming the first network model includes two convolutional layers, a pseudo-quantization node can be inserted at the input node of the first convolutional layer, a pseudo-quantization node at the weight node of the first convolutional layer, and a pseudo-quantization node at the output node of the first convolutional layer. Similarly, a pseudo-quantization node can be inserted at the weight node of the second convolutional layer and a pseudo-quantization node at the output node of the second convolutional layer, for a total of five pseudo-quantization nodes. It should be noted that since the output of the first convolutional layer is the input of the second convolutional layer, no further pseudo-quantization nodes need to be inserted at the input node of the second convolutional layer in this example. Each pseudo-quantization node corresponds to its own quantization parameters.

[0058] It is understood that the above examples are for illustrative purposes only, and the embodiments of the present invention do not limit the position and number of pseudo-quantization nodes inserted. For example, assuming that the output of a certain convolutional layer in the first network model is connected to the input of a batch normalization layer, and the output of the batch normalization layer is connected to the input of the activation function, then a pseudo-quantization node can also be inserted at the output node of the batch normalization layer, and a pseudo-quantization node can also be inserted at the output node of the activation function.

[0059] In this embodiment of the invention, pseudo-quantization nodes are inserted at the target nodes of the first network model. Calibration data from the calibration set is sequentially input into the first network model with the inserted pseudo-quantization nodes, and the quantization information of each pseudo-quantization node is statistically analyzed. Based on the quantization information of each pseudo-quantization node, the parameters of each pseudo-quantization node are initialized to obtain a second network model, which is the quantized first network model. The calibration set can be a subset of the training set. For example, a portion of the data can be extracted from the training set as the calibration set.

[0060] Optionally, the quantization information may include the maximum and minimum floating-point values ​​flowing through each pseudo-quantization node, as well as the preset maximum and minimum fixed-point values ​​corresponding to each pseudo-quantization node.

[0061] In this embodiment of the invention, the inserted pseudo-quantization node can be used to count the maximum and minimum floating-point values ​​flowing through the node, and can also be used to simulate the precision loss caused by quantization. Based on this precision loss, the model parameters are adjusted in reverse to continuously reduce the precision loss caused by quantization, thereby achieving the goal of improving the accuracy of the second network model.

[0062] Based on the quantization information of each pseudo-quantization node obtained through statistics, the quantization parameters of each pseudo-quantization node can be initialized. These quantization parameters include a scaling factor S and a zero point Z. The goal of model quantization is to calculate the scaling factor S and the zero point Z based on the numerical range of the floating-point input and the quantized output. Here, floating-point input refers to floating-point type input data. Quantized output refers to the pseudo-quantization node using the quantization parameters to manipulate the floating-point input, mapping the continuous floating-point input to discrete values, and then mapping them to the preset output range to obtain the quantized output. The value of zero point may be 0 (corresponding to symmetric quantization) or not 0 (corresponding to asymmetric quantization).

[0063] In this embodiment of the invention, the maximum and minimum floating-point values ​​flowing through each pseudo-quantization node refer to the maximum and minimum floating-point values ​​corresponding to the floating-point input of each pseudo-quantization node; the preset maximum and minimum fixed-point values ​​corresponding to each pseudo-quantization node refer to the preset maximum and minimum fixed-point values ​​corresponding to the quantization output of each pseudo-quantization node.

[0064] After using the calibration set to statistically analyze the quantization information of each pseudo-quantized node in the first network model, the quantization parameter of each pseudo-quantized node can be calculated using the following formula:

[0065]

[0066]

[0067] Where Rmax and Rmin represent the maximum and minimum floating-point values ​​obtained statistically, respectively, and Qmax and Qmin represent the maximum and minimum fixed-point values, respectively. The maximum and minimum fixed-point values ​​can be calculated based on the quantization bit width. For example, if the quantization bit width is 8 bits, the range for signed numbers is -127 to 127, then Qmax is 127 and Qmin is 127; the range for unsigned numbers is 0 to 255, then Qmax is 0 and Qmin is 255.

[0068] In one example, assuming we want to quantize a convolutional layer in the first network model, we can insert a pseudo-quantized node at each of the three nodes: the weight w, the input x, and the output y. That is, we insert three pseudo-quantized nodes. Let's denote the pseudo-quantized node for weight w as q1, the pseudo-quantized node for input x as q2, and the pseudo-quantized node for output y as q3. The Rmax and Rmin values ​​corresponding to pseudo-quantized node q1 are the maximum and minimum floating-point parameters of this convolutional layer. The Rmax and Rmin values ​​corresponding to pseudo-quantized nodes q2 and q3 need to be obtained from calibration data in the calibration set. For example, suppose the calibration set includes 100 calibration data points, which are 100 images. These 100 images are input into the first network model with the three pseudo-quantization nodes mentioned above. By statistically analyzing the maximum and minimum floating-point values ​​flowing through the input x after inputting these 100 images, the Rmax and Rmin values ​​corresponding to pseudo-quantization node q2 can be obtained. Similarly, by statistically analyzing the maximum and minimum floating-point values ​​flowing through the output y after inputting these 100 images, the Rmax and Rmin values ​​corresponding to pseudo-quantization node q3 can be obtained. It should be noted that the Rmax and Rmin values ​​for each pseudo-quantization node are obtained by statistically analyzing the maximum and minimum values ​​of all calibration data in the calibration set; that is, Rmax and Rmin are obtained through global statistical analysis of the calibration data.

[0069] In an optional embodiment of the present invention, the method may further include: using the second network model to perform inference, and determining whether the output result of the second network model meets a preset precision.

[0070] After quantizing the first network model to obtain a lightweight second network model, inference can be performed using this second network model. The output of the second network model can then be used to determine if it meets the preset accuracy. If it does, the accuracy loss caused by the quantization process is small, and the second network model still has high accuracy; in this case, subsequent steps 103 to 105 are unnecessary. If the output of the second network model does not meet the preset accuracy, the accuracy loss caused by the quantization process is large, and the accuracy of the second network model no longer meets the preset requirements. In this case, subsequent steps 103 to 105 can be performed to improve the accuracy of the second network model, enabling it to reach the preset accuracy.

[0071] In one example, assuming the first network model contains two convolutional layers, five pseudo-quantization nodes can be inserted into this first network model. First, the quantization parameters of each pseudo-quantization node in the first network model are initialized using a calibration set to obtain the second network model. Then, this second network model is used for inference to determine whether its output meets the preset accuracy. Specifically, the input data is input into the second network model. This input data undergoes quantization and dequantization operations at the pseudo-quantization nodes of the first convolutional layer input node, and quantization and dequantization operations at the pseudo-quantization nodes of the first convolutional layer weight node, resulting in the output data of the first convolutional layer. This output data undergoes quantization and dequantization operations at the pseudo-quantization nodes of the first convolutional layer output node, resulting in the input data of the second convolutional layer. This input data undergoes quantization and dequantization operations at the pseudo-quantization nodes of the second convolutional layer weight node, resulting in the output data of the second convolutional layer. This output data undergoes quantization and dequantization operations at the pseudo-quantization nodes of the second convolutional layer output node, resulting in the output result.

[0072] In this embodiment of the invention, it is determined whether the output result of the second network model meets the preset accuracy, that is, whether the accuracy loss of the second network model is within an acceptable range.

[0073] The preset accuracy can be set according to actual needs. Meeting the preset accuracy means that the accuracy of the output result of the second network model is greater than the preset accuracy. For example, assuming the preset accuracy is 78, the accuracy of the output result of the first network model is 80, and the accuracy of the output result of the second network model after quantization is 78.1, then the accuracy of the output result of the second network model is considered to meet the preset accuracy.

[0074] In an optional embodiment of the present invention, the network layers of the second network model may include weights of different bit widths.

[0075] This invention can use mixed-precision quantization to further improve model accuracy. Mixed-precision quantization means that the network layers of the second network model obtained after quantizing the first network model can include weights with different bit widths. For example, a second network model may include convolutional layer 1, convolutional layer 2, and convolutional layer 3; wherein the weight bit width of convolutional layer 1 is 4 bits, the weight bit width of convolutional layer 2 is 8 bits, and the weight bit width of convolutional layer 3 is 2 bits; in this case, the second network model is of mixed precision.

[0076] Specifically, for a first network model to be quantized, the target bit width of each network layer of the first network model to be quantized can be determined first; then, the first network model can be quantized according to the target bit width of each network layer to obtain the second network model.

[0077] Optionally, the target bit width can be determined based on the quantization sensitivity of each network layer. The step of determining the target bit width to be quantized for each network layer of the first network model based on the quantization sensitivity of each network layer may include: calculating the single-layer quantization loss of each network layer of the first network model; determining the quantization sensitivity of each network layer based on the single-layer quantization loss of each network layer; and determining the target bit width to be quantized for each network layer based on the quantization sensitivity of each network layer.

[0078] The quantization sensitivity of a network layer can be determined by the single-layer quantization loss of that layer. The single-layer quantization loss refers to the accuracy loss incurred during the quantization process, which involves quantizing only the parameters (such as weights and / or activations) of that single layer. A larger single-layer quantization loss indicates higher quantization sensitivity of the network layer, and vice versa.

[0079] For network layers with high quantization sensitivity, a high bit width can be assigned; for network layers with low quantization sensitivity, a low bit width can be assigned. Furthermore, a sensitivity threshold can be set, and the target bit width for quantization of each network layer can be determined based on its quantization sensitivity and the threshold. For example, a first target bit width can be set for network layers with a sensitivity greater than or equal to the threshold; a second target bit width can be set for network layers with a sensitivity less than the threshold.

[0080] For example, for a first network model to be quantized, a mixed-precision quantization method is used for the first network model. The mixed precision includes a first target bit width (e.g., 8 bits) and a second target bit width (e.g., 4 bits). A sensitivity threshold can be set. For network layers that are greater than or equal to the sensitivity threshold, the target bit width to be quantized for the network layer can be set to the first target bit width (e.g., 8 bits); for network layers that are less than the sensitivity threshold, the target bit width to be quantized for the network layer can be set to the second target bit width (e.g., 4 bits).

[0081] It should be noted that the embodiments of the present invention do not limit the method for determining the target bit width to be quantized for each network layer in the first network model. For example, the target bit width can be determined by the Neural Architecture Search (NAS) method.

[0082] The basic idea of ​​Neural Network Architecture (NAS) is to find the optimal network structure from a set of candidate neural network structures, called the search space, using a certain search strategy. In this embodiment, a search space can be set, which includes convolutional layers of different bit widths, such as 8-bit and 4-bit convolutional layers. A neural network structure composed of convolutional layers of different bit widths is obtained from this search space using a certain search strategy, and this neural network structure satisfies a preset constraint. The preset constraint can be the total bit width of the neural network structure; that is, the sum of the bit widths of all network layers in the searched neural network structure does not exceed this total bit width. The search strategy is not limited in this embodiment; any one of random search strategies, modular search strategies, and continuous search strategies can be used.

[0083] After model quantization, a second network model is obtained from the first network model. Compared to the first network model, the second network model reduces the space occupied by parameters and lowers the hardware memory usage during runtime. However, the second network model, obtained by model quantization, suffers from a loss of accuracy compared to the first network model. To reduce the accuracy loss caused by quantization, this embodiment of the invention determines a target network layer in the second network model; and adds a corresponding cross-layer connection module at each target network layer to obtain a third network model. The quantization loss of the target network layer satisfies a preset condition, which indicates that the quantization loss of the target network layer exceeds a preset threshold, i.e., the quantization loss of the target network layer exceeds an acceptable range. The cross-layer connection module is used to compensate for the quantization loss generated by the target network layer.

[0084] The quantization loss refers to the accuracy loss caused by quantization. The target network layer is the network layer in the second network model that suffers a significant accuracy loss due to quantization; the accuracy loss of the target network layer due to quantization exceeds an acceptable range. This embodiment of the invention determines the target network layer in the second network model to compensate for the accuracy loss caused by the target network layer, thereby improving the accuracy of the entire model.

[0085] In an optional embodiment of the present invention, determining the target network layer in the second network model may include:

[0086] Step S21: Input the same input data into the first network model and the second network model respectively for inference;

[0087] Step S22: Extract the output features of each network layer in the first network model and the second network model;

[0088] Step S23: Calculate the similarity between the output features of corresponding network layers in the first network model and the second network model;

[0089] Step S24: In the second network model, determine the network layer with a similarity less than a preset value as the target network layer.

[0090] In this embodiment of the invention, the same input data is input into both the first network model and the second network model for inference. This input data is sequentially processed by each network layer in the first network model, and the output features of each network layer in the first network model are extracted. Similarly, the input data is sequentially processed by each network layer and each pseudo-quantization node in the second network model, and the output features of each network layer in the second network model are extracted. It should be noted that, for each network layer in the second network model, if a pseudo-quantization node has been inserted at the output node of that network layer, extracting the output features of that network layer refers to extracting the output features of the pseudo-quantization node at that output node.

[0091] Next, the similarity between the output features of corresponding network layers in the first network model and the second network model is calculated. This embodiment of the invention does not limit the method for calculating the similarity. For example, the similarity can be cosine similarity or KL divergence (relative entropy), etc. Taking cosine similarity as an example, if the cosine similarity between the output features of a certain network layer in the first network model and the second network model is less than a preset value, it indicates that the output features of that network layer in the second network model after quantization have a large error, resulting in a significant loss of accuracy after quantization. Therefore, this network layer can be identified as the target network layer.

[0092] In this embodiment of the invention, target network layers are determined in the second network model, and corresponding cross-layer connection modules are added to each target network layer to obtain a third network model. The cross-layer connection modules can add a small amount of computation to the second network model, thereby increasing the model's expressive power and compensating for errors caused by quantization in the target network layers. Compared to the second network model, the third network model can improve accuracy to a certain extent through the cross-layer connection modules.

[0093] The present invention does not limit the structure of the cross-layer connection module, which can be defined as a lightweight network layer. In an optional embodiment of the present invention, the cross-layer connection module may include a convolutional layer, a normalization layer, and an activation function layer. The input of the cross-layer connection module is connected to the input of the target network layer, and the output of the cross-layer connection module is connected to the output of the target network layer.

[0094] Reference Figure 2 The diagram illustrates a model structure diagram of an embodiment of the present invention that includes an added cross-layer connection module. Figure 2 The third network model shown includes network layer 1, network layer 2, and network layer 3. Network layer 2 is the target network layer, and a corresponding cross-layer connection module is added to network layer 2. The input of this cross-layer connection module is connected to the input of network layer 2, and the output of this cross-layer connection module is connected to the output of network layer 2. Figure 2 As shown, the output of network layer 1 is input into network layer 2 and the cross-layer connection module for calculation. The output of the cross-layer connection module and the output of network layer 2 are added together and then input into network layer 3 for calculation. Through this cross-layer connection module, network layer 3 can receive richer feature information, thereby compensating for the errors caused by quantization in network layer 2 and improving the model accuracy.

[0095] In an optional embodiment of the present invention, the cross-layer connection module may include pseudo-quantized nodes. In this embodiment, the cross-layer connection module is used to compensate for the quantization loss generated by the target network layer. After training, the cross-layer connection module still needs to be retained for inference; therefore, the added cross-layer connection module also needs to be quantized. The target network model obtained after training can be a fully quantized model, meaning that each network layer of the target network model is quantized. It should be noted that the method of inserting pseudo-quantized nodes into the cross-layer connection module can be the same as the method of inserting pseudo-quantized nodes into the first network model. For example, a pseudo-quantized node can be inserted at the target node of the cross-layer connection module.

[0096] In an optional embodiment of the present invention, the number of channels in the convolutional layer of the cross-layer connection module is less than a preset proportion of the number of channels in the convolutional layer of the second network model. This embodiment adds a cross-layer connection module at the target network layer, which increases the number of parameters and computational load of the model. If the increase in parameters and computational load by the cross-layer connection module is too large, it will affect the overall computational speed of the model, failing to achieve the goal of improving the model's computational speed through quantization. Therefore, to avoid this problem, this embodiment controls the number of channels in the convolutional layer of the cross-layer connection module to minimize the increase in parameters and computational load while still meeting the model's accuracy requirements. For example, the number of channels in the convolutional layer of the cross-layer connection module is controlled to be less than a preset proportion of the number of channels in the convolutional layer of the second network model. The preset proportion can be set according to actual needs; for example, the preset proportion can be one-tenth, meaning the number of channels in the convolutional layer of the cross-layer connection module should be less than one-tenth of the number of channels in the convolutional layer of the second network model.

[0097] After adding cross-layer connection modules to the target network layer in the second network model, a third network model is obtained. The third network model, including the cross-layer connection modules, is then trained to improve model accuracy, and the target network model is obtained upon completion of training. This embodiment of the invention does not limit the method for training the third network model. For example, a calibration set can be used to calibrate the third network model again to further optimize the quantization parameters of each pseudo-quantization node; alternatively, a teacher network model can be used to perform knowledge distillation on the third network model.

[0098] In an optional embodiment of the present invention, training the third network model to obtain the target network model may include: using the first network model as a teacher network model and the third network model as a student network model for knowledge distillation training to obtain the target network model.

[0099] The core idea of ​​Knowledge Distillation (KD) is to improve the performance of a student network model without changing its structure by guiding a lightweight student network model to "imitate" a more complex and better-performing teacher network model.

[0100] In this embodiment of the invention, the first network model is used as the teacher network model, and the third network model is used as the student network model for knowledge distillation training. This allows the output of the third network model to approximate the output of the first network model. When the training is completed, the target network model obtained has the same or similar accuracy as the first network model.

[0101] In an optional embodiment of the present invention, the step of using the first network model as a teacher network model and the third network model as a student network model for knowledge distillation training to obtain the target network model may include:

[0102] The training data from the training set is sequentially input into the teacher network model and the student network model for inference. The parameters of the student network model are adjusted according to a preset loss function. When the iteration stopping condition is met, the target network model is obtained. The preset loss function includes a first loss function and a second loss function. The first loss function is determined based on the output results of the network output layers of the student network model and the network output layers of the teacher network model for the same training data. The second loss function is determined based on the prediction results of the student network model for the training data and the true labels of the training data.

[0103] The purpose of knowledge distillation is to improve the fitting ability, or learning ability, of the student network model (the third network model). Because the smaller model (the student network model) has relatively fewer parameters, its learning ability is relatively weak, and direct training makes it difficult to converge to the same level as the teacher network model (the first network model). In other words, the error between the student network model's output and the actual value is difficult to reduce to the same level as the teacher network model. Therefore, during distillation training, the output of the teacher network model when given the same input data is used as the label of the input data. The student network model's output is then used to fit the output of the larger model (the teacher network model), converging to a level similar to the teacher network model, thus improving the student network model's learning ability.

[0104] In this embodiment of the invention, training data from the training set is sequentially input into the teacher network model and the student network model, with the same training data being input into both models each time. Taking an image recognition scenario as an example, the training data is input into both the teacher and student network models, and calculations are performed layer by layer until the output layer. Finally, a fully connected layer outputs the probability distribution of the recognition categories. For example, if there are N categories, the fully connected layer outputs N-dimensional features, representing the output probability of each category. That is, the fully connected layer of the teacher network model outputs N-dimensional features, and the fully connected layer of the student network model also outputs N-dimensional features.

[0105] The N-dimensional features output by the fully connected layers of the teacher network model are called soft labels. The goal of knowledge distillation is to use the output of the teacher network model to guide the learning of the student network model, completing the transfer of knowledge from the high-precision model (teacher network model) to the low-precision model (student network model), thereby improving the accuracy of the low-precision model. In this embodiment of the invention, during the knowledge distillation training of the student network model using the teacher network model, the parameters of the student network model are adjusted according to a preset loss function in each training round. Specifically, the network parameters of the student network model and the quantization parameters of each pseudo-quantized node are adjusted, while keeping the parameters of the teacher network model unchanged. The preset loss function includes a first loss function and a second loss function.

[0106] The first loss function is determined based on the output results of the student network model's output layer and the teacher network model's output layer on the same training data. The network output layer refers to the last layer of the model's network, such as a fully connected layer. For example, in the example above, the first loss function can be determined based on the N-dimensional features output by the fully connected layer of the student network model and the teacher network model. The output of the fully connected layer is also called Logits, and the first loss function is a loss function based on Logits, which can be used to measure the degree of difference between the probability distributions output by the student network model and the teacher network model. The purpose of the first loss function is to make the output result (Logits) of the student network model's network output layer as close as possible to the output result of the teacher network model's network output layer. Logits contain richer information than the predicted result and positive / negative labels, enabling more accurate back-parameter tuning and further improving model accuracy.

[0107] The second loss function is determined based on the student network model's predictions for the training data and the true labels of the training data. For example, the second loss function is used to calculate the cross-entropy loss determined based on the student network model's predictions for the training data and the true labels of the training data. The purpose of the second loss function is to make the student network model's predictions as close as possible to the true values.

[0108] It should be noted that the embodiments of the present invention do not limit the loss functions used to calculate the first loss and the second loss. For example, any loss function such as cross-entropy loss, KL divergence loss, L2 loss, MGD loss, and FGD loss can be used.

[0109] The preset loss function combines the two loss functions mentioned above. Specifically, the loss values ​​calculated by the first and second loss functions are added (or weighted) to obtain the final loss value. This final loss value is then used to adjust the network parameters of the student network model and the quantization parameters of each pseudo-quantized node. The student network model is iteratively trained by continuously inputting the same training data into both the student and teacher network models until it converges. That is, until the accuracy of the student network model meets the preset accuracy and no longer increases, the iterative training stops, the quantization parameters of the student network model are saved, and the target network model is obtained.

[0110] In an optional embodiment of the present invention, the method may further include: removing pseudo-quantized nodes in the target network model and using the target network model to perform a target task.

[0111] After training the target network model, the cross-layer connection modules can be retained while the pseudo-quantization nodes are removed. This target network model can then be deployed on the target device to perform the target task. The input to this target network model is fixed-point data, and inference is performed based on fixed-point parameters. This target network model can be deployed on resource-constrained hardware devices and can quickly infer correct prediction results with low power consumption. For example, in image recognition applications, the target network model for image recognition can be deployed in an embedded camera, enabling the recognition of real-time captured images and the output of object classifications.

[0112] The target devices include, but are not limited to: smart home terminals (including air conditioners, refrigerators, rice cookers, water heaters, etc.), smart business terminals (including video phones, smart conference desktop terminals, etc.), wearable devices (including smartwatches, smart glasses, etc.), smart financial terminals, as well as smartphones, tablets, personal digital assistants (PDAs), in-vehicle devices, computers, embedded devices, etc.

[0113] In summary, the model compression method provided in this embodiment of the invention, after quantizing the first network model to be compressed to obtain a second network model, determines the target network layer in the second network model and adds a corresponding cross-layer connection module to each target network layer to obtain a third network model. The quantization loss of the target network layer meets a preset condition; that is, the quantization loss of the target network layer exceeds an acceptable range, causing the overall model's accuracy to fail to meet the preset accuracy. This embodiment of the invention adds a corresponding cross-layer connection module to each target network layer. This cross-layer connection module can add a small amount of computation to the second network model, thereby increasing the model's expressive power and compensating for the errors caused by quantization in the target network layer. Finally, training the third network model can improve the model accuracy, resulting in the target network model. The target network model is a lightweight model with high accuracy and speed, which can be deployed on resource-constrained hardware devices and can quickly infer correct prediction results with low power consumption.

[0114] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.

[0115] Reference Figure 3 The diagram illustrates a structural block diagram of an embodiment of a model compression device according to the present invention. The device may include:

[0116] The model acquisition module 301 is used to acquire a first network model, which is a floating-point model trained based on a training set.

[0117] The model quantization module 302 is used to quantize the first network model to obtain the second network model;

[0118] The target determination module 303 is used to determine a target network layer in the second network model; the quantization loss of the target network layer satisfies a preset condition.

[0119] The model processing module 304 is used to add a corresponding cross-layer connection module at each of the target network layers to obtain a third network model; the input of the cross-layer connection module is connected to the input of the target network layer, and the output of the cross-layer connection module is connected to the output of the target network layer.

[0120] The model training module 305 is used to train the third network model to obtain the target network model.

[0121] Optionally, the target determination module includes:

[0122] The input inference submodule is used to input the same input data into the first network model and the second network model respectively for inference;

[0123] The feature extraction submodule is used to extract the output features of each network layer in the first network model and the second network model;

[0124] The feature comparison submodule is used to calculate the similarity between the output features of corresponding network layers in the first network model and the second network model;

[0125] The target determination submodule is used to determine the network layer with a similarity less than a preset value in the second network model as the target network layer.

[0126] Optionally, the model training module is specifically used to perform knowledge distillation training on the first network model as a teacher network model and the third network model as a student network model to obtain the target network model.

[0127] Optionally, the model training module is specifically used to sequentially input the training data from the training set into the teacher network model and the student network model for inference, adjust the parameters of the student network model according to a preset loss function, and obtain the target network model when the iteration stopping condition is met; the preset loss function includes a first loss function and a second loss function, wherein the first loss function is determined based on the output results of the network output layers of the student network model and the network output layers of the teacher network model for the same training data; and the second loss function is determined based on the prediction results of the student network model for the training data and the true labels of the training data.

[0128] Optionally, the model quantization module includes:

[0129] The node insertion submodule is used to insert pseudo-quantized nodes at the target nodes of the first network model;

[0130] A calibration set acquisition submodule is used to acquire a calibration set, which is a subset of the training set;

[0131] The information acquisition submodule is used to sequentially input the calibration data in the calibration set into the first network model that has been inserted into the pseudo-quantization nodes, and to acquire the quantization information of each pseudo-quantization node.

[0132] An initialization submodule is used to initialize the quantization parameters of each pseudo-quantization node based on the quantization information of each pseudo-quantization node, thereby obtaining a second network model.

[0133] Optionally, the cross-layer connection module includes a pseudo-quantization node.

[0134] Optionally, the cross-layer connection module includes a convolutional layer, a normalization layer, and an activation function layer.

[0135] Optionally, the number of channels in the convolutional layer of the cross-layer connection module is less than a preset ratio of the number of channels in the convolutional layer of the second network model.

[0136] The model compression apparatus provided in this embodiment of the invention performs model quantization on a first network model to be compressed to obtain a second network model. Then, it determines target network layers within the second network model and adds corresponding cross-layer connection modules to each target network layer to obtain a third network model. The quantization loss of the target network layer meets a preset condition; that is, the quantization loss of the target network layer exceeds an acceptable range, causing the overall model's accuracy to fail to meet the preset accuracy. This embodiment of the invention adds corresponding cross-layer connection modules to each target network layer. These cross-layer connection modules can add a small amount of computation to the second network model, thereby increasing the model's expressive power and compensating for the errors caused by quantization in the target network layer. Finally, training the third network model improves the model's accuracy, resulting in the target network model. The target network model is a lightweight model with high accuracy and speed, allowing deployment on resource-constrained hardware devices and enabling rapid inference to obtain correct prediction results with low power consumption.

[0137] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0138] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0139] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0140] This invention also provides a non-transitory computer-readable storage medium, wherein when the instructions in the storage medium are executed by a processor of a device (server or terminal), the device is able to execute the foregoing text. Figure 1The description of the model compression method in the corresponding embodiments is already provided and will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated. For any technical details not disclosed in the computer program products or computer program embodiments related to this application, please refer to the description of the method embodiments of this application.

[0141] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the following claims.

[0142] It should be understood that the present invention is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

[0143] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

[0144] The above provides a detailed description of the model compression method, apparatus, and machine-readable storage medium provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A model compression method, characterized in that, The method includes: Obtain a first network model, which is a floating-point model trained on a training set; the training set includes image data or speech data. The first network model is quantized to obtain a second network model; the network layers of the second network model include weights of different bit widths. In the second network model, a target network layer is determined; the quantization loss of the target network layer satisfies a preset condition. A third network model is obtained by adding a corresponding cross-layer connection module at each target network layer; the input of the cross-layer connection module is connected to the input of the target network layer, and the output of the cross-layer connection module is connected to the output of the target network layer. The target network model is obtained by knowledge distillation training using the first network model as the teacher network model and the third network model as the student network model, including: The training data from the training set is sequentially input into the teacher network model and the student network model for inference. The parameters of the student network model are adjusted according to a preset loss function. When the iteration stopping condition is met, the target network model is obtained. The preset loss function includes a first loss function and a second loss function. The first loss function is determined based on the output results of the network output layers of the student network model and the network output layers of the teacher network model for the same training data. The second loss function is determined based on the prediction results of the student network model for the training data and the true labels of the training data. The parameters of the student network model include the network parameters of the student network model and the quantization parameters of each pseudo-quantized node.

2. The method according to claim 1, characterized in that, Determining the target network layer in the second network model includes: The same input data is fed into the first network model and the second network model respectively for inference; Extract the output features of each network layer in the first network model and the second network model; Calculate the similarity between the output features of corresponding network layers in the first network model and the second network model; In the second network model, the network layer with a similarity less than a preset value is identified as the target network layer.

3. The method according to claim 1, characterized in that, The step of quantizing the first network model to obtain the second network model includes: Insert pseudo-quantized nodes at the target nodes of the first network model; Obtain a calibration set, which is a subset of the training set; The calibration data in the calibration set are sequentially input into the first network model that has been inserted with pseudo-quantization nodes, and the quantization information of each pseudo-quantization node is obtained. The quantization parameters of each pseudo-quantization node are initialized based on the quantization information of each pseudo-quantization node to obtain the second network model.

4. The method according to claim 1, characterized in that, The cross-layer connection module includes pseudo-quantization nodes.

5. The method according to claim 1, characterized in that, The cross-layer connection module includes convolutional layers, normalization layers, and activation function layers.

6. The method according to claim 1, characterized in that, The number of channels in the convolutional layer of the cross-layer connection module is less than a preset ratio of the number of channels in the convolutional layer of the second network model.

7. A model compression device, characterized in that, The device includes: The model acquisition module is used to acquire a first network model, which is a floating-point model trained on a training set; the training set includes image data or speech data. The model quantization module is used to quantize the first network model to obtain a second network model; the network layers of the second network model include weights with different bit widths. The target determination module is used to determine the target network layer in the second network model; the quantization loss of the target network layer satisfies a preset condition. The model processing module is used to add a corresponding cross-layer connection module at each of the target network layers to obtain a third network model; the input of the cross-layer connection module is connected to the input of the target network layer, and the output of the cross-layer connection module is connected to the output of the target network layer. The model training module is used to perform knowledge distillation training on the first network model as a teacher network model and the third network model as a student network model to obtain a target network model. This includes: sequentially inputting training data from the training set into the teacher network model and the student network model for inference; adjusting the parameters of the student network model according to a preset loss function; and obtaining the target network model when an iteration stopping condition is met. The preset loss function includes a first loss function and a second loss function. The first loss function is determined based on the output results of the network output layers of the student network model and the teacher network model for the same training data. The second loss function is determined based on the prediction results of the student network model for the training data and the true labels of the training data. The parameters of the student network model include the network parameters of the student network model and the quantization parameters of each pseudo-quantized node.

8. A machine-readable storage medium having instructions stored thereon that, when executed by one or more processors of a device, cause the device to perform the model compression method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Convolutional neural network model compression method combining pruning and knowledge distillation

    CN113159173A

  • Image processing model training method, image processing method and related equipment

    CN113705317A