A model inference method and device for convolutional neural networks

By performing parameter pruning and quantization on the trained convolutional neural network model, removing parameters in small threshold ranges, and employing bi-segmentation and multi-segmentation quantization methods, the problems of high computational complexity and low training efficiency of convolutional neural network models are solved, achieving efficient model training.

CN116957010BActive Publication Date: 2025-11-04SHENZHEN MSU-BIT UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310857190.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-13
Publication Date
2025-11-04
Estimated Expiration
2043-07-13

AI Technical Summary

Technical Problem

Convolutional neural network models have high computational complexity and redundant parameters, resulting in high memory and bandwidth requirements for hardware platforms. This limits their inference deployment on some devices, and the training process is cumbersome and inefficient.

Method used

By pruning and quantizing the parameters of the trained convolutional neural network model, parameters in the small threshold range are removed. Using bi-segmentation and multi-segmentation quantization methods, the flag bits, sign bits, and effective width bits of the fixed-point values ​​are obtained, and the model parameters are replaced to form a new convolutional neural network model.

Benefits of technology

While ensuring model accuracy, it significantly improves the training efficiency of convolutional neural network models, reduces the cost of quantization parameters, and decreases the amount of computation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116957010B_ABST
    Figure CN116957010B_ABST
Patent Text Reader

Abstract

The application discloses a model inference method and device for a convolutional neural network. Firstly, model parameters of a first trained convolutional neural network model are extracted. Then, pruning is performed on the extracted model parameters to remove model parameters with values in a first threshold interval and to retain model parameters with values in a second threshold interval. Then, quantization processing is performed on each model parameter to obtain values of a flag bit, a sign bit and an effective width bit of each model parameter. The value of the effective width bit of the model parameter is a fixed-point value, which is obtained by applying a two-segment quantization method or a multi-segment quantization method to the model parameter. Finally, the model parameters obtained after the quantization processing are substituted for original model parameters to obtain a second convolutional neural network model. Since the model parameters of the new convolutional neural network model are obtained according to the trained model parameters, the training efficiency of the new convolutional neural network model is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of convolutional neural network model training technology, and more specifically to a model inference method and apparatus for convolutional neural networks. Background Technology

[0002] With the increasing demand for neural network-based artificial intelligence solutions, convolutional neural networks (CNNs) are being used in mobile platforms such as drones and robots, profoundly changing human production and lifestyles. However, CNN models are computationally complex and have redundant parameters, placing high demands on hardware platforms in terms of memory and bandwidth, thus limiting their deployment in some scenarios or on some devices. Recent optimization methods for model inference include model compression, software library optimization, heterogeneous computing, and hardware acceleration. However, the diverse structures, large data volumes, and high computational demands of CNN models also pose significant challenges to hardware implementation of neural network algorithm design. In particular, the training process of CNN models is not only crucial but also extremely complex. Improving the training efficiency of CNN models is currently a major research direction. Summary of the Invention

[0003] This application provides a model inference method for convolutional neural networks, which improves the training efficiency of convolutional neural network models.

[0004] According to a first aspect, one embodiment provides a model inference method for a convolutional neural network, comprising:

[0005] Step 101: Extract the model parameters of the first convolutional neural network model that has been trained; the first convolutional neural network model is used to extract the first image features; the model parameters of the first convolutional neural network model include weights.

[0006] Step 102: Prune the model parameters of the extracted first convolutional neural network model to remove model parameters whose values ​​are in the first threshold range and retain model parameters whose values ​​are in the second threshold range; the range of the second threshold range is greater than that of the first threshold range.

[0007] Step 103: Quantize each of the model parameters to obtain the values ​​of the flag bit, sign bit, and effective width bit of each model parameter; wherein, the value of the effective width bit of the model parameter is a fixed-point value, which is obtained by quantizing the model parameter using a two-segment quantization method or a multi-segment quantization method;

[0008] Step 104: Replace the model parameters of the first convolutional neural network model with the model parameters obtained after quantization to obtain a second convolutional neural network model; the second convolutional neural network model is used to extract the second image features.

[0009] In one embodiment, the first image feature is different from the second image feature.

[0010] In one embodiment, the model inference method further includes:

[0011] Step 105: Train the second convolutional neural network model using a training dataset containing the features of the second image to obtain the accuracy loss value of the second convolutional neural network model.

[0012] Step 106: When the accuracy loss value is greater than a first preset value, change the first threshold range or the second threshold range, and repeat steps 102 to 105.

[0013] When the accuracy loss value is not greater than the first preset value, the second convolutional neural network model is output.

[0014] In one embodiment, the model inference method further includes:

[0015] The binary quantization method is applied to quantize model parameters whose absolute values ​​are less than a first preset initial value and have a relatively high proportion.

[0016] In one embodiment, the bi-segment quantization method for quantizing the model parameters includes:

[0017] The model parameter is identified by a flag bit, a sign bit, and a valid bit width, and the middle of the valid bit width is set as a split point; the flag bit is used to indicate whether the value of the model parameter is a high-order bit value or a low-order bit value, the sign bit is used to indicate whether the value of the model parameter is positive or negative, and the valid bit width is used to indicate the valid value of the model parameter;

[0018] The data before the split point of the effective bit width of the model parameters is cleared to reduce the effective bit width, and the data with reduced effective bit width is used as the model parameters of the second convolutional neural network model.

[0019] In one embodiment, the model inference method further includes:

[0020] The multi-segment quantization method is applied to quantize model parameters with a relatively low percentage and an absolute value not less than the first preset initial value.

[0021] In one embodiment, the multi-segment quantization method for quantizing the model parameters includes:

[0022] The model parameter is identified by a flag bit, a sign bit, and an effective bit width, and at least two split points are set on the effective bit width according to a preset step size; the flag bit is used to indicate whether the value of the model parameter is a high-order bit value or a low-order bit value, the sign bit is used to indicate whether the value of the model parameter is positive or negative, and the effective bit width is used to indicate the effective value of the model parameter;

[0023] The data with effective bit width before the last split point of the model parameters are cleared to reduce the effective bit width, and the data with reduced effective bit width is used as the model parameters of the second convolutional neural network model.

[0024] In one embodiment, the model inference method further includes:

[0025] The second convolutional neural network model is then trained using a training dataset containing the features of the second image to obtain the accuracy loss value of the second convolutional neural network model.

[0026] When the accuracy loss value is greater than a second preset value, the position of the segmentation point in the two-segment quantization method or the multi-segment quantization method is changed, and steps 102 to 105 are repeated.

[0027] When the accuracy loss value is not greater than the second preset value, the second convolutional neural network model is output.

[0028] According to a second aspect, one embodiment provides a computer-readable storage medium storing a program that can be executed by a processor to implement the model inference method as described in the first aspect.

[0029] According to a third aspect, one embodiment provides a model inference apparatus for a convolutional neural network, comprising:

[0030] The parameter extraction module is used to extract the model parameters of the first convolutional neural network model that has been trained; the first convolutional neural network model is used to extract the first image features; the model parameters of the first convolutional neural network model include weights.

[0031] The parameter pruning module is used to prune the model parameters of the extracted first convolutional neural network model to remove model parameters whose values ​​are in the first threshold range and retain model parameters whose values ​​are in the second threshold range; the range of the second threshold range is greater than that of the first threshold range.

[0032] The quantization processing module is used to quantize each of the model parameters to obtain the values ​​of the flag bit, sign bit, and effective width bit of each model parameter; wherein, the value of the effective width bit of the model parameter is a fixed-point value, which is obtained by quantizing the model parameter using a two-segment quantization method or a multi-segment quantization method;

[0033] The model acquisition module is used to replace the model parameters of the first convolutional neural network model with the model parameters obtained after quantization to obtain a second convolutional neural network model; the second convolutional neural network model is used to extract second image features.

[0034] The model inference method according to the above embodiments obtains the model parameters of the new convolutional neural network model based on the model parameters of the already trained convolutional neural network model, which can greatly improve the training efficiency of the new convolutional neural network model. Attached Figure Description

[0035] Figure 1 This is a structural diagram of a convolutional neural network;

[0036] Figure 2 This is a diagram of a single-hidden-layer convolutional neural network structure;

[0037] Figure 3 Image data information in another embodiment;

[0038] Figure 4 This is a flowchart illustrating a model inference method in one embodiment;

[0039] Figure 5 This is a schematic diagram illustrating the data format representation of model parameters in one embodiment;

[0040] Figure 6 This is a schematic diagram illustrating the segmentation position setting in a multi-segment quantization method in one embodiment;

[0041] Figure 7 This is a schematic diagram illustrating the data format representation of model parameters after multi-segment quantization in one embodiment;

[0042] Figure 8 This is a structural block diagram of a model inference device in one embodiment. Detailed Implementation

[0043] The present invention will now be described in further detail with reference to specific embodiments and accompanying drawings. Similar elements in different embodiments are referred to by associated similar element reference numerals. In the following embodiments, many details are described to facilitate a better understanding of this application. However, those skilled in the art will readily recognize that some features may be omitted in different situations, or may be replaced by other elements, materials, or methods. In some cases, certain operations related to this application are not shown or described in the specification. This is to avoid obscuring the core parts of this application with excessive description. For those skilled in the art, detailed description of these related operations is not necessary; they can fully understand the related operations based on the description in the specification and general technical knowledge in the art.

[0044] Furthermore, the features, operations, or characteristics described in the specification can be combined in any suitable manner to form various embodiments. At the same time, the steps or actions in the method description can be rearranged or adjusted in a manner obvious to those skilled in the art. Therefore, the various orders in the specification and drawings are only for the clear description of a particular embodiment and do not imply a necessary order, unless otherwise stated that a particular order must be followed.

[0045] The serial numbers assigned to components in this document, such as "first" and "second," are used only to distinguish the described objects and have no sequential or technical meaning. The terms "connection" and "linkage" used in this application, unless otherwise specified, include both direct and indirect connections (linkages).

[0046] A convolutional neural network is a type of feedforward neural network. Its artificial neurons can respond to a portion of the surrounding units within their coverage area. It can generally be divided into an input layer, a hidden layer, and an output layer. The hidden layer can be further divided into convolutional layers and pooling layers.

[0047] The following example illustrates the structure of a convolutional neural network. Figure 1 The diagram shown illustrates the structure of a convolutional neural network. The input to this convolutional neural network is an image with a resolution of a*a, for example, a 28*28 resolution image.

[0048] The convolutional layer C1 uses M n*n convolutional kernels to perform convolution operations on the above image to obtain M b*b resolution images. Generally, bias and activation operations are also added, but these two steps are omitted to facilitate understanding of the structure of the convolutional neural network.

[0049] Sampling layer S2 performs a sampling operation on the M b*b resolution images obtained by convolutional layer C1 to obtain M b / 2*b / 2 resolution images.

[0050] Convolutional layer C3 uses 12 5*5 convolutional kernels to perform convolution operations on the 6 images with a resolution of 12*12 obtained from sampling layer S2, resulting in 12 images with a resolution of 8*8.

[0051] Sampling layer S3 performs a sampling operation on the 12 8*8 resolution images obtained by convolutional layer C3 to obtain 12 4*4 resolution images.

[0052] The output layer is a fully connected layer that outputs 12 4x4 resolution images obtained from the sampling layer S3, resulting in 12 feature information of the images.

[0053] In the above example, the convolutional neural network uses two convolutional layers, and the fully connected output of the output layer is also a special kind of convolution operation. Therefore, convolution calculation is the main task of the convolutional neural network, and the convolution operation parameters are the core of the convolutional neural network.

[0054] Please refer to Figure 2 This is a diagram of a single-hidden-layer convolutional neural network (CNN). The input layer of this CNN takes 7x7 pixel images from the MINIST dataset as input, performs 4-bit quantization, and rounds to two decimal places. The MINIST dataset consists of multiple grayscale images, each represented by a numeric array containing 7x7 pixels. Each image has a corresponding label, which is a numerical value. Figure 3 The image shown represents an image where the image information is the number "0" and the corresponding pixel values. Each pixel value in each image represents the intensity value of that pixel within the image. The training dataset for the MINIST dataset uses numbers from 0 to 9 to describe the numbers represented in a given image. This is called a unique vector; a unique vector has all digits except for one digit (1) as 0. Therefore, in this example, the number N will be represented as a 10-dimensional vector where only the Nth dimension (starting from 0) is 1. For example... Figure 3 The label 0 shown will be represented as ([1, 0, 0, 0, 0, 0, 0, 0, 0, 0]). Its label is a matrix of numbers [10000, 10].

[0055] In this embodiment, six 4x4 convolutional kernels are selected, and the following formula (1) is used to obtain... Y k :

[0056] ;

[0057] in, k Represents the number of convolution kernels (e.g.) k =1, 2, 3, 4, 5, 6). m Represents the kernel size (4-bit quantization in this example). m=4*4), Wc ik These are the weights of the convolutional layer. bc k For kernel bias, x i The input image data.

[0058] The pooling layer's maximum downsampling operation uses dim=2. The fully connected layer has 2*2*10=40 weights, quantized to 4 bits with 0 decimal places. The fully connected layer's bias has 10 ones, quantized to 4 bits with 2 decimal places. The convolution formula for the fully connected layer is:

[0059] ;

[0060] Among them, Z out_channeli It represents the result of a fully connected operation. f k () represents the activation pooling transformation function. Y k It is the output value of the previous convolutional computation layer. k Represents the number of biases in fully connected layers ( k =1.2.3. 4... k (for natural numbers) i This represents the fully connected output category channel number. m Represents the number of channels (e.g.) m =4*4=16), Wd ik These are fully connected weights. bd k This is used for biasing the fully connected layer.

[0061] The training process of a convolutional neural network model is the process of quantizing the parameters of the convolution operation. In the current technology, the only way to quantize the parameters of a convolutional neural network model is to perform long-term effective training on the entire convolutional neural network.

[0062] In this embodiment, a second convolutional neural network model is inferred based on a pre-trained first convolutional neural network model by means of directional parameter tuning and a small amount of training. This eliminates the need to retrain the entire convolutional neural network model, thereby reducing the cost of quantization parameters while ensuring model accuracy and greatly improving the efficiency of convolution operation parameter quantization in the convolutional neural network model.

[0063] Example 1:

[0064] Please refer to Figure 4 This is a flowchart illustrating a model inference method in one embodiment. The model inference method obtains the model parameters of a second convolutional neural network model based on the model parameters of a first convolutional neural network model that has already been trained. Specifically, it includes:

[0065] Step 101: Extract model parameters.

[0066] Extract the model parameters of the first convolutional neural network model that has been trained. The first convolutional neural network model is used to extract the first image features, and the model parameters of the first convolutional neural network model include weights.

[0067] Step 102: Perform the pruning operation.

[0068] The model parameters of the extracted first convolutional neural network model are pruned to remove those with values ​​in the first threshold interval and retain those with values ​​in the second threshold interval. The range of the second threshold interval is larger than that of the first threshold interval.

[0069] Pruning involves setting a large threshold (first threshold interval) and a small threshold (second threshold interval), and then approximating the accuracy loss results of these two thresholds until a point with a small accuracy loss and a large threshold is found. Statistical analysis of the model parameters of multiple convolutional neural networks (CNNs) shows that the distribution of CNN model parameters approximates a normal distribution with a mathematical expectation of 0. From the fact that the parameters approximate a normal distribution, we know that the weights of CNNs are mainly concentrated in the interval [-1, 1] (a few networks may have weights slightly larger than [-1, 1], varying slightly depending on the training method and network structure). Small parameters close to 0 (within the second threshold interval) constitute the majority, while parameters with extremely large absolute values ​​(within the first threshold interval) account for a limited proportion.

[0070] Step 103: Perform quantification.

[0071] Each model parameter is quantized to obtain the values ​​of its flag bit, sign bit, and effective width bit. The effective width bit value is a fixed-point value obtained by quantizing the model parameter using a bi-segment quantization method or a multi-segment quantization method.

[0072] In the process of quantizing network parameters, different quantization strategies should be adopted for network parameters with large and small absolute values ​​of weights. For parameters with large absolute values, their original values ​​should be maintained as much as possible to ensure that their changes are small. For parameters with small absolute values, their numerical changes can be ignored to a certain extent, and parameters with absolute values ​​within a certain threshold can even be directly represented as 0.

[0073] Furthermore, for the same neural network model, there is a significant difference between the weights of convolutional layers and fully connected layers. For example, in AlexNet, the weights of convolutional layers are concentrated in the range of (-0.4, 0.4), while the weights of fully connected layers are concentrated in the range of (-0.04, 0.04). There is an order of magnitude difference between the two; after converting the values ​​to signed fixed-point binary numbers, there is a 3-bit difference (the highest 3 decimal places after the sign bit in fully connected layers are 0). Therefore, different quantization strategies should be chosen when quantizing network parameters for weights with different numerical ranges. Thus, quantizing model parameters requires analyzing the distribution of parameters based on their characteristics, replacing them with appropriate values, and then quantizing the network parameters while ensuring minimal loss of accuracy. For example:

[0074] After pruning the network parameters, it is necessary to select an appropriate bit width representation based on the length of the data. Similarly, the appropriate length division is determined based on the evaluation of accuracy loss, that is, the bit width and splitting position of the fixed point are determined.

[0075] In one embodiment, model parameters with a relatively high percentage whose absolute value is less than a first preset initial value are quantized using a two-piece quantization method. The two-piece quantization method for quantizing model parameters includes:

[0076] The model parameters are identified by their flag, sign, and effective bit width, with the middle of the effective bit width designated as a split point. The flag indicates whether the model parameter value is a high-order or low-order value, the sign indicates whether the model parameter value is positive or negative, and the effective bit width represents the effective value of the model parameter. Data before the split point in the effective bit width is cleared to reduce the effective bit width, and this reduced effective bit width data is used as the model parameters for the second convolutional neural network model.

[0077] In one embodiment, model parameters with a relatively low percentage and an absolute value not less than a first preset initial value are quantized using a multi-segment quantization method. The multi-segment quantization method for quantizing model parameters includes:

[0078] Identify the flag, sign, and effective bit width of the model parameters, and set at least two split points on the effective bit width according to a preset step size. The flag indicates whether the model parameter value is a high-order or low-order value, the sign indicates whether the model parameter value is positive or negative, and the effective bit width indicates the effective value of the model parameter. Clear the data of the effective bit width before the last split point to reduce the effective bit width, and use the data with reduced effective bit width as the model parameters of the second convolutional neural network model.

[0079] The model parameters are set to a width of 16 bits, and the effects of the two-segment quantization method and the multi-segment quantization method are tested respectively. The effective bit width of the model parameter data is gradually decreased starting from 8 bits until the accuracy loss is greater than 2%. The number of segments in the multi-segment quantization is gradually increased, and the number of segments can be an integer power of 2, such as 3 segments, 5 segments, 6 segments, etc.

[0080] Please refer to Figure 5 This diagram illustrates the data format representation of model parameters in one embodiment. The bi-segment quantization method allocates as many quantization points as possible based on the frequency of parameters with higher absolute values ​​and larger absolute values. Similarly, it allocates the same number of quantization center points to parameters with larger absolute values. Taking a 16-bit unsigned number as an example, the 7th and 8th bits are truncated; this position is called the truncation position. The two segments after truncation should have equal widths (8 bits each) to ensure data format uniformity. This width is called the effective bit width. In addition, there should be one sign bit to indicate the sign of the data and one flag bit to indicate whether the data is Q0.8 (high 8 bits) or Q-8.16 (low 8 bits).

[0081] Once the effective bit width and split points of the model parameters are determined, the quantization parameters are also limited to a finite range of values. Bisegmented quantization, starting from the numerical representation range and precision of fixed-point numbers, provides a solution for low-bit-width quantization of convolutional neural networks by using half the effective bit width to represent the original bit-width fixed-point number. The specific split positions and effective bit widths need to be determined through an iterative process.

[0082] Please refer to Figure 6 This diagram illustrates the segmentation position setting in a multi-segment quantization method in one embodiment. Further segmenting the model parameters into multiple segments can further reduce the bit width of the parameters. Correspondingly, the truncation error increases compared to two-segment quantization, and the flag bit will be greater than 1 bit. For example, in four-segment quantization, the parameters can be divided into four evenly spaced segments, each with four bits. The effective bit width of each segment can overlap, thereby increasing the effective bit width, reducing the truncation error, and increasing the accuracy of the quantized parameters. The cost of multi-segment quantization is an increase in the flag bit width. Simultaneously, the segmentation position no longer needs to be evenly divided. The advantage of this is that the quantization details can differ in different layers of different networks, allowing for the selection of a more suitable quantization method for different layers. Please refer to [reference needed]. Figure 7 This is a schematic diagram of the data format representation of model parameters after multi-segment quantization in one embodiment. The data format of multi-segment parameters is similar to that of two-segment parameters, except that the width of the flag bit is greater than 1 bit, and the specific width d is determined by the number of segments N.

[0083] Step 104: Replace the model parameters.

[0084] The model parameters of the first convolutional neural network model are replaced with the model parameters obtained after quantization to obtain a second convolutional neural network model. The second convolutional neural network model is used to extract second image features. The first image features are different from the second image features.

[0085] In one embodiment, the model inference method disclosed in this application further includes:

[0086] Step 105: Obtain the accuracy loss value.

[0087] The second convolutional neural network model is trained using a training dataset containing features of the second image to obtain the accuracy loss value of the second convolutional neural network model.

[0088] Step 106: Reset the threshold range.

[0089] When the accuracy loss value is greater than a first preset value, the first threshold interval or the second threshold interval is changed, and steps 102 to 105 are repeated. When the accuracy loss value is not greater than the first preset value, the second convolutional neural network model is output.

[0090] Step 107: Re-quantize.

[0091] When the accuracy loss value is greater than a second preset value, the position of the segmentation point in the two-segment quantization method or the multi-segment quantization method is changed, and steps 102 to 105 are repeated. When the accuracy loss value is not greater than the second preset value, the second convolutional neural network model is output. In practical applications, during the quantization process of each neural network model, it is necessary to perform an accuracy test on the current quantization result. That is, the quantized neural network model and the original floating-point model are tested for recognition accuracy on the same dataset to determine whether the accuracy loss is within an acceptable range. In one embodiment, the design goal is to keep the accuracy loss within 2%. If the accuracy loss is too high, the quantization strategy is adjusted, the effective bit width is increased, or the segmentation position is changed; if the accuracy loss is acceptable, the effective bit width is further reduced.

[0092] The following is an example of an application scenario for the model inference method disclosed in this application:

[0093] The first convolutional neural network (CNN) model acquires the first image feature to identify "chair," while the second CNN model acquires the second image feature to identify "chair with a backrest." In existing technologies, even if the first and second CNN models have the same structure, the entire second CNN model needs to be retrained. However, the model inference method disclosed in this application leverages the high similarity between the first and second CNN models by quantizing the model parameters of the first CNN model to obtain its own parameters. Then, a precision loss test is performed using a small training dataset. Training of the second CNN model is complete when the precision loss falls below a preset value.

[0094] Please refer to Figure 8 This is a structural block diagram of a model inference device in one embodiment. The model inference device is used to apply the model inference method described above, specifically including a parameter extraction module 100, a parameter pruning module 200, a quantization processing module 300, and a model acquisition module 400. The parameter extraction module 100 is used to extract the model parameters of a first convolutional neural network model that has been trained. The first convolutional neural network model is used to extract first image features. The model parameters of the first convolutional neural network model include weights. The parameter pruning module 200 is used to prune the extracted model parameters of the first convolutional neural network model to remove model parameters whose values ​​are in a first threshold range and retain model parameters whose values ​​are in a second threshold range. The range of the second threshold range is larger than that of the first threshold range. The quantization processing module 300 is used to quantize each model parameter to obtain the values ​​of the flag bit, sign bit, and effective width bit of each model parameter. The effective width bit of the model parameter is a fixed-point value obtained by quantizing the model parameter using a two-segment quantization method or a multi-segment quantization method. The model acquisition module 400 is used to replace the model parameters of the first convolutional neural network model with the model parameters obtained after quantization to obtain the second convolutional neural network model. The second convolutional neural network model is used to extract the second image features.

[0095] The model inference method and apparatus disclosed in this application first extracts the model parameters of a first convolutional neural network model that has already been trained; then, the extracted model parameters are pruned to remove those with values ​​in a first threshold range and retain those with values ​​in a second threshold range; next, each model parameter is quantized to obtain the values ​​of its flag bit, sign bit, and effective width bit; wherein, the effective width bit of the model parameter is a fixed-point value obtained by quantizing the model parameters using a bi-segment quantization method or a multi-segment quantization method; finally, the quantized model parameters are used to replace the original model parameters to obtain a second convolutional neural network model. Since the model parameters of the new convolutional neural network model are obtained based on the already trained model parameters, the training efficiency of the new convolutional neural network model is greatly improved.

[0096] Those skilled in the art will understand that all or part of the functions of the various methods in the above embodiments can be implemented by hardware or by computer programs. When all or part of the functions in the above embodiments are implemented by computer programs, the program can be stored in a computer-readable storage medium, which may include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the program is executed by a computer to achieve the above functions. For example, the program can be stored in the memory of a device, and when the program in the memory is executed by the processor, all or part of the above functions can be achieved. In addition, when all or part of the functions in the above embodiments are implemented by computer programs, the program can also be stored in a server, another computer, disk, optical disk, flash drive, or external hard drive, etc., and can be downloaded or copied to the memory of a local device, or the system of the local device can be updated. When the program in the memory is executed by the processor, all or part of the functions in the above embodiments can be achieved.

[0097] The above examples illustrate the present invention only to aid in understanding it and are not intended to limit the scope of the invention. Those skilled in the art can make various simple deductions, modifications, or substitutions based on the principles of this invention.

Claims

1. A model inference method for convolutional neural networks, characterized in that, include: Step 101: Extract the model parameters of the first convolutional neural network model that has been trained; the first convolutional neural network model is used to extract the first image features; the model parameters of the first convolutional neural network model include weights. Step 102: Prune the model parameters of the extracted first convolutional neural network model to remove model parameters whose values ​​are in the first threshold interval and retain model parameters whose values ​​are in the second threshold interval; the first threshold interval is within the range of the second threshold interval. Step 103: Quantize each of the model parameters to obtain the values ​​of the flag bit, sign bit, and effective width bit of each model parameter; wherein, the value of the effective width bit of the model parameter is a fixed-point value, which is obtained by quantizing the model parameter using a two-segment quantization method or a multi-segment quantization method; Step 104: Replace the model parameters of the first convolutional neural network model with the model parameters obtained after quantization to obtain a second convolutional neural network model; the second convolutional neural network model is used to extract the second image features; The first image feature is different from the second image feature; The binary quantization method is applied to quantize model parameters whose absolute values ​​are less than a first preset initial value and have a relatively high proportion. The bi-segment quantization method quantizes the model parameters by including: The model parameter is identified by a flag bit, a sign bit, and a valid bit width, and the middle of the valid bit width is set as a split point; the flag bit is used to indicate whether the value of the model parameter is a high-order bit value or a low-order bit value, the sign bit is used to indicate whether the value of the model parameter is positive or negative, and the valid bit width is used to indicate the valid value of the model parameter; Clear the data before the split point to reduce the effective bit width, and use the data with reduced effective bit width as the model parameters of the second convolutional neural network model; The multi-segment quantization method is applied to quantize model parameters with a relatively low percentage and an absolute value not less than the first preset initial value. The multi-segment quantization method quantizes the model parameters by including: The model parameter is identified by a flag bit, a sign bit, and an effective bit width, and at least two split points are set on the effective bit width according to a preset step size; the flag bit is used to indicate whether the value of the model parameter is a high-order bit value or a low-order bit value, the sign bit is used to indicate whether the value of the model parameter is positive or negative, and the effective bit width is used to indicate the effective value of the model parameter; The data with effective bit width before the last split point of the model parameters are cleared to reduce the effective bit width, and the data with reduced effective bit width is used as the model parameters of the second convolutional neural network model.

2. The model reasoning method as described in claim 1, characterized in that, Also includes: Step 105: Train the second convolutional neural network model using a training dataset containing the features of the second image to obtain the accuracy loss value of the second convolutional neural network model. Step 106: When the accuracy loss value is greater than a first preset value, change the first threshold range or the second threshold range, and repeat steps 102 to 105. When the accuracy loss value is not greater than the first preset value, the second convolutional neural network model is output.

3. The model reasoning method as described in claim 1, characterized in that, Also includes: The second convolutional neural network model is then trained using a training dataset containing the features of the second image to obtain the accuracy loss value of the second convolutional neural network model. When the accuracy loss value is greater than a second preset value, the position of the segmentation point in the two-segment quantization method or the multi-segment quantization method is changed, and steps 102 to 105 are repeated. When the accuracy loss value is not greater than the second preset value, the second convolutional neural network model is output.

4. A computer-readable storage medium, characterized in that, The medium stores a program that can be executed by a processor to implement the model reasoning method as described in any one of claims 1 to 3.

5. A model inference device for convolutional neural networks, characterized in that, For applying the model inference method as described in any one of claims 1 to 3, the model inference apparatus comprises: The parameter extraction module is used to extract the model parameters of the first convolutional neural network model that has been trained; the first convolutional neural network model is used to extract the first image features; the model parameters of the first convolutional neural network model include weights. The parameter pruning module is used to prune the model parameters of the extracted first convolutional neural network model to remove model parameters whose values ​​are in the first threshold range and retain model parameters whose values ​​are in the second threshold range; the first threshold range is within the range of the second threshold range. The quantization processing module is used to quantize each of the model parameters to obtain the values ​​of the flag bit, sign bit, and effective width bit of each model parameter; wherein, the value of the effective width bit of the model parameter is a fixed-point value, which is obtained by quantizing the model parameter using a two-segment quantization method or a multi-segment quantization method; The model acquisition module is used to replace the model parameters of the first convolutional neural network model with the model parameters obtained after quantization to obtain a second convolutional neural network model; the second convolutional neural network model is used to extract second image features.

Citation Information

Patent Citations

  • Signal recognition convolutional neural network convolution kernel partition pruning method

    CN111222640A

  • Neural network model reasoning acceleration method and device, electronic equipment and medium

    CN115526320A