A Method and Device for Accelerating Dynamic Inference of Convolutional Neural Network

By integrating convolutional and pooling layer operations in CNNs, the method reduces unnecessary computations and improves inference efficiency, addressing inefficiencies in existing dynamic CNN acceleration techniques.

CN114202058BActive Publication Date: 2025-07-15CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111406598.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-24
Publication Date
2025-07-15
Estimated Expiration
2041-11-24

AI Technical Summary

Technical Problem

The existing convolutional neural network inference methods are inefficient in devices with limited computing power and storage space, and the existing dynamic network methods are highly complex in the training and inference stages, requiring additional computational overhead and structural modification.

Method used

By fusing the convolution layer with the adjacent maximum/minimum pooling layer, unnecessary convolution calculations are reduced, key parameters and formulas are used to estimate the sum of the maximum to be calculated, the filter core is sorted, and the output feature map is generated.

Benefits of technology

Without increasing computing overhead, the efficiency and accuracy of network inference are improved, redundant computing is reduced, and the performance of computing devices is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114202058B_ABST
    Figure CN114202058B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and apparatus for accelerating dynamic inference of a convolutional neural network. The method includes: based on key parameters, applying the convolution kernel of a filter to a sliding window corresponding to an input channel of an input feature map and generating a partial sum of the sliding window, and representing the activation value of an activation feature map through a specific formula based on the partial sum, where the output feature map of a convolutional layer is referred to as an activation feature map; estimating the maximum partial sum to be calculated corresponding to the activation value according to the maximum partial sum to be calculated of a corresponding kernel in the filter and the number of non-zero values of each cell in a preset input feature map corresponding to the kernel; and sorting the kernels of the filter based on the maximum partial sum to be calculated corresponding to the activation value to generate an output feature map. The present invention realizes the operation of the maximum / minimum pooling layer while performing the convolution process calculation, so as to reduce unnecessary convolution calculations and improve the efficiency of network inference at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of dynamic acceleration algorithms for convolutional neural network inference, and particularly relates to a method and device for dynamically accelerating the inference of a convolutional neural network. Background Art

[0002] Deep learning methods, especially convolutional neural networks, have been proven to be effective in various computer vision tasks such as image classification, face recognition, and object detection. The reliable results generated by convolutional neural networks rely on many parameters and complex calculations, which limits their use in devices with limited computing power and storage space. Therefore, research on effectively accelerating the calculation of convolutional neural networks and compressing the models has emerged continuously. Some research shows that there are a large number of redundant parameters in deep neural network models, and the calculation acceleration and model compression can be achieved by removing the redundant parameters. Based on this premise, many effective methods for accelerating and compressing by removing the redundant parameters in the neural network have been proposed in the past work. These works achieve model compression and calculation acceleration by identifying and permanently deleting unimportant connections in the neural network, thereby reducing the number of calculations in the training and inference stages.

[0003] However, the above-mentioned methods all perform the inference stage of the neural network in a static manner. Once the training stage is completed, their computational graphs and network parameters are fixed, which will limit their efficiency, representational ability, etc. Contrary to static neural networks, dynamic neural networks can adjust their structures and parameters during the inference stage to adapt to the input, so they can enjoy beneficial features that are not available in static models. Combining the advantages of dynamic execution, applying dynamic neural network methods to accelerate neural networks is an emerging topic in the field of deep learning. Compared with static neural networks, their computational graphs and network parameters can be changed during the inference process, making a better compromise between efficiency and consumption. However, in order to achieve the effects of dynamic parameters and computational graphs during the inference process, most of the existing dynamic network works need to make major modifications to the original network structure or introduce other structures that need to be trained during the training stage, which increases the complexity.

[0004] The channel gating network works by using partial weights in each convolutional layer and only activating the remaining weights in selected regions, which is equivalent to having two computational execution paths. One is the basic path that must perform calculations, while the other is selectively calculated. Dynamic neural architecture search uses neural architecture search to find the optimal architecture among multiple structures with different widths, which can enable any instance to obtain reliable prediction results at an early stage of the inference phase. Different from the method that relies on the activation values of the previous layer, runtime neural pruning adaptively prunes according to the input image and feature map, models layer-by-layer pruning as a Markov decision process, and uses reinforcement learning for training. In dynamic channel pruning, not only a new residual block is introduced to gate the channels in a fine-grained manner, but also a tool - batch reshaping is introduced to force the gating function to be more executable on the feature map.

[0005] Different from the previous work of skipping layers or filters, some work combines the work of skipping layers and filters at the same time, and this method can reduce more computational volume. These works are more flexible to adapt based on the computational graph, and only selectively activate channels when it is determined to execute a specific layer. However, this approach brings more challenges to the training and inference phases and adds more additional structures to the network. The research puts the module into another network with the same structure as the backbone network, which means that the additional network obtains the same input instance as the backbone network and generates all channel selection decisions, and then returns to the backbone network for inference. This method requires an additional training network and a redesigned training phase. Another method based on feature activation does not require additional auxiliary structures and computational overhead, but requires introducing a regularization term in the training phase to increase the sparsity of intermediate features, which brings additional complexity to the training phase.

[0006] While achieving good results, the above work also introduces other problems that need to be solved. Summary of the Invention

[0007] The present invention aims to solve at least one of the technical problems in the related art to some extent.

[0008] To this end, the present invention proposes a method for accelerating the dynamic inference of a convolutional neural network. Aiming at the overhead of retraining introduced by the above-mentioned prior art during inference acceleration, the purpose of the present invention is to achieve the dynamic inference acceleration of the neural network without introducing additional computational overhead and while ensuring the inference accuracy. This method studies the output of the convolutional layer and the input of the adjacent max / min pooling layer, and fuses the calculation of the convolutional layer with the pooling operation of the adjacent max / min pooling layer, realizing the operation of the max / min pooling layer while performing the convolutional process calculation, so as to reduce unnecessary convolutional calculations and improve the efficiency of network inference at the same time.

[0009] Another object of the present invention is to provide a convolutional neural network dynamic inference acceleration device.

[0010] To achieve the above object, on the one hand, the present invention provides a convolutional neural network dynamic inference acceleration method, including: based on key parameters, applying the convolutional kernel of a filter to a sliding window of a corresponding input channel of an input feature map, and generating a partial sum of the sliding window; based on the partial sum, representing the activation value of an activation feature map through a specific formula; wherein, the output feature map of a convolutional layer is referred to as an activation feature map; estimating the maximum partial sum to be calculated corresponding to the activation value according to the maximum partial sum to be calculated of a corresponding kernel in the filter and the number of non-zero values of each cell in a preset input feature map corresponding to the kernel; and sorting the kernels of the filter based on the maximum partial sum to be calculated corresponding to the activation value to generate an output feature map.

[0011] In the convolutional neural network dynamic inference acceleration method according to an embodiment of the present invention, based on key parameters, applying the convolutional kernel of a filter to a sliding window of a corresponding input channel of an input feature map, and generating a partial sum of the sliding window; based on the partial sum, representing the activation value of an activation feature map through a specific formula; wherein, the output feature map of a convolutional layer is referred to as an activation feature map; estimating the maximum partial sum to be calculated corresponding to the activation value according to the maximum partial sum to be calculated of a corresponding kernel in the filter and the number of non-zero values of each cell in a preset input feature map corresponding to the kernel; and sorting the kernels of the filter based on the maximum partial sum to be calculated corresponding to the activation value to generate an output feature map. By studying the output of a convolutional layer and the input of an adjacent max / min pooling layer, the present invention integrates the calculation of the convolutional layer with the pooling operation of the adjacent max / min pooling layer, realizes the operation of the max / min pooling layer while performing the calculation of the convolutional process, reduces unnecessary convolutional calculations, and improves the efficiency of network inference at the same time.

[0012] In addition, the convolutional neural network dynamic inference acceleration method according to the above embodiment of the present invention may further have the following additional technical features:

[0013] Further, in an embodiment of the present invention, the method further includes: predefined the key parameters, including: the number of input channels C I , the number of output channels C O , the width W of the input feature map I , the height H of the input feature map I , the width W of the filter in the weights K , the height H of the filter in the weights K .

[0014] Further, in an embodiment of the present invention, the representing the activation value of the activation feature map through a specific formula includes:

[0015] By separating the cumulative portions of the first c input channels to represent activation values:

[0016]

[0017] where A K is the k-th activation value in the activation feature map, AccuPsum k c is the cumulative portion of the first c computational input channels, and LeftPsum k c is the remaining cumulative portion of more than c input channels, and ReLU is the rectified linear unit.

[0018] Furthermore, in an embodiment of the present invention, the maximum partial sum to be calculated for the kernel corresponds to the input channels from the (c + 1)-th to the C I - 1 input channels; the number of non-zero values is 2 N - 1; the maximum partial sum to be calculated corresponding to the activation value A K is determined by the l1-norm of the kernel of the filter as follows:

[0019]

[0020] where α represents a scaling factor, and L c represents the sum of the l1-norms of the kernel of the filter from the c-th to the C I - 1 input channels.

[0021] Furthermore, in an embodiment of the present invention, the sorting of the kernel of the filter includes:

[0022] Sorting based on the l1-norm values of the kernels from the c-th input channel to the C I - 1 input channels according to the kernel order in the filter.

[0023] Furthermore, in an embodiment of the present invention, the method further includes:

[0024] Setting an index for the input feature map corresponding to each kernel before sorting, and after re-sorting the kernels, each kernel uses the index to find the corresponding input feature map for calculation.

[0025] Furthermore, in an embodiment of the present invention, the method further includes: determining whether to skip the convolution calculation operation of the activation values in the pooling window of the activation feature map through a determination formula, and the determination formula is:

[0026]

[0027] Among them, AccuPsum K C and AccuPsum j C respectively represent the accumulated partial sums from channel 0 to channel c corresponding to the k-th and j-th activation values in a pooling window.

[0028] Furthermore, in an embodiment of the present invention, according to the determination formula, the accumulated partial sums corresponding to each pair of activation values are compared in the same pooling window.

[0029] Furthermore, in an embodiment of the present invention, the comparing the accumulated partial sums corresponding to each pair of activation values in the same pooling window according to the determination formula includes: on each input feature map, a 3*3 convolution operation is performed on a 4*4 sliding window to obtain 4 activation values corresponding to a 2*2 pooling window. The 4*4 sliding window moves from left to right and from top to bottom, with a stride equal to the size of the pooling window. When one input feature map channel is completed, the 4*4 sliding window will move to the next input feature map; wherein, the convolution calculation according to the determination formula in the 4*4 sliding window is skipped.

[0030] To achieve the above object, on the other hand, the present invention proposes a convolutional neural network dynamic inference acceleration device, including: a generation module, configured to act on a sliding window of a corresponding input channel of an input feature map with a convolution kernel of a filter based on key parameters, and generate a partial sum of the sliding window, and represent an activation value of an activation feature map through a specific formula based on the partial sum; wherein, the output feature map of the convolutional layer is called the activation feature map; an estimation module, configured to estimate the maximum partial sum to be calculated corresponding to the activation value according to the maximum partial sum to be calculated of the corresponding kernel in the filter and the number of non-zero values corresponding to the kernel for each cell in the preset input feature map; a sorting module, configured to sort the kernels of the filter based on the maximum partial sum to be calculated corresponding to the activation value to generate an output feature map.

[0031] The convolutional neural network dynamic inference acceleration device according to the embodiment of the present invention includes a generation module, which is used to act on the sliding window of the input feature map corresponding to the input channel of the filter based on key parameters and generate partial sums of the sliding window, and represent the activation values of the activation feature map through a specific formula; wherein, the output feature map of the convolutional layer is called the activation feature map; an estimation module, which is used to estimate the maximum partial sum to be calculated corresponding to the activation value according to the maximum partial sum to be calculated of the corresponding kernel in the filter and the number of non-zero values of each cell in the preset input feature map corresponding to the kernel; a sorting module, which is used to sort the kernels of the filter based on the maximum partial sum to be calculated corresponding to the activation value to generate the output feature map. The present invention studies the output of the convolutional layer and the input of the adjacent max / min pooling layer, and integrates the calculation of the convolutional layer with the pooling operation of the adjacent max / min pooling layer, so as to complete the operation of the max / min pooling layer during the convolutional process, reduce unnecessary convolutional calculations, and improve the efficiency of network inference at the same time.

[0032] The beneficial effects of the present invention are as follows:

[0033] By combining the structure of the network itself and the characteristics of the dynamic network, the acceleration exploration in the inference stage of the convolutional neural network will be realized. From the perspective of calculation, the calculation of the maximum partial sum to be calculated and the order of the kernels in the filter participating in the calculation are explored, and finally the acceleration gain brought by the whole algorithm is improved.

[0034] By exploring the scale factor, different improvements are made to the size of the maximum partial sum to be calculated, so as to find a more favorable solution for different situations in the search space, and finally improve the effect of the whole design.

[0035] The additional aspects and advantages of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. Description of the Drawings

[0036] The above and / or additional aspects and advantages of the present invention will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0037] Figure 1 It is a flowchart of the convolutional neural network dynamic inference acceleration method according to the embodiment of the present invention;

[0038] Figure 2 It is a schematic diagram of the core idea according to the embodiment of the present invention;

[0039] Figure 3 It is a schematic example diagram of sorting inside the filter according to the l1 norm according to the embodiment of the present invention;

[0040] Figure 4 Schematic diagram of the specific execution process according to an embodiment of the present invention;

[0041] Figure 5 Schematic diagram of the structure of the convolutional neural network dynamic inference acceleration device according to an embodiment of the present invention. Detailed implementation manners

[0042] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, in which the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present invention, but should not be construed as limiting the present invention.

[0043] The convolutional neural network dynamic inference acceleration method and device according to an embodiment of the present invention will be described below with reference to the accompanying drawings. First, the convolutional neural network dynamic inference acceleration method according to an embodiment of the present invention will be described with reference to the accompanying drawings.

[0044] The technical solution adopted by the present invention is a dynamic network inference method applied to a convolutional neural network having a convolutional layer - max / min pooling layer.

[0045] Figure 1 It is a flowchart of the convolutional neural network dynamic inference acceleration method according to an embodiment of the present invention.

[0046] As Figure 1 shown, the convolutional neural network dynamic inference acceleration method includes the following steps:

[0047] Step S1, based on key parameters, apply the convolution kernel of the filter to the sliding window corresponding to the input channels of the input feature map, and generate a partial sum of the sliding window. Based on the partial sum, represent the activation value of the activation feature map through a specific formula; wherein, the output feature map of the convolutional layer is called the activation feature map.

[0048] Specifically, the core idea will be described correspondingly first. Before describing the core idea, several important parameters involved in the description need to be defined in advance: the number of input channels C I , the number of output channels C O , the width W of the input feature map I , the height H of the input feature map I , the width W of the filter in the weight K , the height H of the filter in the weight K . To distinguish the output feature maps of different layers, the output feature map of the convolutional layer is called the activation feature map. The core idea is as Figure 2 shown. Each convolution kernel of a filter will act on the W K *H in the corresponding channel of the input feature mapK A sliding window is used to generate the partial sum (Psum) of the sliding window. Then, the partial sums of all input channels are accumulated to generate an activation value in the output channel. After calculating all the activation values in the output feature map, max / min pooling is usually performed on the activation feature map to generate the output feature map. Figure 2 Among them, (a) CI input feature maps, (b) 1 filter, (c) 1 activation feature map, (d) 1 output feature map, (e) the cumulative partial sum and the calculation process of the partial sum to be calculated.

[0049] Figure 3 As an exemplary schematic diagram of sorting inside the filter according to the l1 norm in an embodiment of the present invention, as Figure 3 shown, (a) the original kernel and the corresponding maximum partial sum to be calculated assuming a scale factor of 1; (b) the sorted kernel and the corresponding maximum partial sum to be calculated assuming a scale factor of 1.

[0050] From Figure 3 the convolution pooling operation process, the present invention notes that the value of the output feature map selects a value in the pooling window of the activation map. If the present invention can predict in advance which value in the pooling window is the maximum / minimum value, rather than calculating the partial sums corresponding to all input channels, this will save a lot of calculations and speed up the inference speed of the neural network. By separating the cumulative partial sums of the first c input channels, the activation value can be expressed as the following equation:

[0051]

[0052] In Equation 1, A K , AccuPsum k c , LeftPsum k c and ReLU respectively represent the kth activation value in the activation map, the cumulative partial sums of the first c calculated input channels, the remaining cumulative partial sums of more than c input channels, and the rectified linear unit.

[0053] As Figure 2 shown, as shown in (b), the yellow, red, green, and blue bars below the dotted line respectively represent the cumulative partial sum (AccuPsum) windows of the first three input channels at different positions in the feature map during the 2*2 pooling operation. Since the gaps (gap1, gap2, and gap3) between the cumulative partial sum shown by the yellow bar and the cumulative partial sums shown by the other 3 bars are large enough, and the partial sum to be calculated (LeftPsum) cannot offset these gaps, then for the activation values represented by the non-yellow boxes in the pooling window, the present invention does not need to perform HK*WK multiply-accumulate (MAC) operations in the remaining channels (as Figure 2as shown by the dashed box in

[0054] Step S2: According to the maximum partial sum to be calculated in the corresponding kernel of the filter, estimate the maximum partial sum corresponding to the activation value through the number of non-zero values corresponding to the kernel in each cell of the preset input feature map.

[0055] It can be understood that from Equation 1, it can be seen that the cumulative partial sum needs to be accurately calculated and compared to obtain the gap between different activation values. However, if the partial sum to be calculated can also be accurately calculated, then there is no chance to skip redundant multiply-accumulate (MAC) calculations. The present invention attempts to first calculate the maximum partial sum to be calculated (MaxLeftPsum) of the corresponding kernel in the filter (corresponding to the input channels from the (c + 1)-th input channel to the C I -1 input channels). Since the value ranges of each cell in the input feature map and the filter are limited by the data / weight bit width (N) in the quantized neural network (from 0 to 2 N -1), it is possible to easily estimate the maximum partial sum to be calculated by assuming that the number of non-zero values corresponding to each cell in the input feature map for the kernel is set to 2 N -1. Therefore, the maximum partial sum to be calculated corresponding to any activation value AK can be determined by the l1-norm of the kernel of the filter as follows:

[0056]

[0057] where α and L c respectively represent the scale factor and the sum of the l1-norms of the kernels of the filter from the c-th to the C I -1 input channels. In theory, α should be set to 2 N -1 to ensure that the maximum partial sum to be calculated in Equation (2) is the maximum value within the value range. However, in this case, it is difficult for the gap between the cumulative partial sums to exceed the maximum partial sum to be calculated, which will result in fewer skipped calculations and thus poor acceleration effect. In most cases, the values in the input feature map will not be close to 2 N -1, which depends on the data distribution. Therefore, the present invention will use the validation dataset to explore the scale factor of each structure.

[0058] Since each weight value in the kernel of the filter is fixed during the inference stage, the maximum partial sum to be calculated for each channel can be pre-calculated and stored together with each kernel in the filter, as Figure 2 shown by the rhombus in (b). Therefore, the present invention does not require the additional overhead of calculating the maximum partial sum to be calculated during the inference stage.

[0059] Step S3, sorting the filter kernels based on the maximum to-be-calculated partial sum corresponding to the activation value to generate an output feature map.

[0060] It can be understood that since a smaller maximum partial sum to be calculated helps save more redundant calculations, the present invention will further reduce the maximum partial sum to be calculated according to the kernel order in the filter from the cth input channel to the CI-1th input channel. If the kernel with a smaller l1 norm value is sorted in the filter with a lower calculation priority, the corresponding maximum partial sum to be calculated will be reduced. Therefore, the kernels in the filter are sorted according to their θ1 norm values.

[0061] like Figure 3 As shown in (a), assuming the scale factor is 1, the maximum partial sums to be calculated in channels 0, 1, 2, and 3 of the original filter are 52, 40, 26, and 16, respectively. After the kernels in the filter are reordered, the maximum partial sums to be calculated in channels 0, 1, 2, and 3 become 52, 36, 22, and 10, respectively, as shown in Figure 3 , as shown in (b). Therefore, the gap between the accumulated partial sums will be larger than the corresponding maximum partial sums to be calculated, which can further save more multiplication and addition calculations.

[0062] Since the weights are fixed during the inference phase, kernel reordering can be performed statically without runtime overhead. However, since different filters may have different kernel orderings, requiring different channel orders of the input feature maps, if there is output channel parallelism, it will inevitably destroy the data locality of the input feature map access. Therefore, the method proposed in the present invention has less benefits in a computing system with output channel parallelism. In order to ensure the correctness of the calculation, an index is set for the input feature map corresponding to each kernel before sorting, so that after the kernels are reordered, each kernel can use the index to find the corresponding input feature map for calculation.

[0063] The embodiments of the present invention are further described below in conjunction with the accompanying drawings.

[0064] Step 1: In combination with the calculation method of the maximum partial sum to be calculated described in detail above, the present invention can determine whether the convolution calculation operation of the activation value in the pooling window should be skipped by the following equation:

[0065]

[0066] In this formula and They represent the accumulated partial sums from channels 0 to c corresponding to the kth and jth activation values in a pooling window, respectively. When the condition of equation (3) is met, the remaining uncalculated convolution operation of the jth activation value can be terminated early.

[0067] Step 2: According to the equation in Step 1, the present invention compares the accumulated partial sums corresponding to each pair of activation values in the same pooling window. That is to say, the convolution operation on each input channel of the present invention is based on the pooling window. As Figure 4 shown, on each input feature map, the 3*3 convolution operation is performed on a 4*4 sliding window (represented by the green border) to obtain 4 activation values corresponding to the 2*2 pooling window. Then, the 4*4 sliding window moves from left to right and from top to bottom, with the stride equal to the size of the pooling window, as Figure 4 shown by the green arrow in. When one input feature map channel is completed, the 4*4 sliding window will move to the next input feature map, as Figure 4 shown by the green dashed arrow in. During this process, the convolution calculation in the 4*4 sliding window will be skipped according to Equation 3. Figure 4 In, (a) the 0th input feature map, (b) the 1st input feature map, (c) the 2nd input feature map, (b) the 3rd input feature map, (e) 1 output activation map.

[0068] As an example, the present invention can be applied to speech recognition. Since the concept of deep learning was proposed, neural networks have once again come into people's view. Speech recognition is the first field to achieve a breakthrough. Using neural networks to change the feature extraction method in speech recognition greatly reduces the recognition error rate, and the simultaneous interpretation products applying this technology have amazing effects. After that, researchers further improved the method and used recurrent neural networks to improve the prediction and recognition of speech recognition, further reducing the error rate. This series of successes has made the practical application of speech recognition possible and inspired a large number of commercial applications. Since neural networks are very complex and the resources consumed by inference are extremely amazing, edge devices (such as mobile phones) used to deploy relevant networks often do not have such a large amount of resources. Therefore, after the neural network training is completed, it is often necessary to accelerate the network inference process. So far, the recognition effect of relevant products not only exceeds the level of human stenographers, but also consumes less resources.

[0069] As another example, the present invention can be applied to computer vision, which has always been a popular research field. Traditional research mainly designs different features manually according to the characteristics of images, such as edge features, color features, scale-invariant features, etc., and then uses these features to complete specific computer vision tasks, such as image classification, image clustering, image segmentation, object detection, etc. The features extracted by the above research methods rely on manual design, generally being relatively intuitive primary features with low abstraction level and weak expression ability. While the neural network method uses a large amount of image data to learn features completely automatically. In a deep neural network, the features of each layer form a hierarchical division of edges, lines, contours, shapes, objects, etc., and the abstraction level gradually increases. On the large-scale image dataset ImageNet, the neural network method has made a major breakthrough. And on the authoritative LFW face recognition evaluation database, the face recognition method DeepID based on a deep neural network has far exceeded the accuracy of human recognition. Products related to computer vision tasks often limit the time for the neural network to infer results. For example, a camera used to monitor vehicle flow on a highway needs to obtain recognition results within an extremely short time, and such results have real-time practical application significance, which requires restricting the inference time of the trained neural network and reducing the inference time while ensuring the inference accuracy. Currently, products for related computer vision tasks are being developed one after another.

[0070] According to the convolutional neural network dynamic inference acceleration method of an embodiment of the present invention, by based on key parameters, the convolutional kernel of a filter acts on a sliding window corresponding to an input channel of an input feature map, and a partial sum of the sliding window is generated. Based on the partial sum, the activation value of an activation feature map is represented by a specific formula; wherein, the output feature map of a convolutional layer is called an activation feature map; according to the maximum partial sum to be calculated of the corresponding kernel in the filter, the maximum partial sum to be calculated corresponding to the activation value is estimated through the number of non-zero values of each cell in a preset input feature map corresponding to the kernel; based on the maximum partial sum to be calculated corresponding to the activation value, the kernels of the filter are sorted to generate an output feature map. The present invention studies the output of a convolutional layer and the input of an adjacent maximum / minimum pooling layer, and fuses the calculation of the convolutional layer with the pooling operation of the adjacent maximum / minimum pooling layer, so as to complete the operation of the maximum / minimum pooling layer during the convolutional process calculation, reduce unnecessary convolutional calculations, and improve the efficiency of network inference at the same time.

[0071] Next, a schematic structural diagram of a convolutional neural network dynamic inference acceleration device according to an embodiment of the present invention is described with reference to the accompanying drawings.

[0072] Figure 5 It is a schematic structural diagram of a convolutional neural network dynamic inference acceleration device according to an embodiment of the present invention.

[0073] AsFigure 5 As shown in Figure 5 , the convolutional neural network dynamic inference acceleration device 10 includes: a generation module 100, an estimation module 200, and a sorting module 300.

[0074] The generation module 100 is configured to, based on key parameters, apply the convolution kernel of the filter to the sliding window of the corresponding input channel of the input feature map, generate partial sums of the sliding window, and represent the activation values of the activation feature map through a specific formula; wherein, the output feature map of the convolutional layer is referred to as the activation feature map.

[0075] The estimation module 200 is configured to estimate the maximum partial sum to be calculated corresponding to the activation value according to the maximum partial sum to be calculated of the corresponding kernel in the filter and the number of non-zero values of each cell in the preset input feature map corresponding to the kernel.

[0076] The sorting module 300 is configured to sort the kernels of the filter based on the maximum partial sum to be calculated corresponding to the activation value to generate the output feature map.

[0077] For the convolutional neural network dynamic inference acceleration device according to an embodiment of the present invention, the generation module is configured to, based on key parameters, apply the convolution kernel of the filter to the sliding window of the corresponding input channel of the input feature map, generate partial sums of the sliding window, and represent the activation values of the activation feature map through a specific formula; wherein, the output feature map of the convolutional layer is referred to as the activation feature map; the estimation module is configured to estimate the maximum partial sum to be calculated corresponding to the activation value according to the maximum partial sum to be calculated of the corresponding kernel in the filter and the number of non-zero values of each cell in the preset input feature map corresponding to the kernel; the sorting module is configured to sort the kernels of the filter based on the maximum partial sum to be calculated corresponding to the activation value to generate the output feature map. By studying the output of the convolutional layer and the input of the adjacent max / min pooling layer, the present invention integrates the calculation of the convolutional layer with the pooling operation of the adjacent max / min pooling layer, achieving the operation of the max / min pooling layer while performing the convolutional process calculation, reducing unnecessary convolutional calculations, and improving the efficiency of network inference at the same time.

[0078] It should be noted that the foregoing explanation of the embodiments of the convolutional neural network dynamic inference acceleration method is also applicable to this device and will not be elaborated here.

[0079] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0080] In the present invention, unless otherwise clearly specified or limited, the terms "mounted", "connected", "connected to", "fixed", etc. shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements or the interaction relationship between two elements, unless otherwise clearly limited. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0081] In the present invention, unless otherwise clearly specified or limited, the first feature being "on" or "under" the second feature may be that the first and second features are in direct contact, or the first and second features are indirectly in contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on top of" the second feature may be that the first feature is directly above or obliquely above the second feature, or simply means that the first feature is at a higher horizontal level than the second feature. The first feature being "under", "below" and "beneath" the second feature may be that the first feature is directly below or obliquely below the second feature, or simply means that the first feature is at a lower horizontal level than the second feature.

[0082] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0083] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for accelerating dynamic inference of a convolutional neural network, characterized in that Applied to the field of computer vision, including the following steps: Based on key parameters, the convolution kernel of the filter acts on the sliding window of the corresponding input channel of the input feature map of the image to be classified, and generates a partial sum of the sliding window. Based on the partial sum, the activation value of the activation feature map is represented by a specific formula; wherein, the output feature map of the convolutional layer is called the activation feature map; According to the maximum partial sum to be calculated of the corresponding kernel in the filter, estimate the maximum partial sum to be calculated corresponding to the activation value through the number of non-zero values of each cell in the preset input feature map corresponding to the kernel; Based on the maximum partial sum to be calculated corresponding to the activation value, sort the kernels of the filter to generate an output feature map for image classification of the image to be classified; The activation value of the activation feature map represented by the specific formula includes: By separating the cumulative partial sums of the first c input channels to represent the activation value: Among them, A K is the k-th activation value in the activation feature map, AccuPsum k c is the cumulative partial sum of the first c computational input channels, LeftPsum k c is the remaining cumulative partial sum of more than c input channels, and ReLU is the rectified linear unit; The maximum partial sum to be calculated of the kernel is for the input channels corresponding to the (c + 1)-th to the C I -1 input channel; the number of non-zero values is 2 N -1; the activation value A K The corresponding maximum partial sum to be calculated is determined by the 1-norm 1-norm as follows: where α represents a scaling factor, L c represents the sum of the 1-norms of the kernels of the input channels from the c-th to the C I -1 -th in the filter; 1-norm sum; The sorting of the kernels of the filter includes: From the c-th input channel to the C I -1 input channel based on the kernel order in the filter, and sort according to the L1 norm value of the kernel; The method further includes: Before sorting, set an index for the input feature map corresponding to each kernel. After the kernels are re-sorted, each kernel uses the index to find the corresponding input feature map for calculation.

2. The dynamic inference acceleration method for a convolutional neural network according to claim 1, wherein The method further includes: predefined the key parameters, including: the number of input channels C I , the number of output channels C O , the width of the input feature map W I , the height of the input feature map H I , the width of the filter in the weights W K , the height of the filter in the weights H K .

3. The dynamic inference acceleration method of the convolutional neural network according to claim 1, wherein The method further includes: determining whether to skip the convolution calculation operation of the activation value in the pooling window of the activation feature map through a determination formula, and the determination formula is: Among them, AccuPsum K C and AccuPsum j C respectively represent the cumulative partial sums from channel 0 to j corresponding to the k-th and the c -th activation values in a pooling window.

4. The convolutional neural network dynamic inference acceleration method according to claim 3, wherein According to the determination formula, compare the cumulative partial sums corresponding to each pair of activation values in the same pooling window.

5. The convolutional neural network dynamic inference acceleration method according to claim 4, characterized in that, The comparing the cumulative partial sums corresponding to each pair of activation values in the same pooling window according to the determination formula includes: On each input feature map, a 3*3 convolution operation is performed on a 4*4 sliding window to obtain 4 activation values corresponding to a 2*2 pooling window. The 4*4 sliding window moves from left to right and from top to bottom, and the stride is equal to the size of the pooling window. When an input feature map channel is completed, the 4*4 sliding window will move to the next input feature map; wherein, the convolution calculation in the 4*4 sliding window is skipped according to the determination formula.

6. A convolutional neural network dynamic inference acceleration device using the method according to claim 1, characterized in that, Applied to the field of computer vision, including: A generation module, configured to, based on key parameters, act on the sliding window of the corresponding input channel of the input feature map of the image to be classified with the convolution kernel of the filter, and generate a partial sum of the sliding window. Based on the partial sum, represent the activation value of the activation feature map by a specific formula; wherein, the output feature map of the convolutional layer is called the activation feature map; An estimation module, configured to estimate the maximum partial sum to be calculated corresponding to the activation value according to the maximum partial sum to be calculated of the corresponding kernel in the filter through the number of non-zero values of each cell in the preset input feature map corresponding to the kernel; A sorting module, configured to sort the kernels of the filter based on the maximum partial sum to be calculated corresponding to the activation value to generate an output feature map for image classification of the image to be classified.