A pruning method, device and equipment of a convolutional neural network model and a medium

CN116258173BActive Publication Date: 2026-10-09SHENZHEN SEICHITECH TECHN CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211479732.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-24
Publication Date
2026-10-09
Estimated Expiration
2042-11-24

AI Technical Summary

Technical Problem

[0005]基于神经元权重的剪枝这种剪枝方式将卷积神经网络模型的各个神经元按各自的权重进行排序,把权重排名较后的神经元去除,这种方式虽然利用了神经元权重信息,但是忽略了神经元输入输出信息,即上下文信息,神经元作为输入特征和输出特征的计算函数,其作用及重要性和上下文密切相关,因此单单考虑神经元自身的权重不能很好的反映神经元的重要性

Benefits of technology

本申请中,首先获取输入特征图像和卷积神经网络模型,卷积神经网络模型包括特征提取模块和注意力生成模块,其中注意力生成模块可以是位于卷积神经网络模型内部,与其他工作层连接使用,也可以是位于卷积神经网络模型外部独立使用。输入特征图像为输入卷积神经网络模型中进行训练中的图像。将输入特征图像输入卷积神经网络模型的特征提取模块。通过特征提取模块中的N个卷积核对原始图像进行图像特征处理,将N个卷积核对应的N个特征通道进行通道融合,生成输出特征图。将输出特征图输入卷积神经网络模型的注意力生成模块,再通过注意力生成模块分析输出特征图,分别为N个特征通道生成N个注意力向量。通过注意力生成模块根据N个注意力向量分别计算N个特征通道的注意力值。根据N个特征通道的注意力值对特征提取模块中的N个卷积核做剪枝处理。本申请中,通过注意力生成模块根据N个注意力向量分别计算N个特征通道的注意力值,得到各个特征通道的注意力值,需要将注意力值小于预设阈值的特征通道对应的神经元进行删除,或者是将注意力值排在末尾的几个特征通道对应的神经元进行删除,已完成剪枝处理,这样的剪枝方式考虑到了图像各个通道所生成的数据的重要程度,使用卷积核输出的特征通道的注意力值为判断依据,以整个卷积核作为剪枝单元,将注意力值较小的特征通道对应的卷积核去除,达到神经网络模型压缩的目的。注意力能更好地反映对应卷积核的重要程度,以注意力为依据的剪枝方式,能更好的保留有用卷积核而删除无用卷积核,达到模型的轻量化的目的且尽可能小地影响模型性能,提高剪枝效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116258173B_ABST
    Figure CN116258173B_ABST
Patent Text Reader

Abstract

The application discloses a pruning method and device of a convolutional neural network model, equipment and a medium, and is used for improving pruning efficiency. The pruning method comprises the following steps: acquiring an input feature image and a convolutional neural network model; inputting the input feature image into a feature extraction module of the convolutional neural network model; performing image feature processing on the original image through N convolution kernels in the feature extraction module, fusing N feature channels corresponding to the N convolution kernels, and generating an output feature map; inputting the output feature map into an attention generation module of the convolutional neural network model; analyzing the output feature map through the attention generation module, and generating N attention vectors for the N feature channels respectively; calculating attention values of the N feature channels according to the N attention vectors respectively; and performing pruning processing on the N convolution kernels in the feature extraction module according to the attention values of the N feature channels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of convolutional neural network models, and more particularly to a pruning method, apparatus, device, and medium for convolutional neural network models. Background Technology

[0002] In recent years, deep learning has flourished in the field of image processing as an emerging technology. Its ability to autonomously learn image data features greatly avoids the tediousness of manually designing algorithms. Furthermore, it possesses accurate detection performance, high detection efficiency, and good generalization performance across various image tasks, leading to its widespread application in image processing, including image detection, image classification, and image reconstruction. Convolutional operations, as the core operator of deep learning in image processing, possess three major characteristics: local perception, weight sharing, and downsampling. Due to its excellent image feature extraction performance, it has become the cornerstone of deep learning's success in the image processing field.

[0003] Convolutional neural network (CNN) models are often extremely complex, consuming significant amounts of storage and computational resources, making it difficult to deploy deep learning models across various hardware platforms. For example, the VGG-16 CNN has over 130 million parameters, occupies over 500 MB of storage, and requires over 30 billion floating-point operations to complete a single image recognition task. Research shows that CNN models contain a large number of redundant neurons and weights, with only 5-10% of the weights participating in the main computation and influencing the final result. This provides a theoretical basis for neural network compression. If effective compression methods can be found, deep neural networks can be more widely deployed on lightweight devices such as mobile devices. On the server side, better performance can also be provided. Pruning, as an effective method for CNN model compression, is receiving increasing attention and application. CNN model pruning first filters out unimportant neurons and weights from a large network, then removes them from the network, while preserving as much performance as possible. Pruning techniques can be categorized into structured pruning and unstructured pruning based on their granularity. Structured pruning removes neurons (filters in convolutional neural networks) as the basic unit. Because it directly prunes neurons, structured pruning models can achieve significant inference acceleration and storage advantages under current hardware conditions. However, its drawback is the large granularity of pruning, which often significantly impacts the accuracy of the compressed model. Unstructured pruning removes individual weights as the basic unit. The resulting model experiences less accuracy loss, but it ultimately produces a sparse weight matrix, requiring robust support from underlying hardware and computing libraries to achieve inference acceleration and storage advantages. Existing pruning methods either prune based on the neuron weights of the convolutional neural network model or perform pruning randomly. This approach cannot accurately determine the importance of neurons, and the pruning operation cannot effectively remove useless neurons, leading to a significant performance degradation in convolutional neural networks. Currently, the mainstream neural network pruning methods fall into two categories: random pruning and neuron weight-based pruning.

[0004] Random pruning removes neurons from a convolutional neural network model without considering the importance of each neuron in the overall model. As the model size decreases, the performance of the convolutional neural network model may also decline significantly.

[0005] Pruning based on neuron weights is a pruning method that sorts the neurons in a convolutional neural network model according to their respective weights and removes neurons with lower weight rankings. Although this method utilizes neuron weight information, it ignores the neuron's input and output information, i.e., contextual information. As a computational function of input and output features, the role and importance of neurons are closely related to the context. Therefore, simply considering the weight of a neuron itself cannot well reflect the importance of the neuron.

[0006] In summary, existing convolutional neural network (CNN) pruning methods fail to adequately consider the importance of neurons, leading to performance instability after pruning and necessitating repeated pruning. This increases the construction time of the CNN and reduces pruning efficiency. Summary of the Invention

[0007] This application discloses a pruning method, apparatus, device, and medium for convolutional neural network models, which can improve pruning efficiency.

[0008] Specifically, this application proposes a novel structural pruning method for neural network convolutional kernels based on output feature channel attention. It uses the attention value of the feature channels output by the convolutional kernel as the criterion, treating the entire convolutional kernel as a pruning unit, removing convolutional kernels corresponding to feature channels with smaller attention values, thereby achieving neural network model compression. Attention better reflects the importance of the corresponding convolutional kernel; this attention-based pruning method can better retain useful convolutional kernels while removing useless ones, achieving model lightweighting and minimizing the impact on model performance.

[0009] The first aspect of this application provides a pruning method for a convolutional neural network model, including: The input feature image and convolutional neural network model are obtained, wherein the convolutional neural network model includes a feature extraction module and an attention generation module; The input feature image is input into the feature extraction module of the convolutional neural network model; The original image is processed by N convolutional kernels in the feature extraction module, and the N feature channels corresponding to the N convolutional kernels are fused to generate an output feature map. The output feature map is input into the attention generation module of the convolutional neural network model; The attention generation module analyzes the output feature map and generates N attention vectors for the N feature channels respectively. Calculate the attention values ​​for the N feature channels based on the N attention vectors respectively; The N convolutional kernels in the feature extraction module are pruned based on the attention values ​​of the N feature channels.

[0010] Optionally, the step of analyzing the output feature map through the attention generation module to generate N attention vectors for the N feature channels includes: The attention generation module performs feature channel compression on the N feature channels in the output feature map to generate compressed features for the N feature channels. The attention generation module performs pooling operations on the compressed features to generate pooled data for the N feature channels. The attention generation module calculates the attention vectors for each of the N feature channels based on the pooling data.

[0011] Optionally, the attention generation module includes a Conv-ReLU layer, a global average pooling layer, and a Conv-SigMoid function layer; The step of compressing the N feature channels in the output feature map using the attention generation module to generate compressed features for the N feature channels includes: The Conv-ReLU layer is used to compress the N feature channels in the output feature map to generate compressed features for the N feature channels.

[0012] Optionally, the step of performing pooling operations on the compressed features through the attention generation module to generate pooled data for the N feature channels includes: The compressed features are subjected to global average pooling operations through the global average pooling layer to generate pooled data for the N feature channels.

[0013] Optionally, the step of calculating the attention vectors for the N feature channels respectively by the attention generation module based on the pooled data includes: The Conv-SigMoid function layer calculates the attention vectors for each of the N feature channels based on the pooling data.

[0014] Optionally, the step of pruning the N convolutional kernels in the feature extraction module based on the attention values ​​of the N feature channels includes: Identify the set of feature channels whose attention values ​​are below a preset value; Pruning is performed on the convolution kernels corresponding to the feature channel set below the preset value in the feature extraction module.

[0015] Optionally, after pruning the N convolutional kernels in the feature extraction module based on the attention values ​​of the N feature channels, the pruning method further includes: The input feature image is then re-input into the convolutional neural network model for training.

[0016] A second aspect of this application provides a pruning device for a convolutional neural network model, comprising: An acquisition unit is used to acquire an input feature image and a convolutional neural network model, wherein the convolutional neural network model includes a feature extraction module and an attention generation module; The first input unit is used to input the input feature image into the feature extraction module of the convolutional neural network model; The first generation unit is used to perform image feature processing on the original image through N convolutional kernels in the feature extraction module, and to perform channel fusion on the N feature channels corresponding to the N convolutional kernels to generate an output feature map; The second input unit is used to input the output feature map into the attention generation module of the convolutional neural network model; The second generation unit is used to analyze the output feature map through the attention generation module and generate N attention vectors for the N feature channels respectively; A calculation unit is used to calculate the attention values ​​of the N feature channels based on the N attention vectors respectively; The pruning unit is used to prune the N convolutional kernels in the feature extraction module based on the attention values ​​of the N feature channels.

[0017] Optionally, the second generating unit includes: The compression module compresses the N feature channels in the output feature map by means of the attention generation module, thereby generating compressed features for the N feature channels. The pooling module generates pooled data for the N feature channels by performing pooling operations on the compressed features through the attention generation module. The calculation module calculates the attention vectors for the N feature channels based on the pooled data through the attention generation module.

[0018] Optionally, the attention generation module includes a Conv-ReLU layer, a global average pooling layer, and a Conv-SigMoid function layer; The compression module includes: The Conv-ReLU layer is used to compress the N feature channels in the output feature map to generate compressed features for the N feature channels.

[0019] Optionally, the pooling module includes: The compressed features are subjected to global average pooling operations through the global average pooling layer to generate pooled data for the N feature channels.

[0020] Optionally, the computing module includes: The Conv-SigMoid function layer calculates the attention vectors for each of the N feature channels based on the pooling data.

[0021] Optionally, the pruning unit includes: Identify the set of feature channels whose attention values ​​are below a preset value; Pruning is performed on the convolution kernels corresponding to the feature channel set below the preset value in the feature extraction module.

[0022] Optionally, the pruning device further includes: The third input unit is used to re-input the input feature image into the convolutional neural network model for training.

[0023] A third aspect of this application provides an electronic device, comprising: Processor, memory, input / output units, and bus; The processor is connected to memory, input / output units, and a bus; The memory holds a program, which the processor calls to execute, such as the first aspect and any optional pruning method of the first aspect.

[0024] The fourth aspect of this application provides a computer-readable storage medium on which a program is stored, which, when executed on a computer, performs the first aspect and any optional pruning method of the first aspect.

[0025] As can be seen from the above technical solutions, the embodiments of this application have the following advantages: In this application, the input feature image and a convolutional neural network (CNN) model are first obtained. The CNN model includes a feature extraction module and an attention generation module. The attention generation module can be located inside the CNN model and used in conjunction with other working layers, or it can be located outside the CNN model and used independently. The input feature image is the image used in the training of the CNN model. The input feature image is input into the feature extraction module of the CNN model. The feature extraction module performs image feature processing on the original image using N convolutional kernels, fusing the N feature channels corresponding to the N convolutional kernels to generate an output feature map. The output feature map is input into the attention generation module of the CNN model, which then analyzes the output feature map and generates N attention vectors for each of the N feature channels. The attention generation module calculates the attention values ​​for each of the N feature channels based on the N attention vectors. Pruning is then performed on the N convolutional kernels in the feature extraction module based on the attention values ​​of the N feature channels. In this application, an attention generation module calculates the attention values ​​for N feature channels based on N attention vectors. Neurons corresponding to feature channels with attention values ​​below a preset threshold are then deleted, or neurons corresponding to the last few feature channels with the lowest attention values ​​are deleted, thus completing the pruning process. This pruning method considers the importance of the data generated by each image channel, using the attention values ​​of the feature channels output by the convolution kernel as the criterion. The entire convolution kernel is used as the pruning unit, removing the convolution kernels corresponding to feature channels with smaller attention values, achieving the goal of neural network model compression. Attention better reflects the importance of the corresponding convolution kernel. This attention-based pruning method better preserves useful convolution kernels while deleting useless ones, achieving model lightweighting while minimizing the impact on model performance and improving pruning efficiency. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is a schematic diagram of an embodiment of the pruning method for the convolutional neural network model of this application; Figure 2 This is a schematic diagram of another embodiment of the pruning method for the convolutional neural network model of this application; Figure 3 This is a schematic diagram of an embodiment of the pruning device for the convolutional neural network model of this application; Figure 4 This is a schematic diagram of another embodiment of the pruning device for the convolutional neural network model of this application; Figure 5 This is a schematic diagram of one embodiment of the electronic device of this application. Detailed Implementation

[0028] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0029] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0030] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0031] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrases "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0032] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0033] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0034] In existing technologies, convolutional neural network (CNN) models are often extremely complex, accompanied by high storage space and computational resource consumption, making it difficult to implement deep learning models on various hardware platforms. For example, the VGG-16 CNN has over 130 million parameters, occupies over 500 MB of storage space, and requires over 30 billion floating-point operations to complete a single image recognition task. Research shows that CNN models contain a large number of redundant neurons and weights, with only 5-10% of the weights participating in the main computation and influencing the final result. This provides a theoretical basis for neural network compression. If effective compression methods can be found, deep neural networks can be more widely deployed on lightweight devices such as mobile devices. On the server side, better performance can also be provided. Pruning, as an effective method for CNN model compression, is receiving increasing attention and application. CNN model pruning first filters out unimportant neurons and weights from a large network, then removes them from the network, while preserving as much performance as possible. Pruning techniques can be categorized into structured pruning and unstructured pruning based on their granularity. Structured pruning removes neurons (filters in convolutional neural networks) as the basic unit. Because it directly prunes neurons, structured pruning models can achieve significant inference acceleration and storage advantages under current hardware conditions. However, its drawback is the large granularity of pruning, which often significantly impacts the accuracy of the compressed model. Unstructured pruning removes individual weights as the basic unit. The resulting model experiences less accuracy loss, but it ultimately produces a sparse weight matrix, requiring robust support from underlying hardware and computing libraries to achieve inference acceleration and storage advantages. Existing pruning methods either prune based on the neuron weights of the convolutional neural network model or perform pruning randomly. This approach cannot accurately determine the importance of neurons, and the pruning operation cannot effectively remove useless neurons, leading to a significant performance degradation in convolutional neural networks. Currently, the mainstream neural network pruning methods fall into two categories: random pruning and neuron weight-based pruning.

[0035] Random pruning removes neurons from a convolutional neural network model without considering the importance of each neuron in the overall model. As the model size decreases, the performance of the convolutional neural network model may also decline significantly.

[0036] Pruning based on neuron weights is a pruning method that sorts the neurons in a convolutional neural network model according to their respective weights and removes neurons with lower weight rankings. Although this method utilizes neuron weight information, it ignores the neuron's input and output information, i.e., contextual information. As a computational function of input and output features, the role and importance of neurons are closely related to the context. Therefore, simply considering the weight of a neuron itself cannot well reflect the importance of the neuron.

[0037] In summary, existing pruning methods for convolutional neural network models reduce pruning efficiency because they fail to adequately consider the importance of neurons.

[0038] Based on this, this application discloses a pruning method, apparatus, device, and medium for convolutional neural network models, which is used to improve pruning efficiency.

[0039] The technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0040] The method described in this application can be applied to servers, devices, terminals, or other devices with logical processing capabilities; therefore, this application does not limit its application. For ease of description, the following description uses a terminal as the executing entity.

[0041] Please see Figure 1 This application provides an embodiment of a pruning method for a convolutional neural network model, comprising: 101. Obtain the input feature image and the convolutional neural network model, wherein the convolutional neural network model includes a feature extraction module and an attention generation module; The terminal acquires the input feature image and the convolutional neural network model. The convolutional neural network model includes a feature extraction module and an attention generation module. The attention generation module can be located inside the convolutional neural network model and used in connection with other working layers, or it can be located outside the convolutional neural network model and used independently.

[0042] 102. Input the input feature image into the feature extraction module of the convolutional neural network model; The terminal inputs the input feature image into the feature extraction module of the convolutional neural network model to perform feature extraction processing on the input feature image to generate feature channels.

[0043] 103. The original image is processed by N convolutional kernels in the feature extraction module, and the N feature channels corresponding to the N convolutional kernels are fused to generate an output feature map; After the terminal inputs the input feature image into the feature extraction module of the convolutional neural network model, the terminal performs image feature processing on the original image through N convolutional kernels in the feature extraction module, and performs channel fusion on the N feature channels corresponding to the N convolutional kernels to generate an output feature map.

[0044] 104. Input the output feature map into the attention generation module of the convolutional neural network model; The terminal inputs the output feature map into the attention generation module of the convolutional neural network model, so that the attention generation module can analyze the output feature map.

[0045] 105. Analyze the output feature map through the attention generation module, and generate N attention vectors for the N feature channels respectively; In this embodiment, the deep learning attention mechanism is a biomimetic of the human visual attention mechanism, essentially a resource allocation mechanism. The physiological principle is that human visual attention can receive high-resolution signals from a specific area of ​​an image and perceive its surrounding areas at low resolution, with the viewpoint changing over time. In other words, the human eye quickly scans the entire image to find the target area requiring attention, then allocates more attention to this area to acquire more detailed information and suppress other useless information, thereby improving the efficiency of representation.

[0046] In neural networks, the attention mechanism can be considered a resource allocation mechanism. It can be understood as redistributing resources that were originally evenly distributed according to the importance of the attention object. Important units get more resources, and unimportant or less important units get less resources. In the structural design of deep neural networks, the resources that attention needs to allocate are basically weights.

[0047] After the terminal inputs the output feature map into the attention generation module of the convolutional neural network model, the terminal analyzes the output feature map through the attention generation module and generates N attention vectors for the N feature channels respectively. The attention vector is a vector obtained by analyzing the feature channels of the image through the attention generation module and calculating the parameters of different feature channels.

[0048] 106. Calculate the attention values ​​for the N feature channels based on the N attention vectors respectively; The terminal calculates the attention value of each of the N feature channels based on the N attention vectors. The attention value represents the number of features contained in each feature channel.

[0049] The importance of a convolutional kernel can be represented by the importance of its corresponding output feature channels. There are various ways to measure the importance of output channels. Since all output channels are calculated from the same input feature map, this embodiment assumes that there is a correlation between the output feature channels. Therefore, using only the values ​​of each feature channel individually would miss this correlation. This embodiment uses the attention of each output channel as the pruning criterion. The attention value is a number between 0 and 1; a larger value indicates that the feature channel is more important in the entire feature image, and also represents the importance of its corresponding convolutional kernel in the entire convolutional group.

[0050] For example, an input 3-channel RGB image, after convolution, outputs a 5-channel feature map. The second column contains 5 convolutional kernels, corresponding to the 5 output feature channels in the third column. This demonstrates that for the same input, each convolutional kernel generates a feature channel, which together form the entire output feature map. Therefore, each output feature channel directly reflects the importance of its corresponding convolutional kernel, making it a crucial criterion for determining the importance of convolutional kernels during neural network pruning. After the input feature map undergoes convolution operations with a set of kernels, it outputs a feature image. This output feature image is then fed into an attention calculation module to obtain attention value vectors for each output feature channel. Based on the attention values, the convolutional kernels in the kernel group are pruned.

[0051] 107. Prune the N convolutional kernels in the feature extraction module according to the attention values ​​of the N feature channels.

[0052] After the terminal calculates the attention values ​​of the N feature channels based on the N attention vectors, the terminal performs pruning on the N convolutional kernels in the feature extraction module based on the attention values ​​of the N feature channels.

[0053] In this application, the input feature image and a convolutional neural network (CNN) model are first obtained. The CNN model includes a feature extraction module and an attention generation module. The attention generation module can be located inside the CNN model and used in conjunction with other working layers, or it can be located outside the CNN model and used independently. The input feature image is the image used in the training of the CNN model. The input feature image is input into the feature extraction module of the CNN model. The feature extraction module performs image feature processing on the original image using N convolutional kernels, fusing the N feature channels corresponding to the N convolutional kernels to generate an output feature map. The output feature map is input into the attention generation module of the CNN model, which then analyzes the output feature map and generates N attention vectors for each of the N feature channels. The attention generation module calculates the attention values ​​for each of the N feature channels based on the N attention vectors. Pruning is then performed on the N convolutional kernels in the feature extraction module based on the attention values ​​of the N feature channels. In this application, an attention generation module calculates the attention values ​​for N feature channels based on N attention vectors. Neurons corresponding to feature channels with attention values ​​below a preset threshold are then deleted, or neurons corresponding to the last few feature channels with the lowest attention values ​​are deleted, thus completing the pruning process. This pruning method considers the importance of the data generated by each image channel, using the attention values ​​of the feature channels output by the convolution kernel as the criterion. The entire convolution kernel is used as the pruning unit, removing the convolution kernels corresponding to feature channels with smaller attention values, achieving the goal of neural network model compression. Attention better reflects the importance of the corresponding convolution kernel. This attention-based pruning method better preserves useful convolution kernels while deleting useless ones, achieving model lightweighting while minimizing the impact on model performance and improving pruning efficiency.

[0054] Please see Figure 2 This application provides an embodiment of a pruning method for a convolutional neural network model, comprising: 201. Obtain the input feature image and the convolutional neural network model, wherein the convolutional neural network model includes a feature extraction module and an attention generation module; 202. Input the input feature image into the feature extraction module of the convolutional neural network model; 203. The original image is processed by N convolutional kernels in the feature extraction module, and the N feature channels corresponding to the N convolutional kernels are fused to generate an output feature map; 204. Input the output feature map into the attention generation module of the convolutional neural network model; Steps 201 to 204 in this embodiment are similar to steps 101 to 104 in the foregoing embodiment, and are not repeated herein.

[0055] 205. Perform feature channel compression on the N feature channels in the output feature map respectively through the Conv-ReLU layer to generate compressed features of the N feature channels; 206. Perform global average pooling on the compressed features respectively through the global average pooling layer to generate pooled data of the N feature channels; 207. Calculate attention vectors of the N feature channels respectively according to the pooled data through the Conv-SigMoid function layer; The terminal performs feature channel compression on the N feature channels in the output feature map respectively through the Conv-ReLU layer, and generates compressed features of the N feature channels. The terminal performs global average pooling on the compressed features respectively through the global average pooling layer, and generates pooled data of the N feature channels. The terminal calculates attention vectors of the N feature channels respectively according to the pooled data through the Conv-SigMoid function layer.

[0056] Specifically, the attention value of an output feature channel can be independently learned by a convolutional neural network model in a deep learning manner, and the feature channel of an output feature image (B C H W) is subjected to feature channel compression through 3×3 Conv+ReLU to generate (B C’ H W), where C’<C, then a global average pooling operation is performed to generate (B C’ 1 1), then an attention vector (B C 1 1) corresponding to the output feature channel is obtained through 1×1 convolution+SigMoid, where B is the data volume of this batch, that is, the number of images, C / C' is the number of output feature channels, H is the height of each feature channel, and W is the width of each feature channel.

[0057] 208. Calculate attention values of the N feature channels respectively according to the N attention vectors; Step 208 in this embodiment is similar to step 106 in the foregoing embodiment, and is not repeated herein.

[0058] 209. Determine a feature channel set whose attention values are lower than a preset value; 210. Prune the convolution kernels corresponding to the feature channel set below the preset value in the feature extraction module; The terminal determines the set of feature channels below the preset value in the attention value, and performs pruning on the convolution kernels corresponding to the set of feature channels below the preset value in the feature extraction module. This greatly utilizes the relationship between the feature image and the convolutional neural network model, filters the feature images output by the convolutional neural network by feature channels, reduces unnecessary neurons, and improves the working efficiency of the convolutional neural network model.

[0059] 211. The input feature image is re-inputted into the convolutional neural network model for training.

[0060] The terminal re-inputs the input feature image into the convolutional neural network model for training, indicating that the attention calculation module can be directly added to the convolutional neural network model without affecting the internal operations of the convolutional neural network model. This allows us to learn channel attention simultaneously during the training of the neural network, prune the convolutional kernels of the entire model all at once after the neural network training is completed, or prune in stages during the training process, increasing the flexibility of pruning.

[0061] In this application, an input feature image and a convolutional neural network (CNN) model are first obtained. The CNN model includes a feature extraction module and an attention generation module. The attention generation module can be located inside the CNN model and used in conjunction with other working layers, or it can be located outside the CNN model and used independently. The input feature image is the image used in the training of the CNN model. The input feature image is input into the feature extraction module of the CNN model. The original image is processed by N convolutional kernels in the feature extraction module, and the N feature channels corresponding to the N convolutional kernels are fused to generate an output feature map. The terminal inputs the output feature map into the attention generation module of the CNN model. The terminal compresses the N feature channels in the output feature map using the Conv-ReLU layer to generate compressed features for the N feature channels. The terminal performs global average pooling on the compressed features using the global average pooling layer to generate pooled data for the N feature channels. The terminal calculates the attention vectors for the N feature channels using the Conv-SigMoid function layer based on the pooled data. The terminal calculates attention values ​​for N feature channels based on N attention vectors using an attention generation module. The terminal identifies a set of feature channels with attention values ​​below a preset value and prunes the convolutional kernels corresponding to these channels in the feature extraction module. This process maximizes the relationship between the feature image and the convolutional neural network (CNN) model, filtering the feature images output by the CNN to reduce unnecessary neurons and improve the model's efficiency. Finally, the input feature image is re-input into the CNN model for training.

[0062] In this application, an attention generation module calculates the attention values ​​for N feature channels based on N attention vectors. Neurons corresponding to feature channels with attention values ​​below a preset threshold are then deleted, or neurons corresponding to the last few feature channels with the lowest attention values ​​are deleted, thus completing the pruning process. This pruning method considers the importance of the data generated by each image channel, using the attention values ​​of the feature channels output by the convolution kernel as the criterion. The entire convolution kernel is used as the pruning unit, removing the convolution kernels corresponding to feature channels with smaller attention values, achieving the goal of neural network model compression. Attention better reflects the importance of the corresponding convolution kernel. This attention-based pruning method better preserves useful convolution kernels while deleting useless ones, achieving model lightweighting while minimizing the impact on model performance and improving pruning efficiency.

[0063] Secondly, this embodiment makes great use of the relationship between feature images and convolutional neural network models, filtering the feature images output by the convolutional neural network through feature channels, reducing unnecessary neurons, and improving the working efficiency of the convolutional neural network model.

[0064] Secondly, the terminal re-inputs the input feature image into the convolutional neural network model for training, indicating that the attention calculation module can be directly added to the convolutional neural network model without affecting the internal operations of the convolutional neural network model. This allows us to learn channel attention simultaneously during the training of the neural network, prune the convolutional kernels of the entire model all at once after the neural network training is completed, or prune in stages during the training process, increasing the flexibility of pruning.

[0065] Please see Figure 3 This application provides an embodiment of a pruning device for a convolutional neural network model, comprising: The acquisition unit 301 is used to acquire the input feature image and the convolutional neural network model, wherein the convolutional neural network model includes a feature extraction module and an attention generation module; The first input unit 302 is used to input the input feature image into the feature extraction module of the convolutional neural network model; The first generation unit 303 is used to perform image feature processing on the original image through N convolutional kernels in the feature extraction module, and to perform channel fusion on the N feature channels corresponding to the N convolutional kernels to generate an output feature map; The second input unit 304 is used to input the output feature map into the attention generation module of the convolutional neural network model; The second generation unit 305 is used to analyze the output feature map through the attention generation module and generate N attention vectors for the N feature channels respectively; The calculation unit 306 is used to calculate the attention values ​​of the N feature channels according to the N attention vectors respectively; The pruning unit 307 is used to prune the N convolutional kernels in the feature extraction module according to the attention values ​​of the N feature channels.

[0066] Please see Figure 4 This application provides an embodiment of a pruning device for a convolutional neural network model, comprising: The acquisition unit 401 is used to acquire the input feature image and the convolutional neural network model, wherein the convolutional neural network model includes a feature extraction module and an attention generation module; The first input unit 402 is used to input the input feature image into the feature extraction module of the convolutional neural network model; The first generation unit 403 is used to perform image feature processing on the original image through N convolutional kernels in the feature extraction module, and to perform channel fusion on the N feature channels corresponding to the N convolutional kernels to generate an output feature map; The second input unit 404 is used to input the output feature map into the attention generation module of the convolutional neural network model; The second generation unit 405 is used to analyze the output feature map through the attention generation module and generate N attention vectors for the N feature channels respectively; Optionally, the second generation unit 405 includes: Compression module 4051 generates compressed features for the N feature channels by performing feature channel compression on the N feature channels in the output feature map through the attention generation module. Pooling module 4052 generates pooled data for the N feature channels by performing pooling operations on the compressed features through the attention generation module. The calculation module 4053 calculates the attention vectors of the N feature channels respectively based on the pooling data through the attention generation module.

[0067] Optionally, the attention generation module includes a Conv-ReLU layer, a global average pooling layer, and a Conv-SigMoid function layer; The compression module 4051 includes: The Conv-ReLU layer is used to compress the N feature channels in the output feature map to generate compressed features for the N feature channels.

[0068] Optionally, the pooling module 4052 includes: The compressed features are subjected to global average pooling operations through the global average pooling layer to generate pooled data for the N feature channels.

[0069] Optionally, the computing module 4053 includes: The Conv-SigMoid function layer calculates the attention vectors for each of the N feature channels based on the pooling data.

[0070] The calculation unit 406 is used to calculate the attention values ​​of the N feature channels according to the N attention vectors respectively; The pruning unit 407 is used to prune the N convolutional kernels in the feature extraction module according to the attention values ​​of the N feature channels. Optionally, the pruning unit 407 includes: Identify the set of feature channels whose attention values ​​are below a preset value; Pruning is performed on the convolution kernels corresponding to the feature channel set below the preset value in the feature extraction module.

[0071] The third input unit 408 re-inputs the input feature image into the convolutional neural network model for training.

[0072] Please see Figure 5 This application provides an electronic device, including: Processor 501, memory 502, input / output unit 503, and bus 504.

[0073] The processor 501 is connected to the memory 502, the input / output unit 503, and the bus 504.

[0074] The memory 502 stores a program, and the processor 501 calls the program to execute it, such as... Figure 1 , Figure 2 The pruning method in the text.

[0075] This application provides a computer-readable storage medium on which a program is stored, and when the program is executed on a computer, it performs the following... Figure 1 , Figure 2 The pruning method in the text.

[0076] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0077] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0078] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0079] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0080] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A pruning method for a convolutional neural network model, characterized in that, include: The input feature image and convolutional neural network model are obtained, wherein the convolutional neural network model includes a feature extraction module and an attention generation module; The input feature image is input into the feature extraction module of the convolutional neural network model; The output feature image is processed by N convolutional kernels in the feature extraction module, and the N feature channels corresponding to the N convolutional kernels are fused to generate an output feature map. The output feature map is input into the attention generation module of the convolutional neural network model; The attention generation module analyzes the output feature map and generates N attention vectors for the N feature channels respectively. Calculate the attention values ​​for the N feature channels based on the N attention vectors respectively; The N convolutional kernels in the feature extraction module are pruned based on the attention values ​​of the N feature channels. The step of analyzing the output feature map through the attention generation module to generate N attention vectors for the N feature channels includes: compressing the N feature channels in the output feature map using the attention generation module to generate compressed features for the N feature channels; performing pooling operations on the compressed features using the attention generation module to generate pooled data for the N feature channels; and calculating the attention vectors for the N feature channels based on the pooled data using the attention generation module. The attention generation module includes a Conv-ReLU layer, a global average pooling layer, and a Conv-SigMoid function layer; the step of compressing the N feature channels in the output feature map through the attention generation module to generate compressed features for the N feature channels includes: compressing the N feature channels in the output feature map through the Conv-ReLU layer to generate compressed features for the N feature channels.

2. The pruning method according to claim 1, characterized in that, The step of performing pooling operations on the compressed features through the attention generation module to generate pooled data for the N feature channels includes: The compressed features are subjected to global average pooling operations through the global average pooling layer to generate pooled data for the N feature channels.

3. The pruning method according to claim 1, characterized in that, The step of calculating the attention vectors for the N feature channels respectively by the attention generation module based on the pooled data includes: The Conv-SigMoid function layer calculates the attention vectors for each of the N feature channels based on the pooling data.

4. The pruning method according to any one of claims 1 to 3, characterized in that, The step of pruning the N convolutional kernels in the feature extraction module based on the attention values ​​of the N feature channels includes: Identify the set of feature channels whose attention values ​​are below a preset value; Pruning is performed on the convolution kernels corresponding to the feature channel set below the preset value in the feature extraction module.

5. The pruning method according to any one of claims 1 to 3, characterized in that, After pruning the N convolutional kernels in the feature extraction module based on the attention values ​​of the N feature channels, the pruning method further includes: The input feature image is then re-input into the convolutional neural network model for training.

6. A pruning device for a convolutional neural network model, characterized in that, include: An acquisition unit is used to acquire an input feature image and a convolutional neural network model, wherein the convolutional neural network model includes a feature extraction module and an attention generation module; The first input unit is used to input the input feature image into the feature extraction module of the convolutional neural network model; The first generation unit is used to perform image feature processing on the output feature image through N convolutional kernels in the feature extraction module, and to perform channel fusion on the N feature channels corresponding to the N convolutional kernels to generate an output feature map; The second input unit is used to input the output feature map into the attention generation module of the convolutional neural network model; The second generation unit is used to analyze the output feature map through the attention generation module and generate N attention vectors for the N feature channels respectively; A calculation unit is used to calculate the attention values ​​of the N feature channels based on the N attention vectors respectively; The pruning unit is used to prune the N convolutional kernels in the feature extraction module based on the attention values ​​of the N feature channels. The second generation unit includes: a compression module, which compresses the N feature channels in the output feature map using the attention generation module to generate compressed features for the N feature channels; a pooling module, which performs pooling operations on the compressed features using the attention generation module to generate pooled data for the N feature channels; and a calculation module, which calculates the attention vectors for the N feature channels based on the pooled data using the attention generation module. The attention generation module includes a Conv-ReLU layer, a global average pooling layer, and a Conv-SigMoid function layer; the compression module includes: performing feature channel compression on the N feature channels of the output feature map through the Conv-ReLU layer to generate compressed features for the N feature channels.

7. An electronic device, characterized in that, include: Processor, memory, input / output units, and bus; The processor is connected to the memory, the input / output unit, and the bus; The memory stores a program, which the processor calls to execute the pruning method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains a program that, when executed on a computer, performs the pruning method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Small sample attention mechanism parallel twinning method for eye fundus image classification

    CN114494195A