A Method for Improving the Performance of Convolutional Neural Networks Based on Weighted Attention

By adopting the weight attention mechanism in the convolutional neural network and using normalization and scale factors to generate weight attention, the problem of the existing attention mechanism increasing parameters and slow speed is solved, and the performance improvement is achieved with lightweight performance. It is suitable for a variety of CNN architectures, especially in image classification tasks.

CN115511051BActive Publication Date: 2025-07-22EAST CHINA NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211179190.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-27
Publication Date
2025-07-22
Estimated Expiration
2042-09-27

AI Technical Summary

Technical Problem

The attention mechanism embedding of existing convolutional neural networks will increase the amount of network parameters and slow down the training speed, resulting in insufficient performance improvement.

Method used

Normalized statistical spatial information volume and scale factors are used to generate weight attention, weight factors are generated through Lp-Norm, and the convolutional layer feature map is corrected to pay attention to important spatial channel areas, and 3D attention weights are generated through sigmoid function to guide feature learning, and embedded in different CNN architectures.

Benefits of technology

With almost no increase in network parameters and computing costs, the model training speed and performance are improved, especially in image classification tasks, and have wide application prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115511051B_ABST
    Figure CN115511051B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for improving the performance of a convolutional neural network based on weighted attention. The method is characterized by normalizing the statistical spatial information and using a scale factor to generate weighted attention, embedding the weighted attention module into convolutional neural networks with different CNN architectures to enhance the feature learning ability of the network model, and realizing the optimization of network performance. Specifically, it includes steps such as correcting the weighted attention for each intermediate feature map and extracting important features through a residual convolutional block that repeatedly embeds the weighted attention. Compared with the prior art, the present invention can embed into convolutional neural network models of any architecture to improve the performance of the model in tasks such as image classification. It is plug-and-play and can obtain a large performance gain with almost no increase in network overhead. It is a lightweight and efficient attention mechanism for optimizing neural networks and has broad application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and deep learning, and specifically to a method for improving the performance of a convolutional neural network based on weight attention. Background Art

[0002] In recent years, neural networks and deep learning have been widely applied in the fields of computer vision and image classification. Among them, the convolutional neural network (CNN) has been widely used due to its excellent feature expression ability. Since the convolutional neural network often needs to stack a large number of convolutional layers to extract features, the neural network model will extract more redundant features, making the model unable to distinguish important and unimportant features well. Researchers have borrowed the attention mechanism of human vision, investing more resources in the focused area to obtain detailed information of the target to be concerned, thereby ignoring or suppressing other useless information. In order to improve the network model to concentrate computing resources on extracting important features, various attention mechanisms such as SENet and ECANet have been successively proposed to capture different features, and corresponding importance weights are obtained for channel calibration according to the channel dependence. However, the global average pooling among them loses spatial information, and attention mechanisms such as CBAM consider the pixel dependence in the spatial dimension while strengthening the channel correlation.

[0003] After the attention mechanism of the prior art is embedded in the model, there are situations where the number of network parameters increases and the training speed slows down, resulting in the performance of the network model not being fully improved. Summary of the Invention

[0004] The object of the present invention is to design a method for improving the performance of a convolutional neural network based on weighted attention in view of the deficiencies of the prior art. By using the method of normalizing the statistical spatial information and generating weighted attention with a scale factor, the feature map extracted by the convolutional layer is corrected to focus on important spatial channel regions. This method uses the scale factor in the spatial dimension normalization of the feature map to measure the information volume of the feature map, and after passing through the Lp-Norm as the weight factor of channel importance, the weight factor is assigned to the feature map for recalibration to improve the feature expression ability. After being constrained by the sigmoid function, 3D attention weights are generated to guide the feature map to concentrate attention on specific spatial and channel regions. Without increasing the computational cost of network model parameters almost, the model training speed is accelerated, and it can be embedded as a gain branch module into convolutional neural networks of various different CNN architectures to enhance the feature learning ability of the network model and realize the optimization of network performance. It preferably solves the problems such as the increase of parameter burden and the slowdown of speed when the existing attention mechanism is embedded in the network, and realizes efficient tasks such as image classification and object detection. The weighted attention module is plug-and-play, which can improve the training convergence speed of the neural network and the performance in the field of image classification and other fields without increasing the network parameter burden and computational cost almost, and has broad application prospects.

[0005] The specific technical solution for achieving the object of the present invention is: a method for improving the performance of a convolutional neural network based on weighted attention, which is characterized in that it uses the method of normalizing the statistical spatial information and generating weighted attention with a scale factor to correct the feature map extracted by the convolutional layer to focus on important spatial channel regions, and without increasing the computational cost of network model parameters almost, higher performance of the model is achieved, such as tasks like image classification. The specific optimization process includes the following steps:

[0006] Step a: Make an image training set;

[0007] Step b: Input the image into the convolutional neural network, perform shallow feature extraction through the convolutional layer and the pooling layer, and then perform deep feature extraction through the convolutional block with residual connection to generate an intermediate feature map X;

[0008] Step c: Group the intermediate feature maps X in the convolutional layer of the residual block along the channel dimension;

[0009] Step d: Statistically calculate the spatial information volume of each feature map X1 in each group, and generate a weight factor ω for channel correction to obtain a feature map X2;

[0010] Step e: Use learnable parameters α and β to scale and translate the feature map X2 after channel correction to enhance the feature expression, and obtain a feature map X3;

[0011] Step f: Generate 3D attention weights after constraining the feature map X3 through the sigmoid function, and multiply the attention weights with the feature map X3 pixel by pixel and channel by channel to obtain the attention-corrected feature map Y1;

[0012] Step g: Repeat the convolutional layer with embedded weight attention for feature extraction to guide the network to learn representative important features and suppress unimportant features. Finally, classify and predict the image to be classified with unknown categories according to the features extracted through the fully connected layer to obtain the classification probability.

[0013] In step d, normalizing each feature map X1 in the feature group to obtain the importance weight factor ω specifically includes the following steps:

[0014] Step d1: Normalize each feature map X1 in the feature group in the spatial dimension, and use the scale factor in its affine transformation process to measure the change amount of spatial information of different feature maps;

[0015] Step d2: Perform Lp-Norm on the scale factor to obtain the importance weight factor ω of different feature maps;

[0016] Step d3: Assign the weight factor ω to the corresponding feature map channel by channel to obtain the feature map X2.

[0017] Compared with the prior art, the present invention has the characteristics of lightweight, high efficiency, plug-and-play. Without increasing the parameter burden of the network almost, it accelerates the model training speed, can be embedded as a gain branch module into a variety of different CNN architectures to optimize the network performance, not only reduces the model training time and additional attention calculation cost, but also improves the performance of the convolutional neural network model in tasks such as image classification, and can achieve faster and better effects than the existing attention in different network model architectures. It is a lightweight and efficient attention mechanism to optimize the neural network, further promoting the lightweight exploration scheme of the neural network attention mechanism, which is an important direction for future research and has broad application prospects. The weight attention mechanism is lightweight and efficient, and obtains greater performance gains without increasing the network overhead almost. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is the flow chart of the present invention;

[0019] Figure 2 It is the example diagram of the weight attention module;

[0020] Figure 3 It is the example diagram of embedding the weight attention module into different positions of the CNN architecture. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The present invention takes a deep convolutional neural network model as the target model, normalizes to statistically calculate the spatial information of the feature map, and uses the scale factor therein to generate a channel weight factor, which efficiently guides the extraction of important features in a lightweight manner, thereby improving the performance of the model in tasks such as image classification. The present invention simultaneously considers the attention of the feature map in the channel and spatial dimensions in the network model attention optimization process, and generates a 3D attention weight correction network. The optimization process can improve the performance of the network model with almost no parameters and computational cost.

[0022] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with the following specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0023] Embodiment 1

[0024] Refer to Figure 1 , the performance of the convolutional neural network is optimized by the present invention according to the following steps:

[0025] Step a: Obtaining the intermediate feature map X

[0026] Input the image into the convolutional neural network, perform shallow feature extraction through the convolutional pooling layer, and perform deep feature extraction through the residual convolutional block to obtain the intermediate feature map X. The residual convolutional block includes: a convolutional layer and a weight attention module.

[0027] Step b: Correcting the weight attention of each intermediate feature map X

[0028] Refer to Figure 2 , the weight attention module mainly performs in two steps: grouping and attention generation. First, each intermediate feature map X is grouped along the channel dimension, and each feature map X1 in the feature group is normalized in the spatial dimension to statistically calculate the spatial information of the feature map and extract the scale factor. The Lp-Norm is implemented on the scale factor to generate the weight factor ω of the importance of different channel features and assign them one by one. Subsequently, the feature map X2 after channel correction is recalibrated with the learnable parameters α and β to obtain the feature map X3. Subsequently, the feature map X3 is constrained by the sigmoid activation function to generate 3D weight attention, and the 3D attention weight is multiplied pixel by pixel and channel by channel with each feature map X1 in the feature group to guide the important feature regions of feature learning.

[0029] Step c: Repeating the residual convolutional block with embedded weight attention to extract important features

[0030] Refer to Figure 3, for the position of the embedded weight attention module in the residual convolution blocks of different CNN architectures (MobileNet series, MoblieNeXt), extract important features and suppress unimportant features.

[0031] Step d: Update the model parameters of the deep convolutional neural network using the gradient descent method on the training dataset

[0032] Train for 200 epochs on the CIFAR-10 and CIFAR-100 datasets respectively. Set the batch size to 128, the initial learning rate to 0.05, use the label smoothing loss function and the SGD optimizer, and set the weight decay coefficient to 4e-5.

[0033] Step e: After the model training is completed, fix the model parameters and verify them on the image classification test set to obtain the following comparison table of the classification performance improvement of the CNN architecture with the embedded weight attention module (WA) and other attention models on the Cifar dataset as shown in Table 1 below:

[0034] Table 1 Comparison table of the classification performance improvement of the embedded weight attention module and other attention models on the Cifar dataset

[0035]

[0036] By comparing the classification performance of the MobileNetV2 and MoblieNeXt lightweight network architectures with the embedded weight attention module and other attention modules in Table 1 above on the CIFAR-10 and CIFAR100 datasets, after the weight attention module (WA) of the present invention is added to the lightweight baseline networks MobileNetV2 and MoblieNeXt respectively, the amount of computation and the number of parameters remain the same as the original network, and the classification accuracy is improved without increasing the computational cost. For example, the MoblieNeXt network with the embedded WA achieved a classification accuracy of 60.88%, which is 6.11% higher than the original network (54.77%). Compared with other attention methods (such as SE, CBAM, etc.), the present invention hardly increases the additional number of parameters and computational overhead, and obtains faster and better image classification results than the existing attention methods.

[0037] The above embodiments are only for further illustrating the present invention, and are not intended to limit the patent of the present invention. All equivalent implementations of the present invention should be included within the scope of the claims of the present invention.

Claims

1. A method for improving the performance of a convolutional neural network based on weighted attention, characterized in that A method of adopting normalized statistical spatial information and using a scale factor to generate weight attention is proposed. The weight attention module is embedded into convolutional neural networks with different CNN architectures to enhance the feature learning ability of the network model and optimize the network performance. The specific steps are as follows: Step a: Input an image into the convolutional neural network. Perform shallow feature extraction and downsampling through convolutional layers and pooling layers, and then repeatedly stack convolutional blocks with residual connections for deep feature extraction to obtain an intermediate feature map X with C channels. Step b: Group the intermediate feature map X along the channel dimension to obtain G groups of feature maps X1, each containing C / G channels. Step c: Normalize each feature map X1 in the spatial dimension, and use the scale factor in its affine transformation to evaluate the spatial information of different feature maps to measure the importance of different feature maps. Step d: Apply Lp-Norm to the scale factor to obtain the importance weight factor ω of different feature maps. Step e: Assign the importance weight factor ω to each feature map X1 channel by channel to obtain the feature map X2. Step f: Recalibrate each feature map X2 using learnable parameters α and β to obtain the feature map X3, enhancing the feature expression ability. The learnable parameters α and β are trained along with the network by the following formula (a): X3 = αX2 + β (a); Step g: Constrain the feature map X3 through the sigmoid activation function to generate 3D attention weights, and multiply each pixel channel by the feature map X3 to obtain one group of feature maps Y1. Finally, aggregate the G groups of feature maps to obtain the output feature map Y. The feature map Y1 is calculated by the following formula (b): Y1 = X1 * Sigmoid(X3) (b); Step h: Embed the weight attention mechanism in steps b to g into the residual convolutional layer to learn representative important features and suppress unimportant features. Finally, obtain the classification probability result through the fully connected layer. The whole process is optimized in the training set, and then the classification prediction of the unknown-class test images is performed on the validation set.

2. The method for improving the performance of a convolutional neural network based on weighted attention according to claim 1, wherein The optimization method can be embedded into different convolutional layers of the convolutional neural network as a plug-and-play branch gain module.

3. The method for improving the performance of a convolutional neural network based on weighted attention according to claim 1, wherein Embedding the weight attention mechanism into the residual convolutional layer can correct the feature map to extract important features and suppress unimportant features.

4. The method for improving the performance of a convolutional neural network based on weighted attention according to claim 1, wherein The weight attention mechanism uses the scale factor in the group normalization process to measure the feature map information weight for attention guidance.

Citation Information

Patent Citations

  • Attention weight module and method for convolutional neural network

    CN112801262A

  • Lightweight image super-resolution method based on serial high-frequency attention

    CN114897690A