Semantic Segmentation Method Based on the Fusion of Channel Attention and Pyramid Convolution
By introducing the fusion technology of channel attention and pyramid convolution in the semantic segmentation method, the problem of low accuracy in existing methods when dealing with small objects is solved, and higher segmentation accuracy is achieved.
Patent Information
- Application Number
- CN202111361747.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-17
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-11-17
AI Technical Summary
When existing semantic segmentation methods deal with small or incomplete objects, it is difficult to extract significant features, resulting in low segmentation accuracy.
The semantic segmentation method based on the fusion of channel attention and pyramid convolution is adopted. By adding a pyramid convolution module to the ResNet50 network, local features and global features are extracted, and then input into the channel attention module to enhance the discrimination ability of the feature map.
Effectively enhance the characterization ability of feature maps for specific semantics and improve the accuracy of segmentation, especially when dealing with small or incomplete objects.
Smart Images

Figure CN114155371B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing and semantic segmentation methods, and relates to a semantic segmentation method based on the fusion of channel attention and pyramid convolution. Background Art
[0002] In recent years, computer vision and machine learning technologies have attracted much attention, and at the same time, people are becoming more and more interested in the problem of image semantic segmentation. More and more application scenarios require accurate and efficient segmentation technologies, such as autonomous driving, indoor navigation, virtual reality, and augmented reality.
[0003] Semantic segmentation is a task of predicting the category of individual pixels in an image and has long been one of the key problems in computer vision. Semantic segmentation divides an image into multiple regions according to the different attributes of pixels and extracts meaningful information for analysis.
[0004] With the in-depth study of semantic segmentation, some classic semantic segmentation models have emerged. The fully convolutional neural network structure (Long J, Shelhamer E, Darrell T. Fully Convolutional Networks for Semantic Segmentation[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2015, 39(4): 640-651.) mainly consists of two parts: the fully convolutional part and the deconvolution part. The fully convolutional part borrows some classic CNN networks and replaces the last fully connected layer with a convolution for feature extraction; the deconvolution part upsamples the small-size feature map to obtain the original-size semantic segmentation image. The U-Net network structure (Ronneberger O, Fischer P, Brox T. U-Net: Convolutional Networks for Biomedical Image Segmentation[J]. Springer International Publishing, 2015.) mainly consists of three parts: downsampling, upsampling, and skip connections. The image size is reduced through convolution and downsampling to extract shallow features; deep features are obtained through convolution and upsampling; and the shallow features and deep features are fused through skip connections to refine the image. However, they do not consider global context information and only extract some local features, resulting in limited segmentation performance.
[0005] The PSPNet network structure (Zhao H, Shi J, Qi X, et al. Pyramid Scene Parsing Network [J]. IEEE Computer Society, 2016.) introduces dilated convolutions to extract features and also introduces a pyramid pooling module to aggregate context information based on different regions to improve the ability to obtain global context information. The DeeplabV3+ (Chen LC, Zhu Y, Papandreou G, et al. Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation [J]. Springer, Cham, 2018.) model introduces a spatial pyramid pooling module with dilated convolutions to fuse multi-scale information. At the same time, a decoder module is introduced to further fuse low-level features and high-level features. However, when faced with small-sized objects, it is unable to extract significant features. During the segmentation process, there will be some small-sized or incomplete objects. If only simple context information fusion is used, the small or incomplete objects will be ignored. Therefore, if different-scale features are treated equally to represent different semantics, the segmentation results will not be accurate enough. Summary of the Invention
[0006] The purpose of the present invention is to provide a semantic segmentation method based on the fusion of channel attention and pyramid convolution to solve the problem of low accuracy of existing segmentation methods.
[0007] The technical solution adopted by the present invention is a semantic segmentation method based on the fusion of channel attention and pyramid convolution, which is specifically implemented according to the following steps:
[0008] Step 1: Input the training images in the database into the ResNet50 network to extract the features of the images.
[0009] Step 2: Add a pyramid convolution module to the last layer of the ResNet50 network in Step 1 to capture local features and global features respectively.
[0010] Step 3: Fuse the local features and global features obtained in Step 2 to obtain fused feature information.
[0011] Step 4: Input the fused feature information obtained in Step 3 into the channel attention module to obtain an enhanced feature map.
[0012] Step 5: Fuse the fused features obtained in Step 3 and the enhanced feature map obtained in Step 4.
[0013] Step 6: Upsample the features after fusion in Step 5 to obtain the segmentation image.
[0014] The feature of the present invention also lies in that
[0015] The calculation expression for extracting the features of the image in Step 1 is:
[0016] F = f(W c *X) (1)
[0017] In formula (1): X represents the training image in the database, and W C represents the overall parameters in the ResNet50 network, and f(·) represents extracting features from the image.
[0018] The specific process of Step 2 is:
[0019] Step 2.1: Add a pyramid convolution local feature extraction module to the last layer of the ResNet50 network to capture local features;
[0020] Step 2.1.1: Reduce the dimension of the features of the image extracted in Step 1 to 512 dimensions through a 1*1 convolution;
[0021] Step 2.1.2: Divide the features with reduced dimensions in Step 2.1.1 into different groups and perform convolutions respectively according to the sizes of convolution kernels of 9*9, 7*7, 5*5, and 3*3;
[0022] Step 2.1.3: Perform convolution on the features processed by convolution in Step 2.1.2 with a convolution kernel of 1*1 size to obtain local features;
[0023] Step 2.2: Add a global feature extraction module of pyramid convolution to the last layer of the ResNet50 network to capture global features;
[0024] Step 2.2.1: Use adaptive average pooling to reduce the size of the features of the image extracted in Step 1 to 9*9;
[0025] Step 2.2.2: Reduce the feature map of the features reduced in Step 2.2.1 to 512 dimensions through a 1*1 convolution;
[0026] Step 2.2.3: Divide the features with reduced dimensions in Step 2.2.2 into different groups and perform convolutions respectively according to the sizes of convolution kernels of 9*9, 7*7, 5*5, and 3*3;
[0027] Step 2.2.4: Perform convolution on the features processed by convolution in Step 2.2.3 with a convolution kernel of 1*1 size to obtain global features.
[0028] In Step 2.1.2 and Step 2.2.3, the number of feature groups corresponding to the 9*9 convolution kernel is 16, the number of feature groups corresponding to the 7*7 convolution kernel is 8, the number of feature groups corresponding to the 5*5 convolution kernel is 4, and the number of feature groups corresponding to the 3*3 convolution kernel is 1.
[0029] The expression of the fused feature information in Step 3 is:
[0030]
[0031] In Equation (4): f 1 is the obtained local feature, f 2 is the obtained global feature, and F1 is the fused feature information.
[0032] The specific process of Step 4 is as follows:
[0033] In Step 4.1, the fused feature information obtained in Step 3 is input into the channel attention module to obtain the channel attention map, that is, the relative factor affecting each channel. The expression is:
[0034]
[0035] In Equation (5), x ji represents the influence of the i-th channel on the j-th channel, A i represents the feature map of the i-th channel, and A j represents the feature map of the j-th channel;
[0036] In Step 4.2, the channel attention map obtained in Step 4.1 and the features of the image extracted in Step 1 are used to calculate the enhanced feature map;
[0037]
[0038] In Equation (6), x ji represents the influence of the i-th channel on the j-th channel, A i represents the feature map of the i-th channel, A j represents the feature map of the j-th channel, and β is the weight factor, initialized to 0.
[0039] The fusion method in Step 5 is:
[0040]
[0041] In Equation (7), F 1 is the fused feature information in Step 3, and E is the enhanced feature map in Step 4.
[0042] The specific process of step 6 is as follows: The features fused in step 5 are subjected to deconvolution operation to add empty pixels between every two pixels, so that the size of the processed feature map is the same as that of the training image, and the image segmentation result is obtained.
[0043] The beneficial effect of the present invention is that the semantic segmentation method based on the fusion of channel attention and pyramid convolution of the present invention uses the pyramid convolution module to extract local features and global features, and fuses the local features and global features. By introducing the channel attention mechanism and obtaining the mutual dependence between different channel mappings, the representation ability of the feature map for specific semantics is effectively enhanced, and finally the discrimination ability of the feature map is enhanced, and the accuracy of segmentation is improved. Description of the Drawings
[0044] Figure 1 is a flowchart of the semantic segmentation method based on the fusion of channel attention and pyramid convolution of the present invention. Detailed Embodiments
[0045] The present invention will be described in detail below with reference to the drawings and specific embodiments.
[0046] The present invention provides a semantic segmentation method based on the fusion of channel attention and pyramid convolution, which is specifically implemented according to the following steps:
[0047] Step 1, input the training images in the database into the ResNet50 network to extract the features of the images;
[0048] The ResNet50 network structure includes 5 stages. The first stage: the training image passes through a convolutional layer with a stride of 2 and a convolutional kernel size of 7 and a max-pooling process with a stride of 2 and a size of 3*3; the second stage contains 3 Bottlenecks; the third stage contains 4 Bottlenecks; the fourth stage contains 6 Bottlenecks; the fifth stage contains 3 Bottlenecks; each Bottleneck is composed of convolutional layers of 1*1, 3*3, and 1*1; the first stage is the preprocessing of the training image, and the remaining 4 stages are for feature extraction;
[0049] When the ResNet50 network extracts features, when the size of the feature map is reduced by half, the number of feature maps will double, maintaining the complexity of the network; but when the depth of the model reaches a certain level, a degradation problem will occur. The ResNet50 network adds an identity mapping. After one convolution, if the effect becomes worse, the weight parameters are kept unchanged, thus preventing the model degradation problem;
[0050] After the ResNet50 network extracts the image features, the finally extracted feature size is 7*7*2048;
[0051] Among them, the calculation expression for extracting the features of the image is:
[0052] F = f(W c *X) (1)
[0053] In formula (1): X represents the training image in the database, and W C represents the overall parameters in the ResNet50 network, including weights and biases, and f(·) represents extracting features from the image;
[0054] Step 2: Add a pyramid convolution module to the last layer of the ResNet50 network in Step 1 to capture local features and global features respectively;
[0055] Step 2.1: Add a pyramid convolution local feature extraction module to the last layer of the ResNet50 network to capture local features;
[0056] The pyramid convolution local feature extraction module is mainly divided into three parts: feature dimensionality reduction, local detail acquisition, and feature combination. Feature dimensionality reduction is composed of 1*1 convolution kernels; local detail acquisition is composed of convolution kernels of different sizes of 9*9, 7*7, 5*5, and 3*3. At the same time, in order to use kernels of different depths at each level of the pyramid convolution, the input feature map is divided into different groups for grouped convolution, and kernels are applied independently to each group of input feature maps; feature combination combines the information extracted under different kernel sizes and depths by 1*1 convolution kernels;
[0057] The pyramid convolution local feature extraction module is mainly responsible for smaller objects and capturing local fine details at multiple scales;
[0058] The calculation method of local feature extraction is as follows,
[0059] f 1 = g 1 (W 1 *F) (2)
[0060] In formula (2): f 1 is the extracted local feature, F is the input feature map, and W 1 represents the overall parameters of the pyramid convolution local feature extraction module, and g 1 (·) is the pyramid convolution local feature extraction module;
[0061] Step 2.1.1: Reduce the dimension of the features of the image extracted in Step 1 to 512 dimensions through 1*1 convolution;
[0062] Step 2.1.2: Divide the dimensionality-reduced features in Step 2.1.1 into different groups (divide into different groups according to the number of channels), and perform convolutions respectively with the sizes of convolutional kernels being 9×9, 7×7, 5×5, and 3×3; among them, the number of feature groups corresponding to the convolutional kernel of 9×9 is 16, the number of feature groups corresponding to the convolutional kernel of 7×7 is 8, the number of feature groups corresponding to the convolutional kernel of 5×5 is 4, and the number of feature groups corresponding to the convolutional kernel of 3×3 is 1;
[0063] Step 2.1.3: Perform convolution on the features processed by convolution in Step 2.1.2 with the size of the convolutional kernel being 1×1 to obtain local features;
[0064] Step 2.2: Add a pyramid convolution global feature extraction module to the last layer of the ResNet50 network to capture global features;
[0065] The pyramid convolution global feature extraction module is responsible for capturing the global features of the scene and processing larger objects. It is a multi-scale global aggregation module, mainly composed of adaptive average pooling, feature dimensionality reduction, global feature acquisition, and feature combination; adaptive average pooling reduces the spatial size of the feature map to a fixed size to ensure capturing complete global information; feature dimensionality reduction consists of a 1×1 convolutional kernel to reduce the features to a reasonable dimension; global feature acquisition consists of convolutional kernels with different sizes of 9×9, 7×7, 5×5, and 3×3. At the same time, in order to use kernels with different depths at each level of the pyramid convolution, the input feature map is divided into different groups, and grouped convolution is performed to independently apply the kernel to each input feature map group; feature combination consists of a 1×1 convolutional kernel to combine the information extracted under different kernel sizes and depths;
[0066] The calculation method of global feature extraction is as follows,
[0067] f 2 =g 2 (W 2 *F) (3)
[0068] In formula (3): f 2 is the extracted global feature, F represents the input feature map, W 2 represents the overall parameters of the pyramid convolution global feature extraction module, and g 2 (·) is the pyramid convolution global feature extraction module;
[0069] Step 2.2.1: Use adaptive average pooling to reduce the size of the features extracted from the image in Step 1 to 9×9;
[0070] Step 2.2.2: Pass the features reduced in Step 2.2.1 through a 1×1 convolution to reduce the feature map to 512 dimensions;
[0071] Step 2.2.3: Divide the dimensionality-reduced features obtained in Step 2.2.2 into different groups and perform convolutions respectively according to the sizes of convolutional kernels of 9×9, 7×7, 5×5, and 3×3; among them, the number of feature groups corresponding to the convolutional kernel of 9×9 is 16, the number of feature groups corresponding to the convolutional kernel of 7×7 is 8, the number of feature groups corresponding to the convolutional kernel of 5×5 is 4, and the number of feature groups corresponding to the convolutional kernel of 3×3 is 1.
[0072] Step 2.2.4: Perform convolution on the features processed by convolution in Step 2.2.3 according to the size of the convolutional kernel of 1×1 to obtain global features.
[0073] Step 3: Fuse the local features and global features obtained in Step 2 to obtain fused feature information, thereby obtaining multi-scale features from coarse to fine and obtaining relatively rich feature information.
[0074] Among them, the expression of the fused feature information is:
[0075]
[0076] In formula (4): f 1 is the obtained local feature, f 2 is the obtained global feature, and F 1 is the fused feature information.
[0077] Step 4: Input the fused feature information obtained in Step 3 into the channel attention module to obtain an enhanced feature map.
[0078] The channel attention module is used to explore the similarity relationship between each channel in the image feature map, so that each channel has global semantic features; each channel mapping of the high-level features can be regarded as a response with a clear category, and different semantic responses are related to each other; by obtaining the mutual dependence between different channel mappings, the representational ability of the feature map for specific semantics can be effectively enhanced.
[0079] Step 4.1: Input the fused feature information obtained in Step 3 into the channel attention module to obtain a channel attention map, that is, the relative factor affecting each channel, and the expression is:
[0080]
[0081] In formula (5), x ji represents the influence of the i-th channel on the j-th channel, A i represents the feature map of the i-th channel, and A j represents the feature map of the j-th channel.
[0082] Step 4.2, calculate the enhanced feature map by using the channel attention map obtained in step 4.1 and the features of the image extracted in step 1;
[0083]
[0084] In formula (6), x ji represents the influence of the ith channel on the jth channel, A i Represents the feature map of the i-th channel, A j represents the feature map of the jth channel, β is the weight factor, initialized to 0;
[0085] Step 5, fusing the fused features obtained in step 3 with the enhanced feature map obtained in step 4;
[0086] In the segmentation process, we should not only pay attention to the multi-scale features of the image, but also learn the global semantic dependency between channel feature maps to enhance the discrimination ability of feature maps; we can obtain the multi-scale features of the image from coarse to fine and the long-distance context information through fusion. The fusion method is as follows:
[0087]
[0088] In formula (7), F 1 is the feature information fused in step 3, and E is the enhanced feature map in step 4;
[0089] Step 6, up-sample the features fused in step 5 to obtain a segmented image;
[0090] Semantic segmentation requires restoring the extracted features to the same size as the original image, upsampling the feature map obtained in step 5, and using a deconvolution operation on the features fused in step 5 to add empty pixels between every two pixels so that the size of the processed feature map is the same as the training image size, thus obtaining the image segmentation result.
[0091] The present invention is based on a semantic segmentation method that combines channel attention with pyramid convolution. The processing object is an image in a database. Pyramid convolution is added to a ResNet50 network. Global and local detail features of the image are extracted through pyramid convolution and fused to obtain multi-scale features. The fused features are then input into a channel attention module to mine the similarity relationship between each channel in an image feature map, so that each channel has a global semantic feature and the discrimination ability of the feature map is enhanced. Next, the multi-scale features are fused with the enhanced feature map to capture effective contextual information. Finally, the obtained feature map is upsampled to obtain a segmented image. The global dependency between channels is fully considered, the discrimination ability is enhanced, and the segmentation accuracy of the model is improved.
Claims
1. A semantic segmentation method based on the fusion of channel attention and pyramid convolution, characterized in that, it is specifically implemented according to the following steps: Step 1: Input the training images in the database into the ResNet50 network to extract the features of the images; Step 2: Add a pyramid convolution module to the last layer of the ResNet50 network in Step 1 to capture local features and global features respectively; The specific process of Step 2 is: Step 2.1: Add a pyramid convolution local feature extraction module to the last layer of the ResNet50 network to capture local features; Step 2.1.1: Reduce the dimension of the features of the images extracted in Step 1 to 512 dimensions through a 1*1 convolution; Step 2.1.2: Divide the features with reduced dimensions in Step 2.1.1 into different groups and perform convolutions respectively according to the sizes of convolution kernels of 9*9, 7*7, 5*5, and 3*3; Step 2.1.3: Perform a convolution with a convolution kernel of 1*1 on the features processed by convolution in Step 2.1.2 to obtain local features; Step 2.2: Add a global feature extraction module of pyramid convolution to the last layer of the ResNet50 network to capture global features; Step 2.2.1: Use adaptive average pooling to reduce the size of the features of the images extracted in Step 1 to 9*9; Step 2.2.2: Reduce the feature map to 512 dimensions through a 1*1 convolution on the features reduced in Step 2.2.1; Step 2.2.3: Divide the features with reduced dimensions in Step 2.2.2 into different groups and perform convolutions respectively according to the sizes of convolution kernels of 9*9, 7*7, 5*5, and 3*3; Step 2.2.4: Perform a convolution with a convolution kernel of 1*1 on the features processed by convolution in Step 2.2.3 to obtain global features; Step 3: Fuse the local features and global features obtained in Step 2 to obtain fused feature information; Step 4: Input the fused feature information obtained in Step 3 into the channel attention module to obtain an enhanced feature map; The specific process of Step 4 is: Step 4.1: Input the fused feature information obtained in Step 3 into the channel attention module to obtain a channel attention map, that is, the relative factor affecting each channel, and the expression is: (5) In formula (5), represents the influence of the i -th channel on the j -th channel, represents the feature map of the i -th channel, represents the feature map of the j -th channel; Step 4.2: Calculate the enhanced feature map through the channel attention map obtained in Step 4.1 and the features of the images extracted in Step 1; (6) In formula (6), represents the influence of the i -th channel on the j -th channel, represents the feature map of the i -th channel, represents the feature map of the j -th channel, is the weight factor, initialized to 0; Step 5: Fuse the fused features obtained in Step 3 and the enhanced feature map obtained in Step 4; The fusion method in Step 5 is: (7) In formula (7), is the feature information fused in step 3, is the enhanced feature map in step 4; Step 6: Upsample the features fused in Step 5 to obtain a segmented image.
2. The semantic segmentation method based on the fusion of channel attention and pyramid convolution according to claim 1, characterized in that, the calculation expression of the features of the images extracted in Step 1 is: (1) In formula (1): represents the training images in the database, represents the overall parameters in the ResNet50 network, represents extracting features from the images.
3. The semantic segmentation method based on the fusion of channel attention and pyramid convolution according to claim 2, characterized in that, In Step 2.1.2 and Step 2.2.3, the number of feature groups corresponding to the 9*9 convolution kernel is 16, the number of feature groups corresponding to the 7*7 convolution kernel is 8, the number of feature groups corresponding to the 5*5 convolution kernel is 4, and the number of feature groups corresponding to the 3*3 convolution kernel is 1.
4. The semantic segmentation method based on the fusion of channel attention and pyramid convolution according to claim 1, wherein, the expression of the fused feature information in Step 3 is: (4) In formula (4): is the obtained local feature, is the obtained global feature, is the fused feature information.
5. The semantic segmentation method based on the fusion of channel attention and pyramid convolution according to claim 1, wherein, the specific process of Step 6 is: performing a deconvolution operation on the features fused in Step 5 to add empty pixels between every two pixels, so that the size of the processed feature map is the same as the size of the training image, and obtaining the image segmentation result.