Image semantic segmentation method based on multi-scale attention mechanism
By adopting a multi-scale attention mechanism and self-attention mechanism in image semantic segmentation, the problem of difficulty in classifying targets at different scales in complex scenarios is solved, and more efficient semantic segmentation effect and better edge details are achieved.
Patent Information
- Application Number
- CN202211471248.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-23
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-11-23
AI Technical Summary
In complex scenarios, the existence of different scales of the same object leads to difficulty in classification, and small-scale targets are prone to missed, large-scale targets are prone to missed, and the edges of target segmentation are not clear.
The image semantic segmentation method based on the multi-scale attention mechanism is adopted, and the semantic segmentation network model is constructed, including downsampling module, multi-scale feature extraction module, feature fusion module and upsampling module. The hollow space pyramid structure, channel attention and spatial attention mechanism are used to extract advanced multi-scale semantic features, and the output features of different scales are fused through the self-attention mechanism.
It improves the accuracy and efficiency of image semantic segmentation, can better handle different scale targets in complex scenes, optimize segmentation effects, and preserve edge details and small-objective features.
Smart Images

Figure CN115775316B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision research, and in particular relates to an image semantic segmentation method based on a multi-scale attention mechanism applied in complex scenes. Background Art
[0002] Semantic segmentation is the basis for computers to understand images. Its purpose is to provide a semantic category label for each pixel in the image, realize the segmentation of different categories in the image, and represent different categories with different colors, and the same category with the same color. Image semantic segmentation is one of the mainstream tasks in the field of computer vision technology, and is widely used in driverless cars, medical imaging, remote sensing images and other fields. Traditional semantic segmentation methods design reasonable criteria based on the grayscale, color, structure, texture and other characteristics of the image, and compare the pixels in the image with one or more thresholds one by one, so as to segment the image into non-overlapping areas. However, traditional semantic segmentation methods rely heavily on manual labor and have low segmentation accuracy. In recent years, the application of artificial intelligence has promoted the rapid development of computer vision technology. The emergence of deep learning has provided new research ideas for semantic segmentation technology. At present, the main frameworks of semantic segmentation are the encoder-decoder structure based on convolutional neural networks and the pyramid structure. The encoder obtains deep semantic features through convolution and pooling operations, and the decoder projects the low-resolution features learned by the encoder into the pixel space to obtain dense classification. The spatial pyramid pooling module of the pyramid structure can obtain multi-scale features and more semantic information. Compared with traditional semantic segmentation technologies, these semantic segmentation technologies based on deep learning have faster calculation speed and higher accuracy. FCN realized end-to-end, pixel-to-pixel semantic segmentation for the first time, which greatly improved the accuracy. However, due to the fixed receptive field and simple jump connection structure, large targets were mislabeled, small targets were ignored, and edge details and other information were lost. U-Net is improved on the basis of FCN. It is a completely symmetrical encoding and decoding structure. It is only suitable for medical image segmentation and not for complex street scene segmentation. The network structure of the DeepLab series obtains a larger receptive field and multi-scale information by improving convolution (expanded convolution) and using multi-scale modules (spatial pyramid pooling). DANet pays better attention to target segmentation by calculating the weights of channels and spatial positions. However, due to the limitations of the convolution kernel, the actual receptive field is smaller than the theoretical receptive field, resulting in limited contextual information extracted by the network, which is crucial for scene understanding tasks. In addition, convolution and pooling operations will lose spatial information, resulting in the loss of edge detail information and small target features in the process of upsampling to restore resolution, and simple fusion methods cannot achieve feature fusion well. Therefore, it is necessary to design a semantic segmentation network with a multi-scale attention mechanism. Summary of the invention
[0003] The technical problem to be solved by the present invention is to provide an image semantic segmentation method based on a multi-scale attention mechanism, integrate the three characteristics of network multi-scale, adaptability and globality, and propose a new segmentation network based on a multi-scale attention mechanism to solve the classification difficulties caused by the presence of different scales of the same object in complex scene segmentation, solve the problems of missing small-scale targets, mis-cutting large-scale targets, unclear target segmentation edges, etc., and optimize the segmentation effect.
[0004] The image semantic segmentation method based on the multi-scale attention mechanism includes the following steps, and the following steps are performed in sequence:
[0005] Step 1: Divide the complex scene image dataset into a training set, a test set, and a validation set, and preprocess the images of the training set and the corresponding label data to obtain images of the same size;
[0006] Step 2: construct a semantic segmentation network model, including a downsampling module, a multi-scale feature extraction module, a feature fusion module and an upsampling module; the image obtained in step 1 is subjected to a Gaussian convolution kernel to generate images of different resolutions to form an image pyramid;
[0007] Step 3: The image pyramid obtained in step 2 is used as the input of the segmentation network, and the original image is obtained through the downsampling module. Figure 1 / 4 resolution feature map, the output information of the downsampling module is divided into two paths, one of which is used as the input of the multi-scale feature extraction module, and the other keeps the fine-grained features with high resolution;
[0008] Step 4: The information input into the multi-scale feature extraction module in step 3 is used to obtain multi-scale semantic information through the spatial pyramid, and double attention is added during extraction, and high-level multi-scale semantic information of fused attention is obtained through the feature fusion module;
[0009] Step 5: The high-level multi-scale semantic information obtained by integrating attention in step 4 is fused with high-resolution features through an upsampling module to segment smaller targets, thereby obtaining multi-scale semantic information that optimizes target edges and details;
[0010] Step 6: The multi-scale semantic information of the optimized target edge and details obtained in step 5 and the fine-grained features obtained in step 3 that maintain high resolution are fused into a plurality of feature maps, and the self-attention of the plurality of feature maps is calculated, and a spatial long-distance pixel relationship is established to obtain a plurality of feature maps containing self-attention;
[0011] Step 7: The multiple feature maps containing self-attention obtained in step 6 are fused again through the decoder to generate a segmentation result map;
[0012] At this point, the image semantic segmentation method based on multi-scale attention mechanism is completed.
[0013] In step five, the upsampling module uses an up-pooling method to obtain a sparse high-resolution feature map, and uses a deconvolution method to convert the sparse feature map into a dense feature map; resolution recovery uses a reserved pooling index to retain spatial information, improve segmentation accuracy, and perform small target segmentation.
[0014] In the step 2, the image pyramid adopts a Gaussian pyramid, convolves the source image with a Gaussian convolution kernel, deletes the even rows and even columns after the convolution, and obtains a reduced image.
[0015] Through the above design scheme, the present invention can bring the following beneficial effects: an image semantic segmentation method based on a multi-scale attention mechanism uses a hollow space pyramid structure to extract high-level multi-scale semantic features, adds channel attention and spatial attention to the high-level multi-scale semantic features, and can choose to aggregate context information of different categories, so that the target representation is more efficient. The images in the image pyramid are input into the network in parallel, which ensures the segmentation efficiency of the network. Among them, large-resolution images are conducive to the segmentation of small targets, and small-resolution images are conducive to the segmentation of large targets. The entire segmentation network has multiple outputs. By using the self-attention mechanism to fuse the output features of different scales, the self-attention mechanism models the relationship between global pixels and pixels, obtains more effective context features, and to a certain extent, retains useful information, filters out useless information, and reduces the computational complexity of the network. The upsampling method of up-pooling and deconvolution can well reconstruct the target structure. By using the pooling index, the spatial information is retained, the segmentation accuracy is increased, and the deconvolution makes the sparse low-resolution features dense. The modules cooperate with each other to achieve better segmentation effects in complex street scene segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The present invention is further described below with reference to the accompanying drawings and specific embodiments:
[0017] Figure 1 This is a schematic block diagram of the image semantic segmentation method based on the multi-scale attention mechanism of the present invention. DETAILED DESCRIPTION
[0018] The image semantic segmentation method based on the multi-scale attention mechanism designs an image pyramid to capture the features of different scales in the image. Large-scale images are conducive to capturing small target features, while small-scale images are conducive to capturing large-scale features. Each image in the image pyramid passes through the segmentation network in parallel to produce multiple segmentation results, which are fused using the self-attention mechanism to obtain the final prediction image.
[0019] Specifically, Figure 1As shown, it includes the following steps, which are executed in sequence:
[0020] Step 1: Divide the street scene dataset into a training set, a test set, and a validation set. Preprocess the images in the training set and the corresponding label data to obtain images of the same size, expand the number of training samples, and make the segmentation model more robust. The data preprocessing methods include image cutting, data balancing, and data enhancement.
[0021] Step 2: Build a semantic segmentation network model. This model includes a downsampling module, a multi-scale feature extraction module, a feature fusion module, and an upsampling module.
[0022] Step three: The training images are down-sampled in stages to obtain images of different resolutions, forming an image pyramid to capture target information of different scales.
[0023] The present invention specifically adopts a Gaussian pyramid, which is to convolve the source image with a Gaussian convolution kernel, and delete the even rows and even columns after the convolution to obtain a reduced image, which is expressed as follows:
[0024]
[0025] In the formula, G i , G i+1 Respectively represent the i-th and i+1-th layer Gaussian images, represents convolution, k represents the Gaussian convolution kernel, the Gaussian convolution kernel is generally selected to be 3×3 or 5×5 in size, and D represents the deletion of the even columns and even rows of the image after convolution.
[0026] Step 4: The image pyramid obtained in step 3 is used as the input of the segmentation network, and the original image is obtained through the downsampling module. Figure 1 / 4 resolution feature map, the output of this module can be divided into two paths, one as the input of the multi-scale feature extraction module, and the other to maintain high-resolution fine-grained features.
[0027] Step 5: The core of the multi-scale feature extraction module is the spatial pyramid and dual attention mechanism (channel attention and spatial attention). The shallow multi-scale information of the image is obtained through the input of step 4. The dual attention mechanism ensures the extraction of features and reduces the amount of parameter calculation during fusion. The shallow multi-scale information is used as the input of the multi-scale feature extraction module. The spatial pyramid structure is used to extract deeper and more advanced multi-scale semantic information, and the dual attention mechanism is used to calculate the attention of the original shallow multi-scale information. The attention is then fused with the generated deep multi-scale semantic information to remove redundancy in order to better focus on the target area.
[0028] After the complex scene image is segmented by the network, a deeper semantic feature map is generated. In the field of semantic segmentation, most of the cross entropy losses are used to optimize the network. The cross entropy is usually the pixel-by-pixel prediction classification result. The specific expression is as follows:
[0029]
[0030] Where L represents the cross entropy loss function, N represents the number of pixels, where n∈N, M represents the number of categories, where c∈M, y c Indicates that the pixel belongs to the label of category c, p c Represents the probability of belonging to this category.
[0031] Step 6: The high-level multi-scale semantic information obtained by fusion attention in step 5 is fused with high-resolution features through the upsampling module to achieve segmentation of smaller targets, and the target edge and detail information is optimized. In the upsampling module, we use up-pooling to obtain sparse high-resolution feature maps, and use deconvolution to make the sparse feature maps dense. In the process of resolution restoration, the spatial information is retained by retaining the pooling index, thereby improving the segmentation accuracy.
[0032] Step 7. The image pyramid mentioned in step 3 is input into the segmentation network in parallel. For the multiple outputs of the entire segmentation network, the self-attention mechanism commonly used in Transformer is used to realize the fusion of multi-channel features, realize the interaction of global information in different scale spaces, retain more effective features, and further improve the accuracy of the semantic segmentation network.
[0033] Specifically, the specific working process of the self-attention mechanism is: obtain the query matrix q and the queried matrix k from the input, calculate the attention A1 between the input features, convert the similarity into probability by the softmax function, and multiply it with the key value matrix v to obtain the result O after attention, and realize feature enhancement; the expression is as follows:
[0034] A=K T Q
[0035] A1=softmax(A)
[0036] O=V·A1
[0037] Among them, K, Q, V represent the sets of input k, q, v respectively.
Claims
1. An image semantic segmentation method based on a multi-scale attention mechanism, characterized by: The method comprises the following steps, and the following steps are performed in sequence: Step 1: Divide the complex scene image dataset into a training set, a test set, and a validation set, and preprocess the images of the training set and the corresponding label data to obtain images of the same size; Step 2: construct a semantic segmentation network model, including a downsampling module, a multi-scale feature extraction module, a feature fusion module, an upsampling module and a self-attention module; the image obtained in step 1 is subjected to a Gaussian convolution kernel to generate images of different resolutions to form an image pyramid; Step 3: The image pyramid obtained in step 2 is used as the input of the segmentation network, and a feature map with a resolution of 1 / 4 of the original image is obtained through a downsampling module. The output information of the downsampling module is divided into two paths, one of which is used as the input of the multi-scale feature extraction module, and the other maintains high-resolution fine-grained features; Step 4: The information input into the multi-scale feature extraction module in step 3 is used to obtain multi-scale semantic information through the spatial pyramid, and double attention is added during extraction, and high-level multi-scale semantic information of fused attention is obtained through the feature fusion module; Step 5: The high-level multi-scale semantic information obtained by integrating attention in step 4 is fused with high-resolution features through an upsampling module to segment smaller targets, thereby obtaining multi-scale semantic information that optimizes target edges and details; Step 6: The multi-scale semantic information of the optimized target edge and details obtained in step 5 and the fine-grained features obtained in step 3 that maintain high resolution are fused into a plurality of feature maps, and the self-attention of the plurality of feature maps is calculated, and a spatial long-distance pixel relationship is established to obtain a plurality of feature maps containing self-attention; Step 7: The multiple feature maps containing self-attention obtained in step 6 are fused again through the decoder to generate a segmentation result map; At this point, the image semantic segmentation method based on multi-scale attention mechanism is completed.
2. The image semantic segmentation method based on the multi-scale attention mechanism according to claim 1 is characterized by: In step five, the upsampling module uses an up-pooling method to obtain a sparse high-resolution feature map, and uses a deconvolution method to convert the sparse feature map into a dense feature map; resolution recovery uses a reserved pooling index to retain spatial information, improve segmentation accuracy, and perform small target segmentation.
3. The image semantic segmentation method based on the multi-scale attention mechanism according to claim 1 is characterized in that: In the step 2, the image pyramid adopts a Gaussian pyramid, convolves the source image with a Gaussian convolution kernel, deletes the even rows and even columns after the convolution, and obtains a reduced image.