A multi-scale Transformer image semantic segmentation method based on convolutional local enhancement
Through the multi-scale Transformer image semantic segmentation method, combined with the Conv Stem layer and feature fusion decoder, the problem of inefficient extraction of local visual structures and insufficient multi-scale feature capture capabilities is solved, achieving a more efficient image segmentation effect.
Patent Information
- Application Number
- CN202311105711.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-30
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2043-08-30
AI Technical Summary
The Transformer structure lacks the inductive bias of local visual structures in image semantic segmentation, which makes it less efficient in extracting local visual structures than convolutional structures. Moreover, because the attention calculation of a single scale limits the capture ability of multi-scale features, the segmentation accuracy is not high.
The multi-scale Transformer image semantic segmentation method is adopted, and the Conv Stem layer, multi-scale feature enhancement extraction module, feature fusion decoder and semantic segmentation module are combined with the MSF-PE module and the MST-Transformer module to enhance the extraction of multi-scale information and capture of local information. The multi-scale vector Transformer encoding module and local information enhancement module achieve complementary advantages and improve segmentation performance.
It improves the accuracy of image semantic segmentation and capture ability of multi-scale features, and enhances the segmentation effect of complex real environments.
Smart Images

Figure CN117058392B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image semantic segmentation in computer vision, and in particular to a multi-scale Transformer image semantic segmentation method based on convolutional local enhancement. Background Art
[0002] With the rapid development and application of artificial intelligence (AI) in recent years, an era of AI is imminent. A large number of intelligent scenarios using AI technologies have sprung up, such as face unlocking for mobile phones, autonomous driving, smart outfit recommendations, and intelligent medical imaging diagnosis. These computer vision-based intelligent applications have become inseparable from modern life. Image semantic segmentation, as one of the foundational technologies for scene understanding, is an indispensable component of real-world intelligence. Currently, image semantic segmentation has been widely applied in real-world scenarios such as medical image segmentation, precision agriculture, geological testing, and autonomous driving. Image semantic segmentation involves computers performing pixel-by-pixel classification based on relevant semantic information in an image. This semantic information typically includes low-level semantic information at the visual level and high-level semantic information at the conceptual level. By fully leveraging semantic information at different levels, computers can classify different objects in an image at the pixel level.
[0003] Traditional image segmentation algorithms combine traditional mathematical methods with computer science to perform segmentation. Their primary goal is to partition a digital image into several non-overlapping subregions, ensuring that pixels within each subregion have similar attributes while maintaining distinct attributes across subregions. Since the rise of deep learning, convolutional neural networks (CNNs) have become the dominant neural network model in computer vision due to their image-friendly properties, such as weight sharing, local perception, and translation invariance. However, with the advancement of deep learning and the growth of datasets, the performance of CNNs has begun to be limited by limitations of their convolutional architecture, such as limited effective receptive fields, excessively strong inductive biases, and difficulty handling complex and changing real-world scenarios. Meanwhile, in the field of natural language processing, the Transformer architecture, capable of modeling long-range dependencies, has achieved tremendous success. Because its self-attention mechanism, combined with its powerful global feature extraction capabilities, effectively handles complex problems, the Transformer architecture has been introduced to the field of semantic segmentation, resulting in Transformer segmentation networks that have achieved state-of-the-art performance on image segmentation datasets. However, the Transformer structure lacks a preset inductive bias, making it less efficient than convolutional structures at extracting local visual structure. On the other hand, while the Transformer structure has been widely used in computer vision due to its self-attention mechanism, which can model long-range dependencies, most Transformer visual networks use slices of a single size to obtain the corresponding embedding vectors, and the self-attention layer also uses a single-scale matrix for attention calculations. The scale uniformity brought about by this structure inevitably limits the Transformer visual network's ability to capture multi-scale features, resulting in low segmentation accuracy when processing images with objects of multiple scales. Furthermore, because the Transformer structure lacks an inductive bias to model local visual structure, it requires a large amount of data to learn an effective inductive bias. Therefore, when the amount of training data is small or in the early stages of network training, the Transformer structure is less effective than the convolutional structure at extracting local detail information.
[0004] In summary, the Transformer architecture lacks a pre-set inductive bias, making it less efficient than convolutional architectures at extracting local visual structure. Furthermore, the Transformer visual network utilizes single-sized slices to derive the corresponding embedding vectors and uses only single-scale matrices for attention calculations in the self-attention layer. This results in a weaker ability to capture multi-scale features and low segmentation accuracy when processing images with multi-scale objects. Currently, there is an urgent need for a new semantic segmentation method that achieves better segmentation results and is more advantageous than semantic segmentation networks in complex real-world environments. Summary of the Invention
[0005] To solve the above problems in the prior art, the present invention adopts a multi-scale Transformer image semantic segmentation method based on convolutional local enhancement, including: obtaining original images from the ADE20K and Cityscapes datasets, inputting the original images into a multi-scale Transformer semantic segmentation model to obtain image semantic segmentation results, wherein the multi-scale Transformer semantic segmentation model includes a Conv Stem layer, a multi-scale feature enhancement extraction module, a feature fusion decoder, and a semantic segmentation module;
[0006] The multi-scale Transformer semantic segmentation model processes the original image in the following steps:
[0007] S1. Input the original image into the Conv Stem layer for feature extraction to obtain a feature map, which enhances the model's ability to extract low-level features;
[0008] S2. Input the feature map into the multi-scale feature enhancement extraction module to obtain feature maps of different scales;
[0009] S3. Input the feature maps of different scales into the feature fusion decoder to obtain a fused feature map;
[0010] S4. Input the fused feature map into the semantic segmentation module to obtain the image semantic segmentation result.
[0011] Furthermore, the multi-scale feature enhancement extraction module includes multiple MSF-PE modules and multiple MST-Transformer modules; the MSF-PE module and the MST-Transformer module are interconnected and appear in pairs; the MSF-PE module consists of a multi-scale feature slice embedding layer, convolutional layers of different scales, and a multi-scale feature fusion module; the MST-Transformer module consists of a multi-scale vector Transformer encoding module, a local information enhancement module, and a feature intersection module; among them, MSF-PE represents multi-scale feature embedding, and MST-Transformer represents multi-scale vector Transformer.
[0012] The process of multi-scale feature enhancement extraction module processing the input feature map includes:
[0013] S21, inputting the feature map into the first MSF-PE module to obtain a first multi-scale feature map to enhance the ability to extract multi-scale information;
[0014] S22, inputting the first multi-scale feature map and the feature map into the first MST-Transformer module respectively to obtain a first encoded feature map;
[0015] S23, input the encoding feature map output by the previous MST-Transformer module to the next MSF-PE module to obtain the multi-scale feature map output by the current MSF-PE module, and input the current multi-scale feature map and the encoding feature map output by the previous MST-Transformer module to the next MST-Transformer module respectively to obtain the encoding feature map output by the current MST-Transformer module;
[0016] S24. Repeat step S23 until all MSF-PE modules and MST-Transformer modules are passed.
[0017] The process of the MSF-PE module processing the input feature map includes:
[0018] S211, inputting the feature map into the multi-scale feature slice embedding layer for multi-scale feature slice embedding, and using convolution layers of different scales to convolve the feature map after the multi-scale feature slice embedding to obtain a slice embedding feature map, so as to enhance the network's ability to extract multi-scale information;
[0019] S212: Input the slice embedding feature map into the multi-scale feature fusion module for information aggregation to obtain a multi-scale feature map.
[0020] The process of the MST-Transformer module processing the encoded feature map and the multi-scale feature map includes:
[0021] S221, input the multi-scale feature map into the multi-scale vector Transformer encoding module to perform multi-scale self-attention calculation to obtain a multi-scale feature vector sequence;
[0022] S222, inputting the encoded feature map into a local information enhancement module to perform local information enhancement to obtain a local information enhanced feature map;
[0023] S223, multi-scale feature vector sequence and local information enhancement feature Figure 1 Perform feature fusion with the input feature intersection module to obtain the encoded feature map.
[0024] Preferably, the process of performing self-attention calculation on the multi-scale feature map includes:
[0025] Step 1: Perform linear transformation on the multi-scale feature map to obtain matrices Q, K, and V. are the linearly transformed query, key, and value, respectively, where is the vector space, N is the number of slice image blocks, C hid is the number of channels of the feature map;
[0026] Step 2: Use the multi-scale vector combined with the self-attention resampling method to divide the attention head into multiple equal parts to obtain multiple heads, and input K and V into each head;
[0027] Step 3: Each head adopts its own downsampling rate B i , K and V are input into the resampling module for dimensionality reduction to obtain K′ and V′ of different dimensions;
[0028] Step 4. Perform self-attention calculation on K′ and V′ in each head and Q to obtain the output of each head. The self-attention calculation formula is:
[0029]
[0030] Among them, Attention(Q,K′,V′) is the result of self-attention calculation, d k is the number of columns of the Q and K′ matrices;
[0031] Step 5: Concatenate the outputs of each head to obtain a multi-scale feature vector.
[0032] Preferably, in order to enhance the ability to handle multi-scale objects in the self-attention layer, the present invention adopts a multi-scale vector joint self-attention resampling method. This method is optimized based on the joint self-attention resampling method. This method reduces the computational complexity of self-attention in the Transformer semantic segmentation network by resampling specific dimensions of K and V to reduce the amount of computation.
[0033] The local information enhancement module consists of multiple stacked convolutional layers. The process of processing the encoded feature map by this module includes:
[0034] Step 1: Perform a small-scale convolution operation on the input encoded feature map to provide stronger local continuity for the Transformer semantic segmentation network;
[0035] Step 2: Perform layer normalization on the convolutional feature map;
[0036] Step 3: Input the normalized result of the layer into the GELU activation function to complete the local feature extraction operation of a stacked convolutional layer;
[0037] Step 4: Input the result of the previous stacked convolutional layer into the next stacked convolutional layer;
[0038] Step 5. Repeat step 4 until all the stacked convolutional layers are passed.
[0039] The process of feature intersection of the local information enhanced feature map and the multi-scale feature vector includes:
[0040] Step 1: Use the Seq2Img layer to reconstruct the multi-scale feature vector sequence to obtain a multi-scale reconstructed feature map;
[0041] Step 2: Perform maximum pooling operation on the local information enhancement feature map;
[0042] Step 3: perform a splicing operation on the multi-scale reconstructed feature map and the pooled local information enhanced feature map to obtain a spliced feature map;
[0043] Step 4: Input the concatenated feature map into the 1×1 convolutional layer to obtain the encoded feature map.
[0044] The process of the feature fusion decoder fusing feature maps of different scales includes:
[0045] S31. Convolve the feature maps of different scales and convert the number of channels of the convolved feature maps into C o ;
[0046] S32. Perform bilinear interpolation upsampling on the feature map after channel number conversion to obtain a resolution recovery feature map. The different resolution recovery feature maps are spliced in the channel dimension to obtain a dimension size of H×W×4C. o The splicing feature map of
[0047] S33, convolve the spliced feature map and reduce the output feature channel dimension to C o , get the fusion feature map;
[0048] Where H is the height, W is the width, C o is the number of channels of the feature map.
[0049] The process of the semantic segmentation module processing the fused feature map includes: feature interaction and coloring output of the fused feature map to obtain the segmentation result of semantic category division.
[0050] The present invention adopts a multi-scale Transformer image semantic segmentation network based on convolution-enhanced local information. The network uses the MSF-PE module and the MST-JRSA module to introduce information of different granularities into the Transformer visual network, where MST-JRSA is a multi-scale joint resampling self-attention module; the present invention adopts a local information enhancement module and a feature intersection module to achieve complementary advantages with the Transformer structure, thereby improving the segmentation performance of the network; the feature fusion decoder module adopted by the present invention realizes the effective fusion of low-level spatial detail features and high-level semantic information features. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 A network structure diagram of a multi-scale Transformer semantic segmentation network based on convolutional enhancement of local information provided by an embodiment of the present invention;
[0052] Figure 2 Flowchart for implementing the multi-scale Transformer image semantic segmentation method based on convolutional local enhancement provided by an embodiment of the present invention;
[0053] Figure 3 A structural diagram of a multi-scale feature slice embedding module provided in an embodiment of the present invention;
[0054] Figure 4 A structural diagram of a multi-scale feature fusion module provided in an embodiment of the present invention;
[0055] Figure 5 A schematic diagram of multi-scale self-attention scale information provided by an embodiment of the present invention;
[0056] Figure 6 A structural diagram of the joint resampling self-attention module provided in an embodiment of the present invention;
[0057] Figure 7 A structural diagram of a local information enhancement module provided in an embodiment of the present invention;
[0058] Figure 8 A structural diagram of a feature intersection module provided in an embodiment of the present invention;
[0059] Figure 9 This is a structural diagram of the feature fusion decoder provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0060] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0061] The present invention is based on Figure 1 The network structure shown in the figure is designed to provide a multi-scale Transformer semantic network segmentation method based on convolutional enhanced local features. The implementation flowchart is as follows Figure 2 As shown;
[0062] When an original image of size H×W×3 is input, the present invention downsamples and extracts image features by stacking multiple small-scale convolutional layers, thereby enhancing the model's ability to extract low-level features. The extracted 4x downsampled feature map is input into the multi-scale feature slice embedding module for slice embedding, and then combined with convolutions of different scales for feature extraction, thereby enhancing the network's ability to extract multi-scale information.
[0063] The convolution results of different scales are input into the multi-scale feature fusion module for information aggregation. The aggregation results are then input into the multiple stacked multi-scale vector Transformer encoding modules of this stage for self-attention related calculations to obtain feature vectors containing rich granular information. This vector is then input into the efficient feedforward neural network module to transform the extracted features.
[0064] Preferably, the present invention further includes a convolution branch in the MST-Transformer module for enhancing local information. The input of this branch is the input feature map of the MSF-PE module at this stage. After multiple convolutions, the output result and the calculation result of the multi-scale self-attention are input into the feature intersection module for information interaction.
[0065] After completing the information exchange, the results are sent to the next stage MSF-PE module to continue the subsequent network operations;
[0066] After the MST-Transformer module extracts and expresses features of different scales on the input image, the present invention inputs the encoded feature maps retained in each stage into the feature fusion decoder, uses small-scale convolution, bilinear interpolation upsampling and splicing operations to merge feature maps of different resolutions, and then performs feature interaction and colorization on the merged feature maps to finally obtain the segmentation results divided by semantic categories;
[0067] Preferably, to introduce multi-path, multi-scale convolution into the slice embedding module to enhance the ability to extract multi-scale information, the present invention proposes a multi-scale feature slice embedding method, abbreviated as MSF-PE. When the MSF-PE module obtains an input feature map, it simultaneously feeds the input feature map into multiple convolution paths with different receptive field sizes to extract features from objects of different scales. Even with different convolution kernel sizes, the output feature map of the same resolution can still be obtained by leveraging the influence of the convolution step size and padding range on the output feature map size.
[0068] In order to accurately evaluate the semantic segmentation ability of an image, this paper adopts an index based on IoU to measure the similarity between the model segmentation results and the real annotations. It measures the degree of overlap between the pixel area predicted by the model and the pixel area of the real annotation.
[0069] Specifically, the four prediction results are True Positive (TP), False Negative (FN), False Positive (FP) and True Negative (TN), as shown in the following table:
[0070]
[0071] As can be seen from the table, TP means that the instance is positive and predicted as a positive example; FN means that the instance is positive but predicted as a negative example; FP means that the instance is negative but predicted as a positive example; TN means that the instance is negative and predicted as a negative example. The calculation formula for the intersection of union and classification is:
[0072]
[0073] Where, IoU i Represents the intersection-over-union ratio of the i-th type of objects.
[0074] After obtaining the intersection-over-union (IoU) of a single category, we can calculate the mIoU value by simply adding up the IoUs of each category in the dataset and averaging them. The formula for calculating mIoU is:
[0075]
[0076] Where mIoU represents the average value of the intersection over union (IoU) of multiple categories of objects, n+1 is the total number of categories in the dataset plus the background category; p ij Indicates the number of pixels whose true category is category i and whose predicted category is category j.
[0077] Specifically, the method comprises the following steps:
[0078] Over 5,000 original images were obtained from the authoritative semantic segmentation datasets ADE20K and Cityscapes. These original images were input into a multi-scale Transformer semantic segmentation model to obtain image semantic segmentation results. The multi-scale Transformer semantic segmentation model includes a Conv Stem layer, a multi-scale feature enhancement extraction module, a feature fusion decoder, and a semantic segmentation module.
[0079] The multi-scale Transformer semantic segmentation model processes the original image in the following steps:
[0080] S1. Input the original image into the Conv Stem layer for feature extraction to obtain a feature map, thereby enhancing the model's ability to extract low-level features;
[0081] S2. Input the feature map into the multi-scale feature enhancement extraction module to obtain feature maps of different scales;
[0082] S3. Input the feature maps of different scales into the feature fusion decoder to obtain a fused feature map;
[0083] S4. Input the fused feature map into the semantic segmentation module to obtain the image semantic segmentation result.
[0084] Furthermore, the multi-scale feature enhancement extraction module includes 4 MSF-PE modules and 4 MST-Transformer modules; the MSF-PE modules are interconnected with the MST-Transformer modules; the MSF-PE module consists of a multi-scale feature slice embedding layer, convolutional layers of different scales, and a multi-scale feature fusion module; the MST-Transformer module consists of a multi-scale vector Transformer encoding module, a local information enhancement module, and a feature intersection module; among them, MSF-PE represents multi-scale feature slice embedding, and MST-Transformer represents multi-scale vector Transformer.
[0085] The process of multi-scale feature enhancement extraction module processing the input feature map includes:
[0086] S21, inputting the feature map into the first MSF-PE module to obtain a first multi-scale feature map to enhance the ability to extract multi-scale information;
[0087] S22, inputting the first multi-scale feature map and the feature map into the first MST-Transformer module respectively to obtain a first encoded feature map;
[0088] S23, input the encoding feature map output by the previous MST-Transformer module to the next MSF-PE module to obtain the multi-scale feature map output by the current MSF-PE module, and input the current multi-scale feature map and the encoding feature map output by the previous MST-Transformer module to the next MST-Transformer module respectively to obtain the encoding feature map output by the current MST-Transformer module;
[0089] S24. Repeat step S23 until four MSF-PE modules and four MST-Transformer modules have been passed.
[0090] The process of MSF-PE module processing the input feature map, such as Figure 3 Shown, including:
[0091] S211, inputting the feature map into the multi-scale feature slice embedding layer for multi-scale feature slice embedding, and using convolution layers of different scales to convolve the feature map after the multi-scale feature slice embedding to obtain a slice embedding feature map;
[0092] S212: Input the slice embedding feature map into the multi-scale feature fusion module for information aggregation to obtain a multi-scale feature map.
[0093] Specifically, when the feature map from the i-1 stage When performing multi-scale slice embedding, where R is the vector space, the MSF-PE module feeds the input feature map into multiple overlapping convolutional slice embedding paths. Each path maintains the same number of input channels during slice embedding, modifying only the spatial resolution. The convolutional slice sizes corresponding to the three paths in the first stage of the MSF-PE module are 3×3, 5×5, and 7×7, respectively. If the MSF-PE module does not need to downsample the feature map, the convolution stride is set to 1; otherwise, it is set to the downsampling factor. The convolution padding is flexibly adjusted to ensure that the feature maps output by each convolution path remain consistent in spatial dimensions.
[0094] The multi-scale feature fusion module aggregates information from the input slice embedding feature map, such as Figure 4 Shown, including:
[0095] Step 1: For the multi-scale feature fusion module of stage i, the invention first embeds the slices obtained by three different convolution paths into the feature map F3×3 、F 5×5 and F 7×7 The process of splicing in the channel dimension is expressed as:
[0096] F concat =Concat(F 3×3 ,F 5×5 ,F 7×7 )
[0097] Where, F concat is the concatenated feature map, It is the splicing operation of the feature map channel dimension;
[0098] Step 2: F concat 1×1 convolution, 3×3 convolution and 1×1 convolution are performed in sequence. The first 1×1 convolution changes the channel dimension of the feature map from 3C i-1 Expanded to 4C i-1 , and use the GELU activation function to increase the nonlinear expression ability of the module to obtain the output feature map F mid_1 , the process expression is:
[0099] F mid_1 =GELU(1×1_Conv_1(F concat ))
[0100] Among them, C i-1 is the number of channels of the feature map, GELU(·) is the GELU activation function;
[0101] Step 3: Use 3×3 convolution to make F mid_1 The dimension remains unchanged, and then the layer is normalized to obtain the output feature map F mid_2 , the process expression is:
[0102] F mid_2 =LN(3×3_Conv(F mid_1 ))
[0103] Where LN(·) is the layer normalization operation;
[0104] Step 4: The second 1×1 convolution will mid_2 The channel dimension is reduced to C i , get the i-th output feature map The process expression is:
[0105]
[0106] Among them C i is the number of channels of the feature map.
[0107] Furthermore, in order to improve performance and reduce computational complexity, the convolutions in the MSF-PE modules of the present invention all adopt depthwise separable convolutions, and the convolution modules all include layer normalization operations and GELU activation functions.
[0108] Preferably, in order to improve performance and reduce computational complexity, the process of processing the encoded feature map and the multi-scale feature map by the MST-Transformer module includes:
[0109] S221, input the multi-scale feature map into the multi-scale vector Transformer encoding module to perform multi-scale self-attention calculation to obtain a multi-scale feature vector sequence;
[0110] S222, inputting the encoded feature map into a local information enhancement module to perform local information enhancement to obtain a local information enhanced feature map;
[0111] S223, multi-scale feature vector sequence and local information enhancement feature Figure 1 Perform feature interaction with the input feature intersection module to obtain the encoded feature map.
[0112] The process of calculating self-attention on multi-scale feature maps includes:
[0113] Step 1: Perform linear transformation on the multi-scale feature map to obtain matrices Q, K, and V. are the linearly transformed query, key, and value, respectively, where is the vector space, N is the number of slice image blocks, C hid is the number of channels of the feature map;
[0114] Step 2: Use the multi-scale vector combined with the self-attention resampling method to divide the attention head into multiple equal parts to obtain multiple heads, and input K and V into each head;
[0115] Step 3: Each head adopts its own downsampling rate B i , K and V are input into the resampling module for dimensionality reduction to obtain K′ and V′ of different dimensions;
[0116] Step 4. Perform self-attention calculation on K′ and V′ in each head and Q to obtain the output of each head. The self-attention calculation formula is:
[0117]
[0118] Among them, Attention(Q,K′,V′) is the result of self-attention calculation, d k is the number of columns of the Q and K′ matrices;
[0119] Step 5: Concatenate the outputs of each head to obtain a multi-scale feature vector.
[0120] Preferably, the multi-scale vector joint self-attention resampling method is optimized based on the joint self-attention resampling method, and its multi-scale self-attention scale information diagram is as follows: Figure 5 , for the joint self-attention resampling module, the design of this module is to reduce the amount of calculation by resampling the specific dimensions of K and V, thereby reducing the computational complexity of self-attention in the Transformer semantic segmentation network. The schematic diagram of this module is Figure 6 The specific principles and steps are as follows:
[0121] Step 1: Input the three dimensions Q, K, and V of the vector into the self-attention module, where Q, K, and V are the three input representations of the self-attention mechanism;
[0122] Step 2: Embed Refactored to The feature map of is a vector space, C hid is the number of channels of the feature map, H and W are the height and width of the reconstructed feature map, respectively, and their values are equal to the number of sliced image blocks N;
[0123] Step 3: Use computationally efficient depth-wise separable convolution and pooling to perform K r and V r Resample, the output after convolution resampling is and The output after maximum pooling resampling is and
[0124] Step 4: K r All outputs and Add them together to get X′ K ′, V r All outputs and Add up to get X″ V , put X″ K and X″ V Integrate effective feature information in the input convolution layer and normalization layer to obtain X″′ K and X″′ V ;
[0125] Step 5: To ensure that the overall calculation method of scaled dot product attention remains unchanged, X″′ k and X″′ V The dimension of is reconstructed into the dimension form of the embedding vector, and we get and After that, the self-attention calculation is performed with Q, and the output dimension is the same as the standard self-attention calculation, where is a vector space, C hid is the number of channels of the feature map, and the hyperparameter B i Represents the downsampling rate of the i-th stage.
[0126] Furthermore, in the multi-scale vector joint self-attention mechanism, the basic rules of the vector joint self-attention mechanism remain unchanged. The multi-scale vector joint self-attention mechanism divides the head into multiple equal parts, and only between different attention heads, the hyperparameter B used to control the resampling range is used. i No longer consistent.
[0127] Preferably, in the multi-scale vector joint self-attention mechanism, the self-attention calculations in different heads do not affect each other. i Setting it to a smaller level makes the retained K and V dimensions larger, so that the self-attention layer of the corresponding head can retain more fine-grained information; and setting the sampling rate of some heads to a larger level allows more K and V to be fused, which significantly reduces the amount of self-attention calculation while enhancing the model's ability to capture large-scale objects.
[0128] The local information enhancement module consists of three stacked convolutional layers, such as Figure 7 As shown in Figure 2, the process of processing the encoded feature map by this module includes:
[0129] Step 1: Perform a small-scale convolution operation on the input encoded feature map to provide stronger local continuity for the Transformer semantic segmentation network;
[0130] Step 2: Perform layer normalization on the convolutional feature map;
[0131] Step 3: Input the normalized result of the layer into the GELU activation function to complete the local feature extraction operation of a stacked convolutional layer;
[0132] Step 4: Input the result of the previous stacked convolutional layer into the next stacked convolutional layer;
[0133] Step 5. Repeat step 4 until all the stacked convolutional layers are passed.
[0134] The process of feature intersection of local information enhancement feature map and multi-scale feature vector, such as Figure 8 Shown, including:
[0135] Step 1: Use the Seq2Img layer to reconstruct the multi-scale feature vector sequence to obtain a multi-scale reconstructed feature map, and perform a maximum pooling operation on the local information enhancement feature map;
[0136] Step 2: Concatenate the multi-scale reconstructed feature map with the pooled local information enhanced feature map to obtain a concatenated feature map.
[0137] Step 3: Input the concatenated feature map into the 1×1 convolutional layer to obtain an encoded feature map containing rich multi-scale information and local information.
[0138] Preferably, in order to adapt to the feature map size output by the multi-scale vector Transformer encoding module, when the feature map sizes of the two paths input do not match, the feature intersection module uses the maximum pooling layer to downsample the local information enhancement feature map, and the downsampling ratio is consistent with the downsampling ratio of the MSF-PE module at this stage.
[0139] Furthermore, after the multi-scale feature enhancement extraction module completes the extraction and expression of features at different scales of the input image, the encoded feature maps retained in the previous four stages are input into the feature fusion decoder, and small-scale convolution, bilinear interpolation upsampling and splicing operations are used to merge the four feature maps of different resolutions, thereby improving the model's ability to predict multiple features. The structural diagram of the feature fusion decoder is shown in the figure. Figure 9 As shown, the specific process includes:
[0140] S31. Convolve the feature maps of different scales in the four stages and convert the number of channels of the convolved feature maps into C o ;
[0141] S32. Perform bilinear interpolation upsampling on the feature map after channel number conversion to obtain a resolution recovery feature map. The different resolution recovery feature maps are spliced in the channel dimension to obtain a dimension size of H×W×4C. o The splicing feature map of
[0142] S33, convolve the spliced feature map and reduce the output feature channel dimension to C o , get the fusion feature map;
[0143] Where H is the height, W is the width, C o is the number of channels of the feature map.
[0144] The process of semantic segmentation module processing the fused feature map includes: feature interaction on the fused feature map, that is, combining different features together to generate a richer feature representation; coloring the feature map after feature interaction, and outputting the final pixel-by-pixel segmentation result.
[0145] The above embodiments further illustrate the purpose, technical solutions and advantages of the present invention in detail. It should be understood that the above embodiments are only preferred implementation plans of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made to the present invention within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A multi-scale Transformer image semantic segmentation method based on convolutional local enhancement, characterized in that: include: Obtain the original image and input it into the multi-scale Transformer semantic segmentation model to obtain the image semantic segmentation result. The multi-scale Transformer semantic segmentation model includes the Conv Stem layer, the multi-scale feature enhancement extraction module, the feature fusion decoder, and the semantic segmentation module; The multi-scale Transformer semantic segmentation model processes the original image in the following steps: S1. Input the original image into the Conv Stem layer for feature extraction to obtain a feature map; S2. Input the feature map into the multi-scale feature enhancement extraction module to obtain feature maps of different scales; S3. Input the feature maps of different scales into the feature fusion decoder to obtain a fused feature map; S4. Input the fused feature map into the semantic segmentation module to obtain the image semantic segmentation result; The multi-scale feature enhancement extraction module includes multiple MSF-PE modules and multiple MST-Transformer modules. The MSF-PE modules and MST-Transformer modules are interconnected and appear in pairs. The MSF-PE module consists of a multi-scale feature slice embedding layer, convolutional layers of different scales, and a multi-scale feature fusion module. The MST-Transformer module consists of a multi-scale vector Transformer encoding module, a local information enhancement module, and a feature intersection module. Among them, MSF-PE represents multi-scale feature slice embedding, and MST-Transformer represents multi-scale vector Transformer. The process of multi-scale feature enhancement extraction module processing the input feature map includes: S21, inputting the feature map into the first MSF-PE module to obtain a first multi-scale feature map; S22, inputting the first multi-scale feature map and the feature map into the first MST-Transformer module respectively to obtain a first encoded feature map; S23, input the encoding feature map output by the previous MST-Transformer module to the next MSF-PE module to obtain the multi-scale feature map output by the current MSF-PE module, and input the current multi-scale feature map and the encoding feature map output by the previous MST-Transformer module to the next MST-Transformer module respectively to obtain the encoding feature map output by the current MST-Transformer module; S24. Repeat step S23 until all MSF-PE modules and MST-Transformer modules are passed; The process of the MSF-PE module processing the input feature map includes: S211, inputting the feature map into the multi-scale feature slice embedding layer for multi-scale feature slice embedding, and using convolution layers of different scales to convolve the feature map after the multi-scale feature slice embedding to obtain a slice embedding feature map; S212, inputting the slice embedding feature map into the multi-scale feature fusion module for information aggregation to obtain a multi-scale feature map; The process of the MST-Transformer module processing the encoded feature map and the multi-scale feature map includes: S221, input the multi-scale feature map into the multi-scale vector Transformer encoding module to perform multi-scale self-attention calculation to obtain a multi-scale feature vector sequence; S222, inputting the encoded feature map into a local information enhancement module to perform local information enhancement to obtain a local information enhanced feature map; S223, inputting the multi-scale feature vector sequence and the local information enhanced feature map into the feature intersection module for feature interaction to obtain a coding feature map; The local information enhancement module consists of multiple stacked convolutional layers. The process of processing the encoded feature map by this module includes: Step 1: Perform a small-scale convolution operation on the input encoding feature map; Step 2: Perform layer normalization on the convolutional feature map; Step 3: Input the normalized result of the layer into the GELU activation function to complete the local feature extraction operation of a stacked convolutional layer; Step 4: Input the result of the previous stacked convolutional layer into the next stacked convolutional layer; Step 5. Repeat step 4 until all the stacked convolutional layers are passed.
2. The multi-scale Transformer image semantic segmentation method based on convolutional local enhancement according to claim 1, characterized in that: The process of calculating self-attention on multi-scale feature maps includes: Step 1: Perform linear transformation on the multi-scale feature map to obtain matrices Q, K, and V. are the linearly transformed query, key, and value, respectively, where is the vector space, N is the number of slice image blocks, C hid is the number of channels of the feature map; Step 2: Use the multi-scale vector combined with the self-attention resampling method to divide the attention head into multiple equal parts to obtain multiple heads, and input K and V into each head; Step 3: Each head adopts its own downsampling rate B i , K and V are input into the resampling module for dimensionality reduction to obtain K′ and V′ of different dimensions; Step 4. Perform self-attention calculation on K′ and V′ in each head and Q to obtain the output of each head. The self-attention calculation formula is: Among them, Attention(Q,K′,V′) is the result of self-attention calculation, d k is the number of columns of the Q and K′ matrices; Step 5: Concatenate the outputs of each head to obtain a multi-scale feature vector.
3. The multi-scale Transformer image semantic segmentation method based on convolutional local enhancement according to claim 1, characterized in that: The process of performing feature intersection on the local information enhancement feature map and the multi-scale feature vector in step S223 includes: Step 1: Use the Seq2Img layer to reconstruct the multi-scale feature vector sequence to obtain a multi-scale reconstructed feature map; Step 2: Perform maximum pooling operation on the local information enhancement feature map; Step 3: Concatenate the multi-scale reconstructed feature map with the pooled local information enhanced feature map to obtain a concatenated feature map. Step 4: Input the concatenated feature map into the 1×1 convolutional layer to obtain the encoded feature map.
4. The multi-scale Transformer image semantic segmentation method based on convolutional local enhancement according to claim 1, characterized in that: The process of the feature fusion decoder fusing feature maps of different scales includes: S31. Convolve the feature maps of different scales and convert the number of channels of the convolved feature maps into C o ; S32. Perform bilinear interpolation upsampling on the feature map after channel number conversion to obtain a resolution recovery feature map. The different resolution recovery feature maps are spliced in the channel dimension to obtain a dimension size of H×W×4C. o The splicing feature map of S33, convolve the spliced feature map and reduce the output feature channel dimension to C o , get the fusion feature map; Where H is the height, W is the width, C o is the number of channels of the feature map.
5. The multi-scale Transformer image semantic segmentation method based on convolutional local enhancement according to claim 1, characterized in that: The process of the semantic segmentation module processing the fused feature map includes: feature interaction and coloring output of the fused feature map to obtain the segmentation result of semantic category division.
Citation Information
Patent Citations
Knife switch state identification method and device, computer equipment and storage medium
CN115719416A
Image processing method and system based on double-branch multi-scale semantic segmentation network
CN116580241A