Image semantic segmentation method based on GMDSEg network
By introducing a global-local feature fusion module and a multi-scale dynamic scanning module, the problem that existing image semantic segmentation methods cannot simultaneously achieve global context modeling, local detail encoding, and multi-scale feature extraction is solved, thus achieving efficient image semantic segmentation results.
Patent Information
- Application Number
- CN202511180599.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-12-05
AI Technical Summary
Existing image semantic segmentation methods struggle to simultaneously achieve efficient global context modeling, high-quality local detail encoding, and rich multi-scale feature extraction, resulting in poor segmentation performance.
A global-local feature fusion module and a multi-scale dynamic scanning module are introduced. The global-local feature fusion module fuses local and global information, and the multi-scale dynamic scanning module extracts contextual information at different scales, thereby achieving efficient semantic segmentation feature modeling.
It improves the accuracy and segmentation effect of image semantic segmentation, and can generate more accurate segmentation results, with advanced performance.
Smart Images

Figure CN121074401A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of image semantic segmentation, and particularly relates to an image semantic segmentation method based on a GMDSeg network. BACKGROUND
[0002] Image semantic segmentation is an important task in the field of computer vision, widely applied in many fields such as automatic driving of cars, medical image analysis, and part defect detection. Image semantic segmentation aims to assign semantic labels to each pixel in an image, requiring the identification of different objects in the image while accurately depicting their boundaries and shapes, thereby achieving fine understanding and multi-level analysis of the image. In recent years, due to the availability of large-scale datasets, powerful computing resources, and advanced network architectures, deep learning methods have made significant achievements in image semantic segmentation.
[0003] Early research on image semantic segmentation was mainly based on convolutional neural networks (CNN) for feature extraction of input images. In recent years, with the introduction of the Transformer network, a series of image semantic segmentation methods based on visual Transformer networks have been proposed, significantly improving the accuracy of image semantic segmentation and achieving competitive performance.
[0004] Accurate semantic segmentation requires three key functions: first, global context modeling: establishing rich context dependencies without spatial distance constraints to achieve overall scene understanding; second, local detail encoding: providing fine-grained features and boundary representations, which are crucial for distinguishing semantic categories and locating different semantic region boundaries; third, multi-scale based context modeling: promoting cross-multi-scale semantic representation, addressing intra-class scale variation while enhancing inter-class discriminability. However, existing image semantic segmentation methods are difficult to simultaneously possess all these functions. Therefore, how to enable the segmentation network to simultaneously perform efficient global context modeling, high-quality local detail encoding, and rich multi-scale feature extraction remains a challenging topic. SUMMARY
[0005] The content of the present application is to provide a kind of image semantic segmentation method (Global-Local Feature Fusion and Multi-Scale Dynamic Scanning For Semantic Segmentation, GMDSeg) based on global-local feature fusion and multi-scale dynamic scanning.In the encoding process, introduce global-local feature fusion module (Global-Local Feature Fusion Module, GFM) to fuse local and global information, to realize efficient context learning and scene understanding.In the decoding process, introduce multi-scale dynamic scanning module (Multi-Scale Dynamic Scanning Module, MDSM) to extract context information of different scales, to realize efficient and resolution robust semantic segmentation feature modeling.
[0006] The image semantic segmentation method based on GMDSeg network of the present application comprises the following steps:
[0007] 1) extract the multi-scale semantic feature map of the input image, comprising the following steps:
[0008] 1.1 GMDSeg network uses a standard four-stage step-by-step downsampling encoder to perform step-by-step downsampling on the input network image, generating multi-scale semantic feature maps S1, S2, S3, S4, with scale sizes of:
[0009] H / 4×W / 4, H / 8×W / 8, H / 16×W / 16, H / 32×W / 32
[0010] Wherein: H is the height of the original image, W is the width of the original image; in order to reduce the power consumption and ensure the working efficiency of the entire network, the largest scale semantic feature map S1 is discarded in the work and does not participate in the subsequent decoding process; the remaining three smaller scale semantic feature maps S2, S3, S4 are sent to the decoder for subsequent decoding;
[0011] 1.2 In the encoding process, a global-local feature fusion module is used after each stage of downsampling module to fuse local and global information; the global-local feature fusion module is composed of batch normalization (Batch Normalization, BN), neighborhood attention (Neighborhood Attention, NA), 2D selective scanning block, convolution layer and feed-forward network (Feed-Forward Network, FFN), which limits the attention range of each pixel to its adjacent area, maintains translational equivalence, effectively captures local dependencies, and has linear time complexity;
[0012] Specifically comprising the following steps:
[0013] 1.2.1 For the input feature map, named S in , after batch normalization, local information S l is extracted by neighborhood attention, and the mathematical expression is as follows:
[0014] S l = NA (BN (S in )) (1)
[0015] Wherein: NA is neighborhood attention; BN is batch normalization;
[0016] 1.2.2 The local information S l is input into the 2D selective scanning block for scanning to obtain global information S g , and the mathematical expression is as follows:
[0017] S g = SS2D (S l ) (2)
[0018] Wherein: SS2D is a 2D selective scanning block;
[0019] 1.2.3 The global information S g and the local information S l are fused by bypassing the 2D selective scanning block, and the output after fusion is further subjected to a convolution layer to obtain further global-local fusion information S f , and the mathematical expression is as follows:
[0020] S f = Conv (S g +S l ) (3)
[0021] Wherein: Conv is a convolution layer with a size of 1x1;
[0022] 1.2.4 The fusion information is stabilized by batch normalization and feedforward network to make up for the deficiency of spatial modeling, and finally the output result S out is obtained, and the mathematical expression of the whole process is as follows:
[0023] S out = BF (S f +S in ) + (S f +S in ) (4)
[0024] Wherein: BF is batch normalization BN+feedforward network FFN;
[0025] 2) Aggregating multi-scale semantic feature maps, comprising the following steps:
[0026] 2.1 Take the multi-scale semantic feature maps S2, S3, S4 extracted in step 1) as input, and perform dimension reduction processing through a dimension reduction module; the dimension reduction module is composed of a convolution layer, a normalization layer (Layer Normalization, LN) and a Swish activation function, S3 and S4 are upsampled to the same spatial scale size as S2 after dimension reduction through bilinear interpolation, to obtain feature maps S3' and S4', and the feature map obtained after dimension reduction processing of S2 is S2', and the mathematical expressions of S2', S3', S4' are respectively:
[0027]
[0028] Wherein: CLS is Conv+LN+Swish; Upsampling(x) is an up-sampling operation on x;
[0029] 2.2 Perform a concatenation operation on S2', S3', S4', and then perform channel dimension reduction through a dimension reduction module, to finally obtain an aggregated feature map S5 with more expressive power and a size of H / 8xW / 8, and the mathematical expression is:
[0030] S5 = CLS(Concat(S2', S3', S4') (6)
[0031] Wherein: Concat is a concatenation operation;
[0032] 3) generating an output image after semantic segmentation, including the following steps:
[0033] 3.1 input the aggregated feature map in step 2) into a multi-scale dynamic scanning module for deep feature extraction; the multi-scale dynamic scanning module is composed of a dimension reduction module, a 2D selective scanning block and a convolution layer, uses a convolution with a stride of 2 to aggregate 3x3 neighborhood features, generates a derived feature map with a size of 1 / 2 of S5, uses a convolution with a stride of 4 to aggregate 5x5 neighborhood features, generates a derived feature map with a size of 1 / 4 of S5, and obtains the region aggregation context through the two derived feature maps;
[0034] 3.2 use a lossless down-sampling operation to reduce the size of S5 and the derived feature map with a size of 1 / 2 to the same H / 32xW / 32 as the derived feature map with a size of 1 / 4; this operation rearranges the 4x4 non-overlapping blocks of S5 and the 2x2 blocks of the derived feature map with a size of 1 / 2 to the channel dimension, respectively increasing the channel depth by 16 times and 4 times, while preserving the original scale information, and then reducing the channel number of the three feature maps to the same number through a dimension reduction module, to ensure that the context information of each scale has equal importance;
[0035] The dimensionality-reduced feature maps are concatenated along the channel dimension and input into a 2D selective scan block to achieve multi-scale context extraction in a single scan. The scanned feature maps are then input into a 1×1 convolutional layer for cross-scale context fusion. Finally, the fused features are reduced to the number of input channels by the dimensionality reduction module and upsampled to the size of H / 8×W / 8 to obtain a new feature map S5' that fuses multi-scale context.
[0036] 3.3 S5' is added to S2', S3' and S4' respectively to enhance their multi-scale contextual information. Then, S5' is concatenated with the enhanced S2', S3' and S4' and the global average pooling result of S4'. After being fused by two consecutive layers of multilayer perceptron (MLP), it is upsampled to the original size of the input image to output an accurate semantic segmentation result map.
[0037] The GMDSeg network proposed in this invention can be effectively applied to the field of image semantic segmentation, achieving relatively accurate segmentation results and exhibiting advanced performance. Attached Figure Description
[0038] Figure 1 A schematic diagram of the GMDSeg network framework;
[0039] Figure 2 Here is a structural diagram of the Global-Local Feature Fusion (GFM) module;
[0040] Figure 3 This is a structural diagram of the Multi-Scale Dynamic Scanning Module (MDSM).
[0041] Figure 4 The images are the original image and the segmentation result image.
[0042] Wherein: (a) is the original image, and (b) is the segmentation result. Detailed Implementation
[0043] The present invention will now be described in conjunction with the accompanying drawings.
[0044] like Figure 1 As shown, an image semantic segmentation method based on the GMDSeg network of the present invention includes the following steps:
[0045] 1) Obtain the dataset and parameter settings, including the following steps:
[0046] 1.1 Obtain the public dataset ADE20K, which contains more than 20,000 images from different environments and situations, covering 150 categories.
[0047] 1.2 The iteration size was determined by the amount of data and the GPU: 160K, batch size = 16, initial learning rate = 0.00012, and weight decay = 0.01. The AdamW optimizer was used for training, and the training image size was adjusted to 512×512.
[0048] 2) Extract multi-scale semantic feature maps from the input image, including the following steps:
[0049] 2.1 The GMDSeg network uses a standard four-stage progressive downsampling encoder to progressively downsample the input image, generating multi-scale semantic feature maps S1, S2, S3, and S4, with the following scales in order:
[0050] H / 4×W / 4, H / 8×W / 8, H / 16×W / 16, H / 32×W / 32
[0051] Where: H is the height of the original image and W is the width of the original image; in order to reduce computational overhead and ensure the efficiency of the entire network, the semantic feature map S1 with the largest scale is discarded during the process and does not participate in the subsequent decoding process; the remaining three smaller-scale semantic feature maps S2, S3 and S4 are sent to the decoder for subsequent decoding.
[0052] 2.2 During the encoding process, after each stage of the downsampling module, a global-local feature fusion module is used to fuse local and global information; for example... Figure 2 As shown, the global-local feature fusion module consists of batch normalization (BN), neighborhood attention (NA), 2D selective scan blocks, convolutional layers, and a feed-forward network (FFN). It limits the attention range of each pixel to its neighboring region, maintains translation equivalence, effectively captures local dependencies, and has linear time complexity.
[0053] Specifically, the following steps are included:
[0054] 2.2.1 The input feature map is named S in After batch normalization, local information S is obtained through neighborhood attention extraction. l Its mathematical expression is:
[0055] S l =NA(BN(S) in )) (1)
[0056] Where: NA represents neighborhood attention; BN represents batch normalization;
[0057] 2.2.2 Transfer local information Sl Input a 2D selective scan block to obtain global information S. g Its mathematical expression is:
[0058] S g =SS2D(S l (2)
[0059] Wherein: SS2D is a 2D selective scan block;
[0060] 2.2.3 Global Information S g and local information S l The fusion is performed by bypassing 2D selective scanning blocks, and the fused output is then passed through a convolutional layer to obtain further global-local fusion information S. f Its mathematical expression is:
[0061] S f =Conv(S g +S l (3)
[0062] Wherein: Conv is a convolutional layer of size 1×1;
[0063] 2.2.4 Information fusion compensates for the deficiencies in spatial modeling by stabilizing the feature distribution through batch normalization and feedforward networks, ultimately yielding the output result S. out The mathematical expression for the entire process is:
[0064] S out =BF(S f +S in )+(S f +S in (4)
[0065] Where: BF is batch normalized BN + feedforward network FFN;
[0066] 3) Aggregate multi-scale semantic feature maps, including the following steps:
[0067] 3.1 The multi-scale semantic feature maps S2, S3, and S4 extracted in step 2) are used as input and dimensionality reduction is performed by a dimensionality reduction module. The dimensionality reduction module consists of convolutional layers, layer normalization (LN) layers, and a Swish activation function. After dimensionality reduction, S3 and S4 are upsampled to the same spatial scale as S2 through bilinear interpolation to obtain feature maps S3' and S4'. The feature map obtained after dimensionality reduction of S2 is S2'. The mathematical expressions for S2', S3', and S4' are as follows:
[0068]
[0069] Where: CLS is Conv+LN+Swish; Upsampling(x) is the upsampling operation on x;
[0070] 3.2 S2', S3', and S4' are concatenated, and then subjected to channel dimensionality reduction by a dimensionality reduction module to finally obtain a more expressive aggregated feature map S5 with a size of H / 8 × W / 8. Its mathematical expression is:
[0071] S5=CLS(Concat(S2′,S3′,S4′) (6)
[0072] Where: Concat is the concatenation operation;
[0073] 4) Generate the semantically segmented output image, including the following steps:
[0074] 4.1 Input the aggregated feature map from step 3) into the multi-scale dynamic scanning module for depth feature extraction; such as Figure 3 As shown, the multi-scale dynamic scanning module consists of a dimensionality reduction module, a 2D selective scanning block, and a convolutional layer. It uses a convolution with a stride of 2 to aggregate 3×3 neighborhood features and generate a derived feature map of size 1 / 2 of the S5 scale. It uses a convolution with a stride of 4 to aggregate 5×5 neighborhood features and generate a derived feature map of size 1 / 4 of the S5 scale. The context of region aggregation is obtained through the two derived feature maps.
[0075] 4.2 Using lossless downsampling, the scale of S5 and the derived feature map of size 1 / 2 is reduced to the same H / 32×W / 32 as the derived feature map of size 1 / 4. This operation rearranges the 4×4 non-overlapping blocks of S5 and the 2×2 blocks of the derived feature map of size 1 / 2 to the channel dimension, increasing the channel depth by 16 times and 4 times respectively, while preserving the original scale information. Then, the number of channels of the three feature maps is reduced to the same number through the dimensionality reduction module to ensure that the context information of each scale is equally important.
[0076] The dimensionality-reduced feature maps are concatenated along the channel dimension and input into a 2D selective scan block to achieve multi-scale context extraction in a single scan. The scanned feature maps are then input into a 1×1 convolutional layer for cross-scale context fusion. Finally, the fused features are reduced to the number of input channels by the dimensionality reduction module and upsampled to the size of H / 8×W / 8 to obtain a new feature map S5' that fuses multi-scale context.
[0077] 4.3 S5' is added to S2', S3' and S4' respectively to enhance their multi-scale contextual information. Then, S5' is concatenated with the enhanced S2', S3' and S4' and the global average pooling result of S4'. After being fused by two consecutive layers of multilayer perceptron (MLP), it is upsampled to the original size of the input image to output an accurate semantic segmentation result map.
[0078] 5) Segmentation results are as follows Figure 4 As shown, the model has an mIoU of 50.2 on the mIoU metric, which measures segmentation performance.
Claims
1. A method for image semantic segmentation based on GMDSeg network, characterized in that Comprising the following steps: 1) Extracting the multi-scale semantic feature map of the input image, comprising the following steps: 1.1 The GMDSeg network uses a standard four-stage step-by-step downsampling encoder to perform step-by-step downsampling on the input network image, generating multi-scale semantic feature maps S1, S2, S3, S4, with scale sizes in turn: H / 4×W / 4, H / 8×W / 8, H / 16×W / 16, H / 32×W / 32 Where: H is the height of the original image, W is the width of the original image; The largest scale semantic feature map S1 is discarded in the work and does not participate in the subsequent decoding process; The remaining three smaller scale semantic feature maps S2, S3, S4 are sent into the decoder for subsequent decoding; 1.2 In the encoding process, a global-local feature fusion module is used after each stage of the downsampling module to realize the fusion of local and global information; The global-local feature fusion module is composed of batch normalization, neighborhood attention, 2D selective scanning block, convolution layer and feedforward network, which limits the attention range of each pixel to its adjacent area, maintains translational equivalence, effectively captures local dependencies, and has linear time complexity; Specifically comprising the following steps: 1.2.1 For the input feature map, named S in , after batch normalization, local information S l is extracted by neighborhood attention, and its mathematical expression is as follows: S l = NA(BN(S in )) (1) Where: NA is neighborhood attention; BN is batch normalization; 1.2.2 Derive local information S l Input 2D selective scanning block to scan and get global information S g The mathematical expression is: S g = SS2D(S l ) (2) Where: SS2D is a 2D selective scanning block; 1.2.3 Global information S g and local information S l The fused output is further convolved to obtain further global-local fused information S f The mathematical expression is: S f = Conv(S g +S l ) (3) Where: Conv is a 1×1 convolution layer; 1.2.4 The fusion information stabilizes the feature distribution through batch normalization and feedforward network, makes up for the deficiency of spatial modeling, and finally obtains the output result S out The mathematical expression of the whole process is: S out = BF(S f +S in )+(S f +S in ) (4) Where: BF is batch normalization BN+feedforward network FFN; 2) Aggregating the multi-scale semantic feature map, comprising the following steps: 2.1 The multi-scale semantic feature maps S2, S3, S4 extracted in step 1) are input and processed by a dimension reduction module; The dimension reduction module is composed of a convolution layer, a normalization layer and a Swish activation function, S3 and S4 are upsampled to the same spatial scale size as S2 after dimension reduction by bilinear interpolation, obtaining feature maps S3' and S4', and the feature map obtained after dimension reduction of S2 is S2', the mathematical expressions of S2', S3', S4' are respectively: Where: CLS is Conv+LN+Swish; Upsampling(x) is an up-sampling operation on x; 2.2 S2', S3', S4' are spliced and then subjected to a dimension reduction module for channel dimension reduction, finally obtaining a more expressive aggregated feature map S5 with a size of H / 8×W / 8, and its mathematical expression is: S5=CLS(Concat(S′2,S′3,S′4) (6) Where: Concat is a splicing operation; 3) Generating an output image after semantic segmentation, comprising the following steps: 3.1 The feature map aggregated in step 2) is input into a multi-scale dynamic scanning module for deep feature extraction; The multi-scale dynamic scanning module is composed of a dimension reduction module, a 2D selective scanning block and a convolution layer, which uses a 3×3 neighborhood feature with a stride of 2 to aggregate, generating a derived feature map with a scale size of 1 / 2 of S5; A 5×5 neighborhood feature with a stride of 4 is used to aggregate, generating a derived feature map with a scale size of 1 / 4 of S5, and the regional aggregation context is obtained through the two derived feature maps; 3.2 Using lossless down-sampling operation, the scale size of S5 and the 1 / 2 size derived feature map is reduced to the same H / 32xW / 32 as the 1 / 4 size derived feature map; this operation rearranges the 4x4 non-overlapping blocks of S5 and the 2x2 blocks of the 1 / 2 size derived feature map to the channel dimension, respectively increasing the channel depth by 16 times and 4 times, while retaining the original scale information, and then reducing the channel number of the three feature maps to the same number through a dimension reduction module, ensuring that the context information at each scale has equal importance; The dimension-reduced feature maps are spliced along the channel dimension, input into the 2D selective scanning block to realize multi-scale context extraction in a single scan; the scanned feature maps are input into a 1x1 convolution layer for cross-scale context fusion; finally, the fused features are reduced to the input channel number through a dimension reduction module, and up-sampled to the H / 8xW / 8 scale size, finally obtaining a new feature map S5' that fuses multi-scale context; 3.3 S5' is added to S2', S3' and S4' respectively to enhance their multi-scale context information, and then S5' is spliced with the enhanced S2', S3' and S4' and the global average pooling result of S4', after consecutive two layers of multi-layer perception fusion, up-sampling to the original size of the input picture, outputting an accurate semantic segmentation result map.