Double-branch remote sensing image semantic segmentation method and system
By using a dual-branch remote sensing image semantic segmentation method combined with CNN and Mamba modules, the problem of inaccurate segmentation of remote sensing image segmentation models in areas with large grayscale changes is solved, and efficient frequency domain information introduction and spatial domain feature preservation are achieved, thereby improving segmentation accuracy and robustness.
Patent Information
- Application Number
- CN202510834427.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-09-23
AI Technical Summary
Existing remote sensing image segmentation models cannot accurately segment areas with large grayscale changes, making it difficult to effectively capture the features of land objects with varying scales and directions in remote sensing images. In addition, the computational complexity is high and cannot adapt to the efficient processing requirements of high-resolution remote sensing images.
A dual-branch remote sensing image semantic segmentation method is adopted, combined with CNN and Mamba modules. Local features are extracted through the main branch, and global features are extracted through the auxiliary branch. Frequency domain information is introduced through wavelet transform, and semantic feature fusion is realized using the cross wavelet transform Mamba module. An adaptive convolution module is designed to improve the model's expression ability.
The segmentation accuracy of the remote sensing image segmentation model in areas with large grayscale changes is improved, the segmentation precision and robustness are improved, the computational complexity is reduced, and the areas with large grayscale changes in the image can be considered more comprehensively.
Smart Images

Figure CN120689624A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing image segmentation, and in particular relates to a dual-branch remote sensing image semantic segmentation method and system. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] Semantic segmentation of remote sensing images involves accurately classifying each pixel in a remote sensing image, identifying its semantic category (such as buildings, vegetation, and water bodies), and thus generating a pixel-by-pixel feature distribution map. Unlike traditional remote sensing image classification (which uses the entire image or a region as a unit), semantic segmentation provides more detailed feature boundary information and spatial distribution characteristics. This technology has important applications in disaster response (such as wildfire spread monitoring and flood zone identification), urban planning (land use change analysis), and agricultural yield estimation (fine-grained classification of crop types). It is one of the core technologies for automated and intelligent geographic information processing.
[0004] Among the current mainstream remote sensing image segmentation models, CNNs achieve feature extraction through local receptive fields and spatial invariance. However, their inherent inductive bias limits their ability to model global context, making it difficult to effectively capture the features of objects with varying scales and orientations in remote sensing images. Transformers utilize a self-attention mechanism to model long-range dependencies, but their computational complexity scales quadratically with the input image size, making them inadequate for the efficient processing of high-resolution remote sensing images.
[0005] In the field of remote sensing image segmentation, CNN and Transformer fusion networks can effectively utilize global and local information to a certain extent. However, these methods rely solely on spatial segmentation, without utilizing frequency domain information, and are relatively complex. In remote sensing images with complex backgrounds, grayscale variations pose a challenge to spatial segmentation. In particular, areas with large grayscale variations, such as edges and shadows, can lead to segmentation errors. Frequency domain features are more sensitive to these areas, resulting in inaccurate segmentation in areas with large grayscale variations. Summary of the Invention
[0006] In order to solve the above problems, the present invention proposes a dual-branch remote sensing image semantic segmentation method and system. The present invention introduces a combination of the Mamba module and CNN. By using the CNN-based main branch to extract local features and the Mamba auxiliary branch to extract global features, and at the same time introducing frequency domain information into the network model through wavelet transform, the problem of inaccurate segmentation in areas with large grayscale changes is effectively solved. In this way, the method can effectively introduce frequency domain information and maintain spatial domain features within an acceptable computational complexity to solve the problem of spatial segmentation and improve the accuracy of the segmentation model.
[0007] According to some embodiments, a first solution of the present invention provides a dual-branch remote sensing image semantic segmentation method, which adopts the following technical solutions: A dual-branch remote sensing image semantic segmentation method, comprising: Acquire the original remote sensing image and perform preprocessing to obtain a preprocessed remote sensing image; Based on the preprocessed remote sensing image, semantic segmentation is performed using the pre-trained dual-branch remote sensing image semantic segmentation model to obtain the segmentation results, specifically: The main branch is used to perform four convolution operations on the preprocessed remote sensing image in sequence to obtain the main coding features of different scales; Based on the preprocessed remote sensing image, the auxiliary branch is used to extract global features four times in sequence, and then the four-level global features are modulated in the spatial-channel dual dimension to obtain auxiliary coding features of different scales. The auxiliary coding features of different scales are fused with the main coding features of the corresponding scales to obtain the corresponding fused coding features of different scales; The fused coding features of different scales are decoded to obtain the segmentation results.
[0008] Furthermore, the main branch includes four residual blocks and four cross wavelet transform Mamba modules, and each residual block is connected to a cross wavelet transform Mamba module; The auxiliary branch includes four auxiliary branch encoder network modules and four spatial multi-scale feature adaptive modulation modules, and each auxiliary branch encoder network module is connected to a spatial multi-scale feature adaptive modulation module.
[0009] Furthermore, the cross wavelet transform Mamba module includes a parallel global branch and a local branch; The global branch is used to receive the auxiliary coding features from the auxiliary encoder and obtain four global components through discrete wavelet transform operation; The local branch is used to receive the main coding features from the main encoder and obtain four local components through discrete wavelet transform operation; The low-frequency global component and the corresponding low-frequency local component are merged and then the data is rearranged to convert into a feature sequence. The feature sequence is converted back to a two-dimensional form after layer normalization, linear layer and Mamba block processing to obtain the low-frequency merged component; The three high-frequency global components and the corresponding three high-frequency local components are merged in pairs and then adaptively convolved to obtain three high-frequency merged components; The three high-frequency merged components and the low-frequency merged components are concatenated and then subjected to inverse discrete wavelet transform. The transformed features are adaptively convolved and residually connected with the main encoding features of the main encoder to obtain the output features of the cross wavelet transform Mamba module.
[0010] Furthermore, the spatial multi-scale feature adaptive modulation module includes a spatial interaction branch and a channel interaction branch; The input features are expanded by convolution to increase the number of channels. In the expanded channel dimension, the convolved input features are equally divided into spatial interaction branch input features and channel interaction branch input features. The input features of the spatial interaction branch are processed by the spatial interaction branch to obtain the spatial modulation interaction features, and the input features of the channel interaction branch are processed by the channel interaction branch to obtain the channel interaction local features; The spatial modulation interaction features and channel interaction local features are spliced and then convolved to obtain the output of the spatial multi-scale feature adaptive modulation module.
[0011] Furthermore, the spatial interaction branch adaptively performs maximum pooling on the input features of the spatial interaction branch and then performs depth-wise separable convolution to obtain non-local features, and then combines the non-local features with the variance of the spatial interaction branch input features and modulates them to obtain modulated features. The modulated features are activated and then multiplied element-by-element with the spatial interaction branch input features to obtain spatial modulated features. The channel interaction branch performs deep convolution and convolution processing on the local information of the channel interaction branch input feature to obtain convolution local features; the convolution local features generate channel interaction local features through activation function and convolution processing.
[0012] Furthermore, the processing process of the auxiliary branch encoder network module is specifically as follows: The input features are first normalized to obtain normalized features, and then divided into parallel main branches and bypass branches; The main branch first processes the normalized features through a linear layer, depthwise separable convolution, and activation function to obtain the main convolution features. The main path convolution features are processed by the 2D selective scanning module to obtain the scan merge features; The scanned merged features are normalized at the layer and then multiplied element-wise with the bypass branch output features to obtain the main-side fusion features. The main-side fusion features are processed by the linear layer and then concatenated with the input features to obtain the final output features. The bypass branch performs linear layer and activation function processing on the normalized features to obtain the bypass branch output features.
[0013] According to some embodiments, a second solution of the present invention provides a dual-branch remote sensing image semantic segmentation system, which adopts the following technical solutions: A dual-branch remote sensing image semantic segmentation system, comprising: The remote sensing image processing module is configured to obtain the original remote sensing image and perform preprocessing to obtain a preprocessed remote sensing image; The semantic segmentation module is configured to perform semantic segmentation based on the pre-processed remote sensing image using a pre-trained dual-branch remote sensing image semantic segmentation model to obtain the segmentation results, specifically: The main branch is used to perform four convolution operations on the preprocessed remote sensing image in sequence to obtain the main coding features of different scales; Based on the preprocessed remote sensing image, the auxiliary branch is used to extract global features four times in sequence, and then the four-level global features are modulated in the spatial-channel dual dimension to obtain auxiliary coding features of different scales. The auxiliary coding features of different scales are fused with the main coding features of the corresponding scales to obtain the corresponding fused coding features of different scales; The fused coding features of different scales are decoded to obtain the segmentation results.
[0014] According to some embodiments, a third aspect of the present invention provides a computer-readable storage medium.
[0015] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a dual-branch remote sensing image semantic segmentation method as described in the first solution.
[0016] According to some embodiments, a fourth aspect of the present invention provides a computer device.
[0017] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the dual-branch remote sensing image semantic segmentation method described in the first solution are implemented.
[0018] According to some embodiments, a fifth aspect of the present invention provides a computer program product or computer program.
[0019] A computer program product or computer program includes computer instructions, which are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the steps in a dual-branch remote sensing image semantic segmentation method as described in the first scheme above.
[0020] Compared with the prior art, the present invention has the following beneficial effects: The present invention constructs a dual-branch encoder through CNN and Mamba to extract features from the local and global levels respectively. The network realizes the effective fusion of two semantic features through the Cross Wavelet Transform Mamba (CWM) module, and uses Haar wavelet transform to perform frequency domain feature mapping to introduce frequency domain information, so as to effectively improve the segmentation accuracy of areas with large grayscale changes such as edges and shadows. At the same time, an adaptive convolution module (AC) is designed to use multiple convolution kernels and dynamically select weighted combinations according to the input data, rather than fixed convolution kernels, to improve the model's expressiveness and performance.
[0021] The present invention can fully utilize the advantages of CNN and Mamba in extracting local and global information. Compared with a single spatial segmentation network or a frequency domain segmentation network, it can not only effectively introduce frequency domain information, but also maintain spatial domain features, so that the segmentation model can more comprehensively consider areas with large grayscale changes in the image, such as shadows, edges and areas with large texture changes, thereby improving segmentation accuracy and robustness. In addition, the adaptive convolution module designed by the present invention effectively solves the problem that the weights of traditional convolution are fixed and cannot change adaptively according to the input. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0023] Figure 1 This is an overall training flow chart of the dual-branch remote sensing image semantic segmentation method in an embodiment of the present invention; Figure 2 This is a diagram of the overall model framework of the dual-branch remote sensing image semantic segmentation model in an embodiment of the present invention; Figure 3 Schematic diagram of the auxiliary branch encoder network structure (VSS) in an embodiment of the present invention; Figure 4 Schematic diagram of a spatial multi-scale feature adaptive modulation module (SMFA) in an embodiment of the present invention; Figure 5 Schematic diagram of the structure of the Cross Wavelet Transform Mamba module (CWM) in an embodiment of the present invention; Figure 6 Schematic diagram of the adaptive convolution module structure in the CWM module in an embodiment of the present invention (AC); Figure 7 : is a line graph of evaluation indicators of the dual-branch remote sensing image semantic segmentation method according to an embodiment of the present invention; Figure 8 Graph showing the segmentation results of the dual-branch remote sensing image semantic segmentation method according to an embodiment of the present invention; Figure 9 This is a comparative experimental data table of the method in the embodiment of the present invention and other advanced semantic segmentation methods. DETAILED DESCRIPTION
[0024] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0025] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0026] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0027] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.
[0028] Example 1 This embodiment provides a dual-branch remote sensing image semantic segmentation method. This embodiment uses the method applied to a server as an example for illustration. It is understandable that the method can also be applied to a terminal, and can also be applied to a terminal, a server, and a system, and implemented through the interaction between the terminal and the server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network servers, cloud communications, middleware services, domain name services, security services CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in this application. In this embodiment, the method includes the following steps: Acquire the original remote sensing image and perform preprocessing to obtain a preprocessed remote sensing image; Based on the preprocessed remote sensing image, semantic segmentation is performed using the pre-trained dual-branch remote sensing image semantic segmentation model to obtain the segmentation results, specifically: The main branch is used to perform four convolution operations on the preprocessed remote sensing image in sequence to obtain the main coding features of different scales; Based on the preprocessed remote sensing image, the auxiliary branch is used to extract global features four times in sequence, and then the four-level global features are modulated in the spatial-channel dual dimension to obtain auxiliary coding features of different scales. The auxiliary coding features of different scales are fused with the main coding features of the corresponding scales to obtain the corresponding fused coding features of different scales; The fused coding features of different scales are decoded to obtain the segmentation results.
[0029] like Figure 1 As shown, the training process of the method described in this embodiment includes: Step S1: Dataset preprocessing: preprocess the high-resolution remote sensing images in the dataset, specifically: The ISPRS Vaihingen dataset is composed of 33 high-resolution remote sensing images with an average size of 2500 × 2000 pixels. These images are composed of three channels: near-infrared, red, and green, with a ground sampling distance of 9 cm. The dataset contains five foreground categories: ground, buildings, low vegetation, trees, and cars, and one background category (clutter).
[0030] The dataset is then cropped with a step size of 256 pixels to obtain images of 256×256 pixels each, thus obtaining the preprocessed remote sensing image dataset.
[0031] Step S2: Divide the preprocessed data set into training set, validation set and test set; Sixteen images from the preprocessed remote sensing image dataset are used for training, and the remaining 17 images are used for validation and testing.
[0032] Step S3: Construct a dual-branch remote sensing image semantic segmentation model based on CNN and Mamba, specifically: First, use represents the input preprocessed remote sensing image, represents the number of image channels, and and Represents the height and width of the image respectively.
[0033] like Figure 2 As shown in the figure, the dual-branch remote sensing image semantic segmentation model includes a dual-branch encoder and decoder. In the encoding stage, image features are extracted through two parallel branches: the main encoder and the auxiliary encoder. Each branch includes four corresponding module groups.
[0034] The main encoder uses a pre-trained ResNet18 network to learn multi-scale local feature representations. It contains four residual blocks (ResBlock) and four CWM modules (Cross Wavelet Transform Mamba Modules). The four residual blocks perform convolution operations and generate The multi-scale main coding features represented by ; wherein, each residual block is connected to a CWM module to form a main module group.
[0035] The auxiliary encoder is based on the VSS module that models long-range dependencies (such as Figure 3 The VSS module (auxiliary branch encoder network module) is used to extract the global features of the four levels, and then the SMFA module (spatial multi-scale feature adaptive modulation module) is used to perform spatial-channel dual-dimensional collaborative modulation on the global features of the four levels to obtain multi-scale auxiliary coding features. ; Among them, each VSS module is connected to an SMFA module to form an auxiliary module group, and each auxiliary module group corresponds to a main module group.
[0036] The auxiliary coding features generated by the auxiliary branch are then fed into the CWM of the corresponding scale in the main branch, and fused with the main coding features to obtain the fused coding features. , and finally sent to the decoder for the corresponding scale decoding block to make the final segmentation prediction.
[0037] The ResNet18 network described in step S3 first extracts image features through a 7×7 convolutional layer (with a stride of 2), followed by a 3×3 maximum pooling layer (with a stride of 2) for downsampling, and then passes through four stages of residual modules (Res Block). Each residual block includes two convolutional layers and a residual connection, which solves the problems of gradient vanishing and gradient exploding in deep network training.
[0038] like Figure 3 As shown, the processing process of the VSS module described in step S3 is specifically as follows: The input features are first normalized to obtain normalized features, and then divided into parallel main branches and bypass branches; The main branch first processes the normalized features through a linear layer, depthwise separable convolution, and activation function to obtain the main convolution features. The main path convolution features are processed by the 2D selective scanning module to obtain the scan merge features; The scanned merged features are normalized at the layer and then multiplied element-wise with the bypass branch output features to obtain the main-side fusion features. The main-side fusion features are processed by the linear layer and then concatenated with the input features to obtain the final output features. The bypass branch performs linear layer and activation function processing on the normalized features to obtain the bypass branch output features.
[0039] Specifically, for the input features , firstly, the normalized features are obtained through layer normalization (Layer Norm), and then divided into two branches. In the main branch, the normalized features are further processed by linear layer (Linear), and then a 3×3 depthwise separable convolution (DWConv) is performed, and the main convolution features are obtained after the SiLU activation function. : (1); in, Representation layer normalization, Represents a 3×3 depthwise separable convolution.
[0040] Next, the core computing unit of the VSS module, the 2D Selective Scan (SS2D) module, scans the main convolution features in four different directions. , decodes semantic information under global receptive field and linear complexity, and creates four independent sequences, which are then processed separately by SSM (State Space Model). Finally, the generated features are combined to generate a complete 2D feature map: (2); (3); (4); in, Indicates four different scanning directions, Indicates scanning extended features, represents sequence modeling features, Indicates a scan merge feature. and The S6 operation represents the sequential modeling of the selective scan space state (SSM) in Mamba, which promotes the interaction between each element in the 1D array and any previous scan sample.
[0041] Next, scan and merge features After layer normalization again and element-by-element multiplication with the bypass branch output feature, the main-side fusion feature is obtained. The main-side fusion feature is processed by a linear layer with residual connection and then combined with the input feature. The final output features are obtained by concatenation; for the bypass branch, the normalized features undergo the same linear embedding and SiLU activation to obtain the bypass branch output features.
[0042] To address the challenges of adaptively fusing multi-scale features in remote sensing image semantic segmentation, this implementation proposes a Spatial Multi-scale Feature Adaptation (SMFA) module. This module is embedded in the multi-scale feature extraction stage of the VMamba encoder in a cascaded manner. By co-modulating the feature maps at four levels in both spatial and channel dimensions, it significantly improves the model's ability to represent complex remote sensing features. The core innovations of SMFA are: (1) Dual-path dynamic feature interaction: The features are split into a spatial statistical modeling path (X-path) and a channel interaction path (Y-path) through a channel grouping strategy. The former constructs multi-scale spatial statistics through an adaptive pooling layer, and the latter uses a deep separable dynamic MLP (DMlp) to enhance channel correlation. (2) Spatial statistical perception modulation: innovatively integrates local maximum response (achieved through 3×3 deep convolution) with global variance features, using learnable parameters α 、 β Adaptive weighting is performed to construct a spatial attention map; this module dynamically perceives spatial context information at different scales, significantly improving the model's ability to segment multi-scale targets (such as building edges, road textures, etc.) in remote sensing images while maintaining computational efficiency.
[0043] like Figure 4 As shown in step S3, the SMFA module includes two branches, the spatial interaction branch and the channel interaction branch. First, the input features are expanded by 1×1 convolution to increase the number of channels, and then the channels are divided into two parts as the input features of the spatial interaction branch. Interacting with channel branch input features .
[0044] The input features are expanded by convolution to increase the number of channels. In the expanded channel dimension, the convolved input features are equally divided into spatial interaction branch input features and channel interaction branch input features. The input features of the spatial interaction branch are processed by the spatial interaction branch to obtain the spatial modulation interaction features, and the input features of the channel interaction branch are processed by the channel interaction branch to obtain the channel interaction local features; The spatial modulation interaction features and channel interaction local features are spliced and then convolved to obtain the output of the spatial multi-scale feature adaptive modulation module.
[0045] The spatial interaction branch performs adaptive maximum pooling on the input features of the spatial interaction branch and then performs depth-wise separable convolution to obtain non-local features. The non-local features are merged with the variance of the spatial interaction branch input features and then modulated to obtain modulated features. The modulated features are activated and then multiplied element-by-element with the spatial interaction branch input features to obtain spatial modulated features.
[0046] Specifically, in the spatial interaction branch, the spatial interaction branch input features First, adaptive maximum pooling is performed and non-local features are obtained through 3×3 depth convolution. To embed the modulation space interaction branch input features A global description of Variance As the statistical divergence of spatial information, and through 1 × 1 convolution and Merge. The specific calculation formula is as follows: (5); (6); (7); in, represents adaptive maximum pooling with a scaling factor of 8, represents a 3×3 depth convolutional layer, Represents the input features of the spatial interaction branch The variance of is the total number of pixels, is the average value of all pixels, is the modulation characteristic.
[0047] Final modulation characteristics After the GELU activation function, the input features are branched with spatial interaction Perform element-by-element multiplication operations to obtain spatially modulated features.
[0048] The channel interaction branch performs deep convolution and convolution processing on the local information of the input features of the channel interaction branch to obtain convolution local features; the convolution local features generate channel interaction local features through activation function and convolution processing.
[0049] Specifically, in the channel interaction branch, the input feature First, the local information is encoded through a 3×3 depth convolution, and then a 1×1 convolution is used to obtain the convolution local features. ; Convolution local features Generate channel interaction local features through GELU activation function and 1×1 convolution , the specific calculation formula is as follows: (8); (9); in, represents the GELU activation function, Represents a 1×1 convolutional layer.
[0050] like Figure 5 As shown, the CWM (Cross Wavelet Mamba) module described in step S3 includes two parallel branches: a global branch and a local branch. The former is used to receive multi-scale auxiliary coding features from the auxiliary encoder. , which receives the multi-scale main encoding features from the main encoder .
[0051] The cross wavelet transform Mamba module includes parallel global branches and local branches; The global branch is used to receive the auxiliary coding features from the auxiliary encoder and obtain four global components through discrete wavelet transform operation; The local branch is used to receive the main coding features from the main encoder and obtain four local components through discrete wavelet transform operation; The low-frequency global component and the corresponding low-frequency local component are merged and then the data is rearranged to convert into a feature sequence. The feature sequence is converted back to a two-dimensional form after layer normalization, linear layer and Mamba block processing to obtain the low-frequency merged component; The three high-frequency global components and the corresponding three high-frequency local components are merged in pairs and then adaptively convolved to obtain three high-frequency merged components; The three high-frequency merged components and the low-frequency merged components are concatenated and then subjected to inverse discrete wavelet transform, and adaptive convolution is performed on the transformed features.
[0052] Specifically, the two branches are subjected to discrete wavelet transform (DWT) operations, and each obtains four global components ( , , , ) and four local components ( , , , ). LL is the low-frequency component, which mainly retains smooth information and is essentially a low-resolution version of the input image. LH, HL, and HH are high-frequency components, which contain details such as edges and textures in different directions. , The merged tensor is first converted into a serialized input through the rearrange operation, and then the feature sequence is processed through layer normalization, linear layer and Mamba block to learn long-range dependency features. The processed sequence is then converted back to 2D form to obtain the low-frequency merge component. Other high frequency components are also merged (splicing , , and , , After merging, the high-frequency merged component is obtained through AC module (adaptive convolution). 、 、 Combine these three components with the low-frequency component The splicing is performed, followed by an inverse discrete wavelet transform (IWT), and then passed through an adaptive convolution (AC) module and a residual connection is made with the main coding features of the main encoder to obtain the output features of the CWM module, and finally sent to the decoder for final segmentation prediction.
[0053] like Figure 6 As shown, the AC (adaptive convolution) module adaptive convolution has convolution kernels that share the same kernel size and input / output dimensions. They are connected by using attention weights Aggregation is performed. Batch normalization and activation functions (such as ReLU) are used after the aggregate convolution to construct the adaptive convolution layer. Compression and excitation mechanism is applied to calculate the attention weight of the convolution kernel. . Global space information First, it is compressed by global average pooling. Then, it is generated by two fully connected layers (using ReLU and Softmax activation functions respectively). Normalized attention weights of convolution kernels. Unlike SENet, which calculates the attention of output channels, this embodiment calculates the attention weights of convolution kernels.
[0054] The decoder described in step S3 uses a plug-and-play feature pyramid network head (FPNHead) to fuse feature maps of different scales and ultimately generate a segmentation result.
[0055] Step S4: Use the training set and validation set in step S2 to train, validate, and optimize the network model constructed in step S3, and save the best model weights obtained, specifically: Step S41: Use loss function To supervise the model, it is the cross entropy loss function Combined with the dice loss function, the expression is: (10); (11); (12); in, represents the number of samples, represents the number of categories, and Respectively represent The one-hot encoding of the true semantic label of each sample and the corresponding softmax result of the network output, , Representation sample Belong to Confidence of the class.
[0056] Step S42: Set model training parameters. The number of training epochs is set to 100, and the batch size is set to 4. To accelerate model convergence, the AdamW optimizer is used, with an initial learning rate of 0.0001, a weight decay of 0.01, and a cosine learning rate strategy. To prevent overfitting during training, the training data is cropped to a size of 256 × 256 using a sliding window operation with a step size of 128. Data augmentation is then performed using random horizontal flips, random vertical flips, and random multi-scale cropping of [0.5, 0.75, 1.0, 1.25, 1.5].
[0057] Step S43: Figure 7 As shown in the figure, the mean intersection over union (mIoU) and mean F1 score (mF1) are used as the main evaluation indicators to evaluate the model segmentation performance. These two indicators are commonly used in the semantic segmentation of remote sensing images. The calculation formulas of mIoU and mF1 are: (13); (14); in, Indicates the categories, Indicates that the sample that is actually positive is correctly predicted as positive; The samples that are actually negative are mistakenly predicted as positive; The samples that are actually negative are correctly predicted as negative; is the accuracy, Indicates how many of the samples predicted by the model as positive are actually positive; is the recall rate, Indicates how many positive samples the model correctly identifies among all samples that are actually positive.
[0058] Step S5: Load the model weights saved in step S4, input the test set in S2 into the network model constructed in step S3 to segment the images in the test set, such as Figure 8 The following is the image segmentation result.
[0059] like Figure 9 As shown in the figure, compared with the spatial network model that does not introduce frequency domain information, the model proposed in the present invention combines spatial information with frequency domain information, so that the model has a significant improvement in both mIoU and mF1 indicators compared with the comparison model.
[0060] Example 2 This embodiment provides a dual-branch remote sensing image semantic segmentation system, including: The remote sensing image processing module is configured to obtain the original remote sensing image and perform preprocessing to obtain a preprocessed remote sensing image; The semantic segmentation module is configured to perform semantic segmentation based on the pre-processed remote sensing image using a pre-trained dual-branch remote sensing image semantic segmentation model to obtain the segmentation results, specifically: The main branch is used to perform four convolution operations on the preprocessed remote sensing image in sequence to obtain the main coding features of different scales; Based on the preprocessed remote sensing image, the auxiliary branch is used to extract global features four times in sequence, and then the four-level global features are modulated in the spatial-channel dual dimension to obtain auxiliary coding features of different scales. The auxiliary coding features of different scales are fused with the main coding features of the corresponding scales to obtain the corresponding fused coding features of different scales; The fused coding features of different scales are decoded to obtain the segmentation results.
[0061] The examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the contents disclosed in the above embodiment 1. It should be noted that the above modules as part of the system can be executed in a computer system such as a set of computer executable instructions.
[0062] The descriptions of the various embodiments in the above embodiments have different focuses. For parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0063] The proposed system can be implemented in other ways. For example, the system embodiment described above is merely illustrative. For example, the above module division is only a logical function division. In actual implementation, other division methods may be used. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not implemented.
[0064] Example 3 This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the steps of the dual-branch remote sensing image semantic segmentation method described in the first embodiment are implemented.
[0065] Example 4 This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the dual-branch remote sensing image semantic segmentation method described in the first embodiment are implemented.
[0066] Example 5 This embodiment provides a computer program product or computer program, including computer instructions, which are stored in a computer-readable storage medium. The processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the steps in the dual-branch remote sensing image semantic segmentation method described in the above embodiment one.
[0067] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage and optical storage, etc.) containing computer-usable program code.
[0068] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0069] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1The function specified in one or more boxes.
[0070] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0071] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0072] Although the above describes the specific embodiments of the present invention in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without any creative work are still within the scope of protection of the present invention.
Claims
1. A dual-branch remote sensing image semantic segmentation method, characterized in that: include: Acquire the original remote sensing image and perform preprocessing to obtain a preprocessed remote sensing image; Based on the preprocessed remote sensing image, semantic segmentation is performed using the pre-trained dual-branch remote sensing image semantic segmentation model to obtain the segmentation results, specifically: The main branch is used to perform four convolution operations on the preprocessed remote sensing image in sequence to obtain the main coding features of different scales; Based on the preprocessed remote sensing image, the auxiliary branch is used to extract global features four times in sequence, and then the four-level global features are modulated in the spatial-channel dual dimension to obtain auxiliary coding features of different scales. The auxiliary coding features of different scales are fused with the main coding features of the corresponding scales to obtain the corresponding fused coding features of different scales; The fused coding features of different scales are decoded to obtain the segmentation results.
2. A dual-branch remote sensing image semantic segmentation method according to claim 1, characterized in that: The main branch includes four residual blocks and four cross wavelet transform Mamba modules, and each residual block is connected to a cross wavelet transform Mamba module; The auxiliary branch includes four auxiliary branch encoder network modules and four spatial multi-scale feature adaptive modulation modules, and each auxiliary branch encoder network module is connected to a spatial multi-scale feature adaptive modulation module.
3. A dual-branch remote sensing image semantic segmentation method according to claim 2, characterized in that: The cross wavelet transform Mamba module includes a parallel global branch and a local branch; The global branch is used to receive the auxiliary coding features from the auxiliary encoder and obtain four global components through discrete wavelet transform operation; The local branch is used to receive the main coding features from the main encoder and obtain four local components through discrete wavelet transform operation; The low-frequency global component and the corresponding low-frequency local component are merged and then the data is rearranged to convert into a feature sequence. The feature sequence is converted back to a two-dimensional form after layer normalization, linear layer and Mamba block processing to obtain the low-frequency merged component; The three high-frequency global components and the corresponding three high-frequency local components are merged in pairs and then adaptively convolved to obtain three high-frequency merged components; The three high-frequency merged components and the low-frequency merged components are concatenated and then subjected to inverse discrete wavelet transform. The transformed features are adaptively convolved and residually connected with the main encoding features of the main encoder to obtain the output features of the cross wavelet transform Mamba module.
4. A dual-branch remote sensing image semantic segmentation method according to claim 2, characterized in that: The spatial multi-scale feature adaptive modulation module includes a spatial interaction branch and a channel interaction branch; The input features are expanded by convolution to increase the number of channels. In the expanded channel dimension, the convolved input features are equally divided into spatial interaction branch input features and channel interaction branch input features. The input features of the spatial interaction branch are processed by the spatial interaction branch to obtain the spatial modulation interaction features, and the input features of the channel interaction branch are processed by the channel interaction branch to obtain the channel interaction local features; The spatial modulation interaction features and channel interaction local features are spliced and then convolved to obtain the output of the spatial multi-scale feature adaptive modulation module.
5. A dual-branch remote sensing image semantic segmentation method according to claim 4, characterized in that: The spatial interaction branch performs adaptive maximum pooling on the input features of the spatial interaction branch and then performs depth-wise separable convolution to obtain non-local features, and then combines the non-local features with the variance of the spatial interaction branch input features and modulates them to obtain modulated features. The modulated features are activated and then multiplied element-by-element with the spatial interaction branch input features to obtain spatial modulated features; The channel interaction branch performs deep convolution and convolution processing on the local information of the channel interaction branch input feature to obtain convolution local features; the convolution local features generate channel interaction local features through activation function and convolution processing.
6. A dual-branch remote sensing image semantic segmentation method according to claim 2, characterized in that: The processing process of the auxiliary branch encoder network module is specifically as follows: The input features are first normalized to obtain normalized features, and then divided into parallel main branches and bypass branches; The main branch first processes the normalized features through a linear layer, depthwise separable convolution, and activation function to obtain the main convolution features. The main path convolution features are processed by the 2D selective scanning module to obtain the scan merge features; The scanned merged features are normalized at the layer and then multiplied element-wise with the bypass branch output features to obtain the main-side fusion features. The main-side fusion features are processed by the linear layer and then concatenated with the input features to obtain the final output features. The bypass branch performs linear layer and activation function processing on the normalized features to obtain the bypass branch output features.
7. A dual-branch remote sensing image semantic segmentation system, characterized in that: include: The remote sensing image processing module is configured to obtain the original remote sensing image and perform preprocessing to obtain a preprocessed remote sensing image; The semantic segmentation module is configured to perform semantic segmentation based on the pre-processed remote sensing image using a pre-trained dual-branch remote sensing image semantic segmentation model to obtain the segmentation results, specifically: The main branch is used to perform four convolution operations on the preprocessed remote sensing image in sequence to obtain the main coding features of different scales; Based on the preprocessed remote sensing image, the auxiliary branch is used to extract global features four times in sequence, and then the four-level global features are modulated in the spatial-channel dual dimension to obtain auxiliary coding features of different scales. The auxiliary coding features of different scales are fused with the main coding features of the corresponding scales to obtain the corresponding fused coding features of different scales; The fused coding features of different scales are decoded to obtain the segmentation results.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of a dual-branch remote sensing image semantic segmentation method according to any one of claims 1 to 6 are implemented.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the dual-branch remote sensing image semantic segmentation method according to any one of claims 1 to 6 are implemented.
10. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, the computer program implements the steps in the dual-branch remote sensing image semantic segmentation method according to any one of claims 1 to 6.
Citation Information
Cited By
Pericardial adipose tissue segmentation method, device and equipment based on MR image and medium
CN120997236A
Method and device for segmenting pericardial adipose tissue based on MR images, equipment and medium
CN120997236B
Grape fruit thinning period fruit grain segmentation method based on combination of wavelet transform and visual state space model
CN121767732A
A Grape Berry Segmentation Method Based on Wavelet Transform Combined with Visual State Space Model during Fruit Thinning
CN121767732B