CNN and Mama-based remote sensing image semantic segmentation method and system
By combining the hybrid neural network architecture of CNN encoder and Mamba decoder, the problem of high global information capture and computing complexity of remote sensing image semantic segmentation model is solved, and the remote sensing image segmentation with higher accuracy and efficiency is achieved.
Patent Information
- Application Number
- CN202510332157.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-25
AI Technical Summary
The existing remote sensing image semantic segmentation model has insufficient global information capture capabilities, high computational complexity and poor scale adaptability, making it difficult to adapt to large-scale remote sensing image scenarios with significant target changes.
A hybrid neural network architecture is constructed using CNN-based encoder and Mamba decoder. Through the ResASPP module and the CSMamba module, multi-scale feature extraction and global semantic modeling are realized, local features are extracted using residual networks and long-distance connections are established through the Mamba decoder to enhance global semantic associations.
It improves the accuracy and computing efficiency of the semantic segmentation model of remote sensing images, can better capture global context information, adapt to different target scales, and reduces the computational complexity.
Smart Images

Figure CN120375039A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for semantic segmentation of remote sensing images, and more particularly to a method and system for semantic segmentation of remote sensing images based on CNN and Mamba. Background Art
[0002] Semantic segmentation of remote sensing images divides large-scale pixels in remote sensing images into different categories to enhance the analysis and interpretation of remote sensing data. This large-scale semantic segmentation is crucial for many practical applications such as autonomous driving, urban planning, and environmental protection. However, in the field of remote sensing, images often contain scenes with large scales and significant target variations. The UNet network based on convolutional neural networks is an encoder-decoder structure. The encoding part consists of multiple convolutional and max-pooling layers, gradually extracting high-level features while increasing the receptive field. The decoding part restores the image resolution through transposed convolution or upsampling steps, and uses skip connections to fuse the features of the corresponding layers of the encoder into the decoding process to retain detailed information. The number of channels is adjusted through 1×1 convolution to generate a semantic segmentation result with the same resolution as the input image. Although UNet performs well in semantic segmentation tasks, it has insufficient global information capture ability, high computational complexity, and poor scale adaptability. Therefore, it is crucial to develop a remote sensing image semantic segmentation architecture that can more effectively capture global information and thus adapt to changing targets. Summary of the Invention
[0003] Object of the Invention: Aiming at the above problems, the present invention proposes a method and system for semantic segmentation of remote sensing images based on CNN and Mamba, which can more effectively capture global information, thus adapting to changing targets and improving the accuracy of the remote sensing image semantic segmentation model.
[0004] Technical Solution: The technical solution adopted by the present invention is a method for semantic segmentation of remote sensing images based on CNN and Mamba, including: preprocessing the obtained remote sensing image, and then using a remote sensing image semantic segmentation model to perform semantic segmentation on image patches to obtain a pixel-level classification result;
[0005] The remote sensing image semantic segmentation model includes a CNN encoder, a residual connection layer, and a CSMamba decoder connected in sequence; the preprocessed image is input into the CNN encoder, and the CNN encoder includes several stacked ResBlock modules; the residual connection layer includes several multi-scale residual spatial pyramid pooling ResASPP modules. The residual connection layers of each level of ResBlock modules in the CNN encoder first undergo cascade processing to obtain combined features, and then the combined features are respectively connected to the corresponding levels of CSMamba modules in the CSMamba decoder through the ResASPP modules; the ResASPP module includes multiple convolutional attention units and an adaptive pooling layer. The convolutional kernels of the multiple convolutional attention units have the same size, and the dilation rate increases one by one. Each convolutional attention unit uses a residual connection network; a supervision module is added at each level of CSMamba module. The supervision module processes the intermediate feature information of the i-th CSMamba module to obtain an intermediate supervision signal, and the i-th CSMamba module adds the intermediate supervision signal to the corresponding elements of the intermediate feature information and then outputs.
[0006] The calculation formula for obtaining combined features through cascade processing is:
[0007]
[0008] In the formula, is the combined feature, Concat is the Concat function in SQL, F i , F i-1 , F i+1 are the features output by the residual connection layers of three adjacent ResBlock modules respectively.
[0009] The CSMamba decoder includes several stacked CSMamba modules. Each layer of CSMamba module obtains the output features of the previous layer through upsampling and combines the output features of the previous layer with the features output from the ResASPP module; the outputs of each level of CSMamba module are subjected to feature mapping through 1×1Conv to output semantic segmentation results of different resolutions.
[0010] The CSMamba decoder designs a dual-branch feature processing. The first branch obtains the long-range dependence of image features through a 2D selective scanning module, and the second branch obtains features that enhance important channels and spatial positions through a channel attention and spatial attention mechanism. The output results of the two branches are aggregated and output through a Hadamard product.
[0011] The 2D selective scanning module first flattens the 2D image features into a 1D sequence; then scans the sequence in four directions of the 2D image, and uses the selective state space model SSM to capture the long-term correlations in each direction as the scanning results; finally, combines the scanning results in the four directions and restores them to the 2D image structure.
[0012] The supervision module processes the intermediate output of the i-th CSMamba module through the following formula:
[0013]
[0014] where, P i represents the intermediate supervision signal output by the supervision module, F cs i is the feature of the i-th CSMamba block, and Conv represents the convolution calculation.
[0015] During the training process of the remote sensing image semantic segmentation model, the expression of the cross-entropy loss function is as follows:
[0016]
[0017] where, MFB_CE loss represents the cross-entropy loss function, n represents the samples in the dataset, c represents the semantic categories of the images, N represents the number of samples, C represents the number of categories, w c represents the weights of each category, l c represents the true label of sample n, and p c represents the probability that the model predicts that sample n belongs to category c.
[0018] The present invention provides a remote sensing image semantic segmentation system based on CNN and Mamba, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the above-mentioned remote sensing image semantic segmentation method based on CNN and Mamba.
[0019] The present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the above-mentioned remote sensing image semantic segmentation method based on CNN and Mamba.
[0020] The present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the above-mentioned remote sensing image semantic segmentation method based on CNN and Mamba.
[0021] The present invention provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the remote sensing image semantic segmentation method based on CNN and Mamba as described above.
[0022] Beneficial effects: Compared with the prior art, the present invention has the following advantages: The present invention adopts a hybrid neural network architecture that integrates multi-scale feature extraction and global semantic modeling, and is applicable to semantic segmentation scenarios with large scale and significant target changes. Aiming at the problems in the prior art that the Unet model has insufficient long-distance dependence modeling, poor adaptability to multi-scale targets, and high computational complexity of traditional Transformer structures, the present invention proposes an innovative encoding and decoding architecture, that is, an encoder based on CNN constructs a remote sensing image semantic segmentation model architecture in combination with a decoder based on CSMamba through a ResASPP module. By using an encoder based on CNN, local features of the image are extracted through a residual network, the features output by the residual connection layers of three adjacent ResBlock modules are cascaded to form a combined feature, and then through a ResidualAtrous-Spatial Pyramid Pooling module architecture combined with a Mamba architecture with dual-branch processing features, the local features can be better associated with the global semantics, and long-distance connections can be adaptively established between different feature layers, enabling the deep fusion of local texture information and global semantic information, thereby enhancing the detail expression ability of the target area and the boundary segmentation accuracy, and making up for the deficiency of ordinary image semantic segmentation models in long-distance information modeling. In addition, the present invention uses ResASPP as the feature extraction module of the CNN encoder, and can also perform feature sampling from multiple different scales, enhance the adaptability to different target scales, and effectively improve the ability to extract global context information. The architecture of the present invention can better aggregate and integrate global information than traditional semantic segmentation architectures, fully capture global context, improve the accuracy of the association between local features and global semantics, make the model's understanding of images more comprehensive and in-depth, and further improve the accuracy of the segmentation model. In addition, the linear complexity of the Mamba decoder does not face excessive computational burden when processing large-size remote sensing images, enabling the model to reduce the amount of calculation. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is a schematic structural diagram of the remote sensing image semantic segmentation model based on CNN and Mamba of the present invention;
[0024] Figure 2 is a schematic structural diagram of the 2D-SSM module of the present invention;
[0025] Figure 3 is a schematic structural diagram of the CSMamba decoder of the present invention;
[0026] Figure 4 It is a schematic structural diagram of the ResASPP module described in the present invention. Specific implementation manners
[0027] The technical solution of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0028] Embodiment 1
[0029] The remote sensing image semantic segmentation method based on CNN and Mamba described in the present invention includes: preprocessing the acquired remote sensing image, and then using a remote sensing image semantic segmentation model to perform semantic segmentation on image patches to obtain a pixel-level classification result. Among them, the remote sensing image semantic segmentation model includes a convolutional neural network (CNN) encoder, a multi-scale residual spatial pyramid pooling (ResASPP) module, and a decoder containing CSMamba blocks.
[0030] The structure of the remote sensing image semantic segmentation model based on CNN and Mamba described in the present invention is as Figure 1 shown. The preprocessed image is input into the CNN encoder. The structure of the CNN encoder adopts a residual network, including several stacked ResBlock modules for extracting local features of the image. Each residual connection layer of the previous several levels of ResBlock modules is respectively connected to the corresponding levels of CSMamba modules in the CSMamba decoder through an atrous spatial pyramid pooling module. The C in Skip Connection refers to a fusion operation, which is used to fuse the features in the encoder with the features in the decoder to better restore high-resolution information. The structure of the CSMamba decoder includes several stacked CSMamba modules. Each CSMamba module gradually restores the information from low resolution to high resolution through upsampling and feature transfer. Through upsampling, each level of CSMamba module receives the output of the previous layer and combines the ResASPP processed features from the residual connection (Skip Connection) to gradually enhance the feature representation. Each level of CSMamba module performs feature mapping through 1×1Conv and outputs semantic segmentation results of different resolutions to improve the overall segmentation performance.
[0031] The remote sensing image semantic segmentation method based on CNN and Mamba provided in this embodiment specifically includes 7 steps:
[0032] Step 1: Obtain a remote sensing image
[0033] As an implementation manner, the original data set is divided into a training set and a test set according to a ratio of 7:3. The training set is used to train the remote sensing image semantic segmentation model.
[0034] During the specific use process, the obtained remote sensing image or the remote sensing image in the test set can be directly input into the trained remote sensing image semantic segmentation model after preprocessing.
[0035] Step 2: Image preprocessing
[0036] Remote sensing images usually have large sizes. Limited by computing resources, remote sensing images are generally cropped before being fed into the feature extraction network.
[0037] (1) Use the sliding window method to crop the remote sensing image into multiple initial image patches. Among them, the size of the sliding window is larger than the moving step of the sliding window, so that two adjacent initial image patches have a common area;
[0038] As an implementation, the size of the sliding window is 896*896 pixels, and it slides 512 pixels each time, that is, the moving step of the sliding window is 512 pixels.
[0039] (2) Perform data augmentation operations of random horizontal and vertical flipping and random rotation by 90 degrees on the initial image patches to obtain the augmented image patches.
[0040] (3) Calculate the mean and standard deviation of the RGB three channels of all augmented image patches, and standardize the pixel values of the augmented image patches to obtain the image patches for inputting into the encoder of the remote sensing image semantic segmentation model.
[0041] Step 3: Encoding process
[0042] The present invention selects an encoder based on CNN, uses the residual network ResNet as the multi-scale feature extraction network to obtain multi-scale features. Among them, the residual network is divided into four stages for feature extraction, and each stage corresponds to residual features of different scales. For example Figure 1As shown in the figure, the CNN encoder includes: the first residual feature Resblok1. In this stage, relatively small convolutional kernels are mainly used to perform preliminary feature extraction on the image. Since the convolutional kernels are small, it can capture very fine detail information in the image. At the same time, through the residual connection, it is ensured that important information will not be lost in this stage. The second residual feature Resblok2. In this stage, the number of convolutional kernels will increase, and downsampling operations are performed. While retaining important features, it integrates the detail information extracted in the first stage and pays attention to local features at a slightly larger scale. The third residual feature Resblok3. The number of convolutional kernels will further increase, and downsampling operations continue. By continuously deepening the network and increasing the number of convolutional kernels, the model can learn more complex feature representations and maintain the stability and trainability of the model through residual connections. The fourth residual feature Resblok4. Through larger-scale convolutional operations and downsampling, the model extracts the global features at the largest scale in the image. These global features contain the overall semantic information of the image. By learning these global features, the model can understand the image from a macroscopic perspective.
[0043] Step 4, ResASPP module
[0044] In the present invention, a multi-scale residual spatial pyramid pooling module (ResASPP) is introduced on the decoder. ResASPP is based on the ASPP framework. As an implementation, as Figure 4 shown, the ResASPP module is used to refine the multi-scale features from the ResNet encoder in the semantic segmentation task of remote sensing images. Specifically, the output features F1, F2, F3 at different levels of the ResNet encoder are first concatenated and processed to form a combined feature Then the combined feature is sent to the ResASPP module for further refinement.
[0045] The ResASPP module consists of four extended convolutional attention units and an adaptive pooling layer. Among them, the dilation rates of the four dilated convolutional attention units are set to 3, 6, and 12 respectively, and the convolutional kernel size is uniformly 3×3. In the present invention, the input image features enter each dilated convolutional attention unit. Each unit first adjusts the weights of the input features according to the attention mechanism to highlight the features of the key regions, and then performs a convolutional operation. Subsequently, the output of each extended attention convolutional unit is added to the output of the corresponding residual unit to obtain four remote sensing image feature maps of different scales. These feature maps respectively reflect the details and structures of the ground objects in the remote sensing image from different scales. For example, the small-scale feature map can clearly present details such as small buildings and road signs, while the large-scale feature map can show the distribution of macroscopic ground objects such as large lakes and urban areas. The adaptive pooling layer processes the original input remote sensing image features, effectively retaining the global context information of the image, and helping the model understand the spatial position relationship and semantic association between different ground objects, such as identifying the relative position of farmland and surrounding irrigation facilities. The ResASPP module deeply optimizes the combined features from the ResNet encoder, combines local features and global context features of different scales, enables the model to more accurately segment various ground objects from complex and changing remote sensing images, and improves the efficiency of the semantic segmentation task.
[0046] Step 5, Decoding process
[0047] Based on the linear complexity advantage of the Mamba model in remote modeling, the present invention introduces a visual state space module into the field of remote sensing semantic segmentation. In specific implementation, as Figure 2 , a 2D selective scanning module (2D-SSM) is adopted. First, the 2D image features are flattened into a 1D sequence; then, the sequence is scanned in four directions (from top left to bottom right, from bottom right to top left, from top right to bottom left, and from bottom left to top right) of the 2D image; then, the selective state space model (SSM) is used to capture the long-term correlation in each direction; finally, the scanning results in the four directions are combined and restored to the 2D image structure. This method can effectively capture the long-range dependence relationship in the remote sensing image, and at the same time reduces the computational cost through the linear complexity of the Mamba model.
[0048] The present invention designs a dual-branch feature processing module in the decoding process, which processes the input features differently and finally generates an output through feature aggregation. As Figure 3 , after completing the relevant operations in the early stage, the input feature X ∈ R H×W×CPass through two parallel branches. In the first branch, the input features are first channel-expanded by a linear layer. Specifically, the number of channels of the input features is expanded from the original C to λC, where λ is a predefined channel expansion factor. Next, the expanded features undergo a depthwise convolution (DWConv) operation to extract spatial information while reducing the computational cost. Then, the SiLU (Sigmoid Linear Unit) activation function is applied to enhance the non-linear representation ability of the features. After that, the features are fed into a 2D selective scanning module (2D-SSM) to capture long-range dependencies through four-directional scanning, enhancing the modeling ability of global context information. Finally, layer normalization (LayerNorm) is used to normalize the features, stabilizing the training process and accelerating convergence.
[0049] In the second branch, the input features are first processed by a channel and spatial attention (CS) module. This module enhances the features of important channels and spatial positions through channel attention and spatial attention mechanisms, suppressing redundant information. Then, the SiLU activation function is applied to further enhance the non-linear representation ability of the features. Next, the output features of the first branch and the second branch are aggregated through the Hadamard product (element-wise product) to fuse the information of the two branches. Finally, a linear layer projects the number of channels of the aggregated features from the expanded λC back to the original number of channels C to generate an output Xout with the same shape as the input:
[0050] X1 = LN(2D-SSM(SiLU(DWConv(Linear(X)))))
[0051] X2 = SiLU(CS(X))
[0052] X out = Linear(X1 ⊙ X2)
[0053] where DWConv represents depthwise convolution, CS represents the channel and spatial attention module, 2D-SSM is the 2D selective scanning module, and D-S represents the Hadamard product. After this series of operations, the model can more accurately classify each pixel in the image during the decoding process, thereby achieving effective segmentation of different semantic contents in the image and achieving the purpose of extracting global context information.
[0054] To supervise the decoder to generate a semantic segmentation map of the remote sensing image, the CM-UNet architecture of the present invention incorporates intermediate supervision at each CSMamba decoder module. This makes each stage of the network contribute to the final segmentation result, ultimately achieving a more refined and accurate output. Specifically, the intermediate output of the i-th CSMamba block is processed by the following formula:
[0055]
[0056] Among them, F cs i is the feature of the i-th CSMamba block. Conv is a convolutional module used to convert the feature into a feature map P with the number of channels C i This feature map P i As an intermediate supervision signal, it will be fed back to the subsequent network layers of the current stage. During the processing of the subsequent network layers, the model will fuse the intermediate supervision signal with the original feature information. During the processing of the subsequent network layers, the model will directly add the corresponding elements of the intermediate supervision signal and the original feature information. Here, the original feature information also comes from the i-th CSMamba block and is the feature output by this block before passing through the Conv module. Through this fusion operation, the model can adjust the feature extraction and processing methods of the subsequent layers according to the information carried by the supervision signal, and finally achieve more accurate semantic segmentation. Through this fusion operation, the model can adjust the feature extraction and processing methods of the subsequent layers according to the information carried by the supervision signal, and finally achieve more accurate semantic segmentation.
[0057] Step 6, Model training and testing
[0058] The present invention uses the multi-scale fusion information as the final feature map for prediction. As an implementation, during the training process, the prediction result is compared with the image patch label, and the cross-entropy function is used as the loss function to calculate the loss value. Further, the expression of the cross-entropy loss function is as follows:
[0059]
[0060] Among them represents the weights of each category, and f c represents the pixel frequency of a certain category c, and media(f c ) represents finding the median of f c .
[0061] The present invention uses the SGD optimizer during the training process, with the momentum set to 0.9, the weight decay coefficient set to 0.0001, the initial learning rate set to 0.007, the learning rate gradually decreased through the polynomial decay strategy, and the batch size set to 4, that is, four image patches are read simultaneously each time for training, and a total of 60,000 iterations are performed. The loss function is calculated at each step and the gradient is backpropagated. By observing the change curve of the loss function, the model is selected as the final model after the loss function becomes stable.
[0062] In the testing stage, the image cropping of the present invention is consistent with the training process, that is, the test image is cropped into multiple image patches by means of a sliding window, where the size of the sliding window is 896*896 pixels and it slides 512 pixels each time. Given any test image I, during the testing process, the position information of each image patch relative to the image I is recorded, and then each image patch is fed into the trained remote sensing image semantic segmentation model. The model will output the prediction results of each pixel belonging to various classes in the form of probabilities. For the overlapping pixel points between two image patches, the present invention calculates the mean value of the probabilities of each class for each pixel point according to the position information of the image patch relative to the image I, and takes the mean value as the final prediction result of the pixel point. If the pixel point is covered by multiple image patches, the final prediction result is also calculated according to the principle of taking the mean value. Further, all the image patches cropped from the image I are combined according to the above principle to form the final segmentation result of the image I.
[0063] Embodiment 2
[0064] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above-mentioned remote sensing image semantic segmentation method based on CNN and Mamba is implemented.
[0065] Embodiment 3
[0066] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned remote sensing image semantic segmentation method based on CNN and Mamba is implemented.
[0067] Embodiment 4
[0068] In one embodiment, a computer program product is provided, including a computer program / instructions. When the computer program / instructions are executed by a processor, the above-mentioned remote sensing image semantic segmentation method based on CNN and Mamba is implemented.
[0069] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0070] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0071] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0072] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
Claims
1. A semantic segmentation method for remote sensing images based on CNN and Mamba, characterized in that, Including: Preprocess the acquired remote sensing image, and then use a remote sensing image semantic segmentation model to perform semantic segmentation on the image patches to obtain a pixel-level classification result; The remote sensing image semantic segmentation model includes a sequentially connected CNN encoder, a residual connection layer, and a CSMamba decoder; the preprocessed image is input into the CNN encoder, and the CNN encoder includes several stacked ResBlock modules; the residual connection layer includes several multi-scale residual spatial pyramid pooling ResASPP modules. The residual connection layers of each level of ResBlock modules in the CNN encoder are first cascade-processed to obtain combined features, and then the combined features are respectively connected to the corresponding levels of CSMamba modules in the CSMamba decoder through the ResASPP modules; the ResASPP module includes multiple convolutional attention units and an adaptive pooling layer. The convolutional kernels of the multiple convolutional attention units have the same size, and the dilation rate increases one by one. Each convolutional attention unit uses a residual connection network; a supervision module is added at each level of CSMamba module. The supervision module processes the intermediate feature information of the i-th CSMamba module to obtain an intermediate supervision signal, and the i-th CSMamba module adds the intermediate supervision signal to the corresponding elements of the intermediate feature information and then outputs; Among them, the calculation formula for obtaining combined features by cascade processing is: In the formula, is a combined feature, Concat is the Concat function in SQL, and F i , F i-1 , F i+1 are the features output by the residual connection layers of three adjacent ResBlock modules respectively.
2. The remote sensing image semantic segmentation method based on CNN and Mamba according to claim 1, characterized in that: The CSMamba decoder includes several stacked CSMamba modules. Each layer of CSMamba module obtains the output features of the previous layer through upsampling and combines the output features of the previous layer with the features output from the ResASPP module; The outputs of each level of CSMamba module are feature-mapped through 1×1Conv to output semantic segmentation results with different resolutions.
3. The remote sensing image semantic segmentation method based on CNN and Mamba according to claim 2, characterized in that: The CSMamba decoder designs a dual-branch feature processing. The first branch obtains the long-range dependence of image features through a 2D selective scanning module, and the second branch obtains features that enhance important channels and spatial positions through a channel attention and spatial attention mechanism. The output results of the two branches are aggregated and output through a Hadamard product.
4. The remote sensing image semantic segmentation method based on CNN and Mamba according to claim 1, characterized in that: The 2D selective scanning module first flattens the 2D image features into a 1D sequence; then scans the sequence in four directions of the 2D image, and uses the selective state space model SSM to capture the long-term correlation in each direction as the scanning result; finally, the scanning results in the four directions are combined and restored to the 2D image structure.
5. The remote sensing image semantic segmentation method based on CNN and Mamba according to claim 1, characterized in that: The supervision module processes the intermediate output of the i-th CSMamba module through the following formula: Among them, P i represents the intermediate supervision signal output by the supervision module, is the feature of the i-th CSMamba block, and Conv represents the convolution calculation.
6. The remote sensing image semantic segmentation method based on CNN and Mamba according to claim 1, characterized in that: During the training process of the remote sensing image semantic segmentation model, the expression of the cross-entropy loss function is as follows: Among them, MFB_CE loss represents the cross-entropy loss function, n represents the samples in the dataset, c represents the semantic categories of the images, N represents the number of samples, C represents the number of categories, and w c represents the weights of each category, and l c represents the true label of sample n, and p c represents the probability that the model predicts that sample n belongs to category c.
7. A remote sensing image semantic segmentation system based on CNN and Mamba, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the CNN and Mamba-based remote sensing image semantic segmentation method according to any one of claims 1 to 6.
8. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the CNN and Mamba-based remote sensing image semantic segmentation method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the remote sensing image semantic segmentation method based on CNN and Mamba described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program and / or instructions, characterized in that, When the computer program and / or instructions are executed by a processor, they implement the remote sensing image semantic segmentation method based on CNN and Mamba described in any one of claims 1 to 6.
Citation Information
Cited By
Photoacoustic image enhancement method and device combining Mama and CNN (Convolutional Neural Network)
CN120707580A
2D medical image segmentation method and system based on Mama and UNet
CN120997233A
Remote sensing change detection method and device for region-space-semantic modeling based on Mamba
CN120997707A
Mamba-based region-space-semantic modeling method and device for remote sensing change detection
CN120997707B
Remote sensing image semantic segmentation method and device, equipment and medium
CN121213935A