A Remote Sensing Image Segmentation Method and System Based on the Mamba Architecture

Through the remote sensing image segmentation method based on Mamba architecture and a special ship data set, the problem of extracting ship shapes in small and medium-sized complex images is solved, and high-precision ship segmentation and model performance are achieved.

CN119850654BActive Publication Date: 2025-06-03NANJING TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510321980.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-03
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

The prior art is difficult to accurately obtain the shape of small target ships in complex images, and the ship segmentation accuracy and generalization capabilities are insufficient, and there is a lack of high-quality data sets specifically for ship segmentation and high-precision high-resolution remote sensing image extraction network.

Method used

Using the remote sensing image segmentation method based on Mamba architecture, a high-quality data set is constructed specifically for ship segmentation, and through feature extraction modules and segmentation model training, the segmentation and extraction accuracy of small-target ships is improved.

Benefits of technology

It realizes high-precision ship segmentation in complex images, improves the model's performance in long-range dependency modeling and complex scene processing, and provides a reliable data set basis for subsequent research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119850654B_ABST
    Figure CN119850654B_ABST
Patent Text Reader

Abstract

The present invention discloses a remote sensing image segmentation method and system based on the Mamba architecture, which relates to the field of computer vision technology. The method includes: obtaining a ship remote sensing image and performing preprocessing to obtain an optimized image set; training a segmentation model based on the optimized image set to obtain a trained segmentation model; obtaining an image to be processed and sequentially inputting it into a partition layer and a first convolutional unit in the trained segmentation model to obtain a first feature sequence; inputting the first feature sequence into a downsampling unit to obtain a second feature sequence and a third feature sequence; inputting the third feature sequence into a bottleneck unit to obtain a fourth feature sequence; inputting the second feature sequence, the third feature sequence, and the fourth feature sequence into an upsampling unit to obtain a fifth feature sequence; inputting the first feature sequence and the fifth feature sequence into a second convolutional unit for processing and performing linear projection to obtain a ship segmentation result. The segmentation and extraction accuracy of small target ships in complex images is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and more specifically, to a remote sensing image segmentation method and system based on the Mamba architecture. Background Art

[0002] Currently, ship extraction technology is the core foundation of ocean monitoring, maritime traffic management, and ocean resource development. With the popularization of high-resolution satellite remote sensing technology, ship detection and segmentation based on remote sensing images have become a key means, providing rich data support for ocean dynamic monitoring. Among the existing ship extraction methods, they can be mainly divided into two categories: object detection and semantic segmentation.

[0003] Among them, object detection technology quickly identifies by locating the target area and has high timeliness. However, traditional object detection technology has a weak ability to accurately depict the ship boundary, pays more attention to target positioning rather than the precise description of the target geometry, and this characteristic often cannot meet the actual needs in tasks that require fine-grained boundary information; semantic segmentation technology provides a new idea for ship extraction through pixel-level prediction and the accurate depiction of ground object boundaries. It can not only locate the target but also completely describe its shape and details, thus showing unique advantages in refinement tasks. However, the existing semantic segmentation technology has problems such as being sensitive to complex ocean noise, having a high missed detection rate for small target ships, and insufficient generalization ability due to relying on limited datasets.

[0004] In recent years, the Mamba structure has received extensive attention due to its high precision and powerful long-range dependence modeling ability. The existing technology Mamba-Transformer backbone network combines the characteristics of convolution and Transformer, breaking through the bottleneck of the degradation of segmentation performance of traditional methods in complex backgrounds. The existing technology proposes the general network U-Mamba for biomedical image segmentation, opening up a new way for long-range dependence modeling in biomedical image analysis. Although these methods have made important progress in some fields, due to the diversity, randomness of position, and uncertainty of quantity of ship images, it poses great challenges to the improvement of model performance and the generalization ability of practical applications. In addition, the existing research lacks high-quality datasets specifically for ship segmentation and high-precision high-resolution remote sensing image extraction networks, and there are still significant gaps in the ship research field.

[0005] Therefore, how to accurately obtain the shape of small target ships in complex images and effectively improve the ship segmentation accuracy is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0006] In view of this, the present invention provides a remote sensing image segmentation method based on the Mamba architecture, constructs a high-quality dataset specifically for ship segmentation, and improves the segmentation and extraction accuracy of small target ships in complex images.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] A remote sensing image segmentation method based on the Mamba architecture, comprising:

[0009] Obtain a plurality of ship remote sensing images and perform segmentation and optimization processing to obtain an optimized image set;

[0010] Train a segmentation model based on the optimized image set to obtain a trained segmentation model;

[0011] Obtain an image to be processed and input it into the trained segmentation model;

[0012] The image to be processed is input into a partitioning layer to obtain a plurality of image blocks;

[0013] Based on the plurality of image blocks, input them into a first convolutional unit to obtain a first feature sequence;

[0014] The first feature sequence is input into a downsampling unit to obtain a second feature sequence and a third feature sequence;

[0015] The third feature sequence is input into a bottleneck unit to obtain a fourth feature sequence;

[0016] The second feature sequence, the third feature sequence, and the fourth feature sequence are input into an upsampling unit to obtain a fifth feature sequence;

[0017] The first feature sequence and the fifth feature sequence are input into a second convolutional unit to obtain a sixth feature sequence and perform linear projection to obtain a ship segmentation result.

[0018] Preferably, the segmentation specifically includes:

[0019] Frame the ship remote sensing image based on the ship label coordinates to obtain a framed area;

[0020] Perform masking processing based on the framed area and cut to generate a corresponding segmentation mask;

[0021] Cut the ship remote sensing image based on the segmentation mask to obtain a segmented image.

[0022] Preferably, the optimization processing specifically includes:

[0023] Apply morphological opening operation based on the segmented image to remove small white noises in the image to obtain a denoised image;

[0024] Apply morphological closing operation to the denoised image to fill holes and obtain an optimized image;

[0025] Obtain the set of optimized images based on all the optimized images.

[0026] Preferably, the first convolution unit includes: a first depth convolution and a first processing unit;

[0027] The multiple image patches are input into the first depth convolution to obtain multiple corresponding short sequences;

[0028] Concatenate all the short sequences to obtain a one-dimensional long sequence;

[0029] The one-dimensional long sequence is input into the first processing unit to obtain the first feature sequence.

[0030] Preferably, the downsampling unit includes: a second processing unit and a third processing unit;

[0031] The first feature sequence is input into the second processing unit after patch merging operation to obtain the second feature sequence;

[0032] The second feature sequence is input into the third processing unit after patch merging operation to obtain the third feature sequence;

[0033] The bottleneck unit includes: a fourth processing unit and a fifth processing unit;

[0034] The third feature sequence is sequentially input into the fourth processing unit and the fifth processing unit to obtain the fourth feature sequence.

[0035] Preferably, the upsampling unit includes: a sixth processing unit and a seventh processing unit;

[0036] The fourth feature sequence is input into the sixth processing unit to obtain a first processing sequence;

[0037] The first processing sequence and the third feature sequence are fused and then patch expansion operation is performed to obtain a first expansion sequence;

[0038] The first expansion sequence is input into the seventh processing unit to obtain a second processing sequence;

[0039] The second processing sequence and the second feature sequence are fused and then patch expansion operation is performed to obtain the fifth feature sequence.

[0040] Preferably, the second convolution unit includes: an eighth processing unit and a second depth convolution;

[0041] The fifth feature sequence is input into the eighth processing unit to obtain a third processing sequence;

[0042] The third processing sequence and the first feature sequence are fused and then input into the second depth convolution to obtain the sixth feature sequence.

[0043] Preferably, the structures of all processing units are the same, including two ResMambaBlock modules connected in series;

[0044] Each ResMambaBlock module includes: a first normalization layer, a first linear transformation layer, a first standardization layer, an activation layer, a MambaLayer module, a second normalization layer, a third normalization layer, and a second linear transformation layer;

[0045] The input sequence is input into the first normalization layer to obtain a first intermediate sequence;

[0046] The first intermediate sequence is sequentially input into the first linear transformation layer, the first standardization layer, and the activation layer to obtain a second intermediate sequence;

[0047] The second intermediate sequence is input into the MambaLayer module to obtain a third intermediate sequence;

[0048] The third intermediate sequence is input into the second normalization layer to obtain a fourth intermediate sequence;

[0049] The first intermediate sequence is input into the second linear transformation layer to obtain a fifth intermediate sequence;

[0050] The fifth intermediate sequence and the fourth intermediate sequence are fused and then input into the third normalization layer to obtain a sixth intermediate sequence;

[0051] The sixth intermediate sequence and the input sequence are fused to obtain an output sequence.

[0052] Preferably, the MambaLayer module includes: a fourth normalization layer, a second standardization layer, an SSM model, a fifth normalization layer, and a third linear transformation layer;

[0053] The second intermediate sequence is sequentially input into the fourth normalization layer, the second standardization layer, and the SSM model to obtain a process feature sequence;

[0054] The process feature sequence is fused with learnable parameters and then sequentially input into the fifth normalization layer and the third linear transformation layer to obtain the third intermediate sequence.

[0055] A remote sensing image segmentation system based on the Mamba architecture, comprising: a data acquisition module, a model training module, an image segmentation module, a feature extraction module, and a result output module;

[0056] The data acquisition module is used to acquire multiple ship remote sensing images, perform segmentation and optimization processing, and obtain an optimized image set;

[0057] The model training module is used to train a segmentation model based on the optimized image set to obtain a trained segmentation model;

[0058] The image segmentation module is used to acquire an image to be processed and input it into the trained segmentation model; the image to be processed is input into a partitioning layer to obtain a plurality of image blocks;

[0059] The feature extraction module is used to input the plurality of image blocks into a first convolutional unit to obtain a first feature sequence; the first feature sequence is input into a downsampling unit to obtain a second feature sequence and a third feature sequence; the third feature sequence is input into a bottleneck unit to obtain a fourth feature sequence; the second feature sequence, the third feature sequence, and the fourth feature sequence are input into an upsampling unit to obtain a fifth feature sequence;

[0060] The result output module is used to input the first feature sequence and the fifth feature sequence into a second convolutional unit to obtain a sixth feature sequence and perform linear projection to obtain a ship segmentation result.

[0061] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a remote sensing image segmentation method based on the Mamba architecture, which has the following beneficial effects:

[0062] 1. The dataset constructed by the present invention not only restores the image data of the original size, but also ensures the integrity and clarity of the ship boundary during the cutting process. Its high-quality annotation and detailed segmentation provide a reliable basis for subsequent research and strong support for improving the performance of the ship target detection model.

[0063] 2. The image segmentation model disclosed by the present invention can effectively restore the spatial resolution of the image, avoid the gradient disappearance problem in the traditional network, and aims to comprehensively improve the performance of the model in long-range dependence modeling and complex scene processing.

[0064] 3. The present invention effectively improves the feature representation ability and spatial detail restoration ability by constructing a ResMambaBlock module based on the state space, thereby enhancing the model's ability to process complex images and small targets. Description of the Drawings

[0065] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0066] Figure 1 It is a flowchart of a remote sensing image segmentation method based on the Mamba architecture provided by the present invention.

[0067] Figure 2 It is a flowchart of an optimized image set construction method provided by the present invention.

[0068] Figure 3 It is a schematic diagram of the segmentation model structure provided by the present invention.

[0069] Figure 4 It is a schematic diagram of the ResMambaBlock module structure provided by the present invention.

[0070] Figure 5 It is a schematic diagram of the MambaLayer module structure provided by the present invention.

[0071] Figure 6 It is a schematic diagram of the segmentation result of the segmentation model for processing other target remote sensing images provided by the present invention.

[0072] Figure 7 It is a schematic diagram of the structure of a remote sensing image segmentation system based on the Mamba architecture provided by the present invention. Detailed implementation manners

[0073] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0074] Embodiment 1

[0075] As Figure 1 shown, the embodiment of the present invention discloses a remote sensing image segmentation method based on the Mamba architecture, including:

[0076] Obtain multiple ship remote sensing images and perform segmentation and optimization processing to obtain an optimized image set;

[0077] Train the segmentation model based on the optimized image set to obtain a trained segmentation model;

[0078] Obtain the image to be processed and input it into the trained segmentation model;

[0079] Input the image to be processed into the partitioning layer to obtain multiple image patches;

[0080] Based on the multiple image patches, input them into the first convolutional unit to obtain the first feature sequence;

[0081] Input the first feature sequence into the downsampling unit to obtain the second feature sequence and the third feature sequence;

[0082] Input the third feature sequence into the bottleneck unit to obtain the fourth feature sequence;

[0083] Input the second feature sequence, the third feature sequence, and the fourth feature sequence into the upsampling unit to obtain the fifth feature sequence;

[0084] Input the first feature sequence and the fifth feature sequence into the second convolutional unit to obtain the sixth feature sequence and perform linear projection to obtain the ship segmentation result.

[0085] Embodiment 2

[0086] The embodiment of the present invention discloses a remote sensing image segmentation method based on the Mamba architecture, including:

[0087] As Figure 2 shown, obtain multiple ship remote sensing images, perform segmentation and optimization processing on them to obtain an optimized image set.

[0088] Preferably, in this embodiment, satellite remote sensing images on Google Earth are selected. These satellite remote sensing image datasets cover various objects such as airplanes, bridges, chimneys, dams, ports, ships, storage tanks, etc. To focus on the ship extraction task, the images and labels in the dataset are screened, only the images related to ships are retained, and those unclear or not containing ships are excluded to ensure that the screened dataset only contains accurate and high-quality ship images.

[0089] Preferably, the segmentation specifically includes:

[0090] Based on the ship label coordinates, frame the ship remote sensing image to obtain a framed area;

[0091] Based on the framed area, perform masking processing and cut to generate corresponding segmentation masks;

[0092] Based on the segmentation masks, cut the ship remote sensing image to obtain segmented images.

[0093] Preferably, through the segmentation process, ship targets can be effectively extracted. However, the images after cutting may still have defects and noises, which will directly affect the subsequent training effect. Therefore, morphological repair operations are introduced to optimize the segmentation results.

[0094] Preferably, in this embodiment, the SAM model is used to segment the ship remote sensing image.

[0095] Preferably, the optimization process specifically includes:

[0096] Applying morphological opening operation based on the segmented image to remove the small white noise in the image, obtaining a denoised image;

[0097] Applying morphological closing operation based on the denoised image to fill the holes, obtaining an optimized image;

[0098] Based on all the optimized images, an optimized image set is obtained as the Boat R dataset.

[0099] Preferably, the above segmentation and optimization processing effectively improve the quality of the image, can generate a set of high-quality ship segmentation datasets, and provide a solid foundation for the subsequent training of the ship segmentation model.

[0100] Preferably, there is a dilemma in the field of satellite remote sensing ships that there is a lack of high-quality datasets specifically for ship segmentation. Some of the currently existing publicly available remote sensing image segmentation datasets are shown in Table 1:

[0101] Table 1 Comparison between Boat R dataset and some publicly available remote sensing image segmentation datasets

[0102]

[0103] As can be seen from Table 1, most of the remote sensing ship target datasets are limited to the visible light band, and only a very small number combine data from other bands. In these datasets, the original remote sensing images often have a small image size after being cropped, resulting in problems such as incomplete boundaries and unclear images of the segmented ship images, seriously affecting the accuracy of subsequent analysis and model training.

[0104] In contrast, the Boat R dataset established by the present invention not only restores the image data of the original size, but also ensures the integrity and clarity of the ship boundaries during the segmentation process, making it have important application value in the training of the ship segmentation model. Its high-quality annotation and detailed segmentation provide a reliable basis for subsequent research and strong support for improving the performance of the ship target detection and segmentation model.

[0105] Based on the optimized image set, the segmentation model is trained to obtain a trained segmentation model, as Figure 3 shown.

[0106] Preferably, the input size of the segmentation model is a three-dimensional color image of H×W×C. Each image is first segmented into multiple image patches and converted into a one-dimensional sequence of dimension H / 4×W / 4×16. The feature dimension is adjusted to any size through an initial linear embedding layer for subsequent processing.

[0107] Preferably, during the training process, the patch tokens after being processed by the embedding layer gradually extract features through multiple ResMambaBlocks and patch merging layers. The segmentation model MambaSegNet gradually reduces the spatial resolution through downsampling operations and increases the feature dimension.

[0108] Each ResMambaBlock module can efficiently learn the features of complex images, especially showing outstanding performance in the processing of small targets and complex scenes. The output resolution gradually decreases from H×W×C to H / 4×W / 4×C, H / 8×W / 8×2C, H / 16×W / 16×4C, and finally to H / 32×W / 32×8C. The spatial resolution of the image is gradually restored through upsampling operations, and residual connections are used to combine shallow features with deep features to restore the lost spatial details. Residual connections effectively avoid the problem of gradient disappearance and improve the stability and convergence speed during the training process.

[0109] Preferably, the convolutional layer of the segmentation model MambaSegNet uses depthwise convolution operations, significantly reducing the number of parameters and improving the computational efficiency. Finally, the high-dimensional feature is mapped to the target number of channels through a 1×1 convolution to generate the final segmentation result.

[0110] Preferably, in this embodiment, during the training process, a combined loss function is adopted, which combines cross-entropy loss and IoU loss. Among them, the cross-entropy loss L CE is:

[0111] ;

[0112] where, y i represents the true label value, taking values of 0 or 1, p i represents the probability value predicted by the model, and N represents the total number of pixels

[0113] The IoU loss L IoU is:

[0114] ;

[0115] where, P represents the predicted region, and G represents the true region;

[0116] Combined loss function formula:

[0117] ;

[0118] Among them, α and β control the weights of the cross-entropy loss and the intersection over union (IoU) loss respectively. This combined strategy balances the influence of the two, enabling the model to accurately classify and effectively capture the target area during training to ensure the balance between classification accuracy and region overlap.

[0119] Preferably, during the optimization process of this embodiment, the Adam optimizer is used, the learning rate is set to 0.0001, the batch size is 4, and the ratio of the training dataset to the validation dataset is 7:2.

[0120] Preferably, the trained segmentation model MambaSegNet has achieved excellent results in the ship segmentation task of high-resolution remote sensing images. Especially in the segmentation of complex backgrounds and small targets, it has shown higher accuracy and computational efficiency compared with existing models. The fully trained model can accurately extract the boundaries and details of ship targets, has high robustness, and is applicable to various practical application scenarios such as marine monitoring and ship emission monitoring.

[0121] Preferably, an accuracy evaluation index is constructed by calculating the loss function:

[0122] The accuracy evaluation index is a key criterion for evaluating the segmentation performance of a classification model on the test set, especially significant in classification tasks. Common evaluation indexes include recall, precision, accuracy, F1-score, and intersection over union (IoU). Each index has unique characteristics and can comprehensively evaluate the performance of the model from different perspectives, suitable for different task requirements.

[0123] Precision measures the proportion of samples predicted as positive that are actually positive, focusing on evaluating the prediction ability of the model for positive classes, and is particularly suitable for tasks with more false positives.

[0124] ;

[0125] Among them, TP represents the number of pixels predicted as positive and actually positive, and FP represents the number of pixels predicted as positive but actually negative.

[0126] Recall focuses on how many samples that are actually positive are correctly identified as positive, emphasizing the capture ability of the model for positive samples, and plays an important role especially when it is necessary to identify as many positive samples as possible.

[0127] ;

[0128] Among them, FN represents the number of pixels predicted as negative but actually positive.

[0129] Accuracy reflects the overall classification accuracy of the model for all samples. Although it can provide an overview of the overall performance of the model, it may not fully reflect the actual capabilities of the model when the sample distribution is uneven.

[0130] 。

[0131] The F1 value, as the harmonic mean of precision and recall, is suitable for tasks that require balancing these two, especially important in small object segmentation tasks.

[0132] 。

[0133] The IoU metric measures the degree of overlap between the predicted region and the ground truth region, and is particularly suitable for object detection and segmentation tasks, capable of evaluating the accuracy of the model in localizing and segmenting the target region.

[0134] 。

[0135] These evaluation metrics complement each other, providing the ability to consider the model performance from multiple dimensions, ensuring that the evaluation results of the model are more comprehensive and accurate in different tasks and scenarios.

[0136] Obtain the image to be processed and input it into the trained segmentation model.

[0137] The image to be processed is input into the partitioning layer to obtain multiple image patches.

[0138] Based on the multiple image patches input into the first convolutional unit, a first feature sequence is obtained.

[0139] Preferably, the first convolutional unit includes: a first depth convolution and a first processing unit;

[0140] The multiple image patches are input into the first depth convolution to obtain multiple corresponding short sequences;

[0141] Based on the splicing of all the short sequences, a one-dimensional long sequence is obtained;

[0142] The one-dimensional long sequence is input into the first processing unit to obtain a first feature sequence.

[0143] The first feature sequence is input into the downsampling unit to obtain a second feature sequence and a third feature sequence.

[0144] Preferably, the downsampling unit includes: a second processing unit and a third processing unit;

[0145] After the first feature sequence undergoes a patch merging operation, it is input into the second processing unit to obtain a second feature sequence;

[0146] After the second feature sequence undergoes patch merging operation, it is input into the third processing unit to obtain the third feature sequence.

[0147] The third feature sequence is input into the bottleneck unit to obtain the fourth feature sequence.

[0148] Preferably, the bottleneck unit includes: a fourth processing unit and a fifth processing unit;

[0149] The third feature sequence is sequentially input into the fourth processing unit and the fifth processing unit to obtain the fourth feature sequence.

[0150] The second feature sequence, the third feature sequence, and the fourth feature sequence are input into the upsampling unit to obtain the fifth feature sequence.

[0151] Preferably, the upsampling unit includes: a sixth processing unit and a seventh processing unit;

[0152] The fourth feature sequence is input into the sixth processing unit to obtain the first processing sequence;

[0153] After the first processing sequence and the third feature sequence are fused, a patch expansion operation is performed to obtain the first expansion sequence;

[0154] The first expansion sequence is input into the seventh processing unit to obtain the second processing sequence;

[0155] After the second processing sequence and the second feature sequence are fused, a patch expansion operation is performed to obtain the fifth feature sequence.

[0156] The first feature sequence and the fifth feature sequence are input into the second convolutional unit to obtain the sixth feature sequence and perform linear projection to obtain the ship segmentation result.

[0157] Preferably, the second convolutional unit includes: an eighth processing unit and a second depth convolution;

[0158] The fifth feature sequence is input into the eighth processing unit to obtain the third processing sequence;

[0159] After the third processing sequence and the first feature sequence are fused, they are input into the second depth convolution to obtain the sixth feature sequence.

[0160] Preferably, the structures of all processing units are the same, that is, the first processing unit, the second processing unit, the third processing unit, the fourth processing unit, the fifth processing unit, the sixth processing unit, the seventh processing unit, and the eighth processing unit have the same structure, and all include two serially connected ResMambaBlock modules;

[0161] Such as Figure 4As shown, each ResMambaBlock module includes: a first normalization layer, a first linear transformation layer, a first standardization layer, an activation layer, a MambaLayer module, a second normalization layer, a third normalization layer, and a second linear transformation layer;

[0162] The input sequence is input to the first normalization layer to obtain a first intermediate sequence;

[0163] The first intermediate sequence is successively input to the first linear transformation layer, the first standardization layer, and the activation layer to obtain a second intermediate sequence;

[0164] The second intermediate sequence is input to the MambaLayer module to obtain a third intermediate sequence;

[0165] The third intermediate sequence is input to the second normalization layer to obtain a fourth intermediate sequence;

[0166] The first intermediate sequence is input to the second linear transformation layer to obtain a fifth intermediate sequence;

[0167] The fifth intermediate sequence and the fourth intermediate sequence are fused and then input to the third normalization layer to obtain a sixth intermediate sequence;

[0168] The sixth intermediate sequence and the input sequence are fused to obtain an output sequence.

[0169] Preferably, different from the traditional vision Transformer, the ResMambaBlock module adopted in the present invention abandons the positional embedding and selects a more simplified structure, removing the MLP stage, enabling more modules to be stacked under the same computational budget, improving the computational efficiency and feature expression ability. In addition, by introducing the MambaLayer module, the network can effectively learn complex image features when processing small targets and complex scenes in remote sensing images.

[0170] Preferably, by constructing a ResMambaBlock module based on the state space, the feature representation ability and the spatial detail restoration ability are effectively improved, thereby enhancing the model's ability to process complex images and small targets.

[0171] Preferably, as Figure 5 shown, the MambaLayer module includes: a fourth normalization layer, a second standardization layer, an SSM model, a fifth normalization layer, and a third linear transformation layer;

[0172] The second intermediate sequence is successively input to the fourth normalization layer, the second standardization layer, and the SSM model to obtain a process feature sequence;

[0173] The process feature sequence and the learnable parameters are fused and then sequentially input into the fifth normalization layer and the third linear transformation layer to obtain a third intermediate sequence.

[0174] Preferably, the SSM model is a Mamba state space model-based model, and the learning parameters are used to adjust the proportion of skip connections.

[0175] Preferably, by training and testing on road and building datasets, the segmentation model MambaSegNet of the present invention can be applied to the segmentation and extraction of roads and buildings. The segmentation and extraction results of remote sensing images for other targets are as Figure 6 shown. The results prove that the segmentation model MambaSegNet of the present invention shows strong adaptability and stability in both the extraction of road networks and the precise segmentation of building forms.

[0176] Example 3

[0177] As Figure 7 shown, a remote sensing image segmentation system based on the Mamba architecture includes: a data acquisition module, a model training module, an image segmentation module, a feature extraction module, and a result output module;

[0178] The data acquisition module is used to acquire multiple ship remote sensing images, perform segmentation and optimization processing, and obtain an optimized image set;

[0179] The model training module is used to train the segmentation model based on the optimized image set to obtain a trained segmentation model;

[0180] The image segmentation module is used to acquire the image to be processed and input it into the trained segmentation model; the image to be processed is input into the partitioning layer to obtain multiple image patches;

[0181] The feature extraction module is used to input multiple image patches into the first convolutional unit to obtain a first feature sequence; the first feature sequence is input into the downsampling unit to obtain a second feature sequence and a third feature sequence; the third feature sequence is input into the bottleneck unit to obtain a fourth feature sequence; the second feature sequence, the third feature sequence, and the fourth feature sequence are input into the upsampling unit to obtain a fifth feature sequence;

[0182] The result output module is used to input the first feature sequence and the fifth feature sequence into the second convolutional unit to obtain a sixth feature sequence and perform linear projection to obtain the ship segmentation result.

[0183] Preferably, the functional implementation methods of the modules in this embodiment correspond one by one to the above methods, and will not be elaborated here one by one.

[0184] Example 4

[0185] Based on the same inventive concept, the present invention also provides a computer device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;

[0186] The memory is used to store computer programs;

[0187] When the processor is used to execute the program stored in the memory, it can implement a remote sensing image segmentation method based on the Mamba architecture as described in Embodiment 1 or 2.

[0188] The electronic device may include: a processor, a communications interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus.

[0189] The processor can call the logical instructions in the memory to execute a remote sensing image segmentation method based on the Mamba architecture as described in Embodiment 1 or 2.

[0190] In addition, when the above-mentioned logical instructions in the memory are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present invention. And the aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0191] Through the above technical solutions, it can be seen that compared with the prior art, the present invention discloses a remote sensing image segmentation method based on the Mamba architecture, which has the following beneficial effects:

[0192] 1. The dataset constructed by the present invention not only restores the image data of the original size, but also ensures the integrity and clarity of the ship boundary during the cutting process. Its high-quality annotation and detailed segmentation provide a reliable basis for subsequent research and strong support for improving the performance of the ship target detection model.

[0193] 2. The image segmentation model disclosed by the present invention can effectively restore the spatial resolution of images, avoiding the problem of gradient disappearance in traditional networks, and aims to comprehensively improve the performance of the model in long-range dependence modeling and complex scene processing.

[0194] 3. By constructing a ResMambaBlock module based on the state space, the present invention effectively improves the feature representation ability and the ability to restore spatial details, thereby enhancing the model's ability to process complex images and small targets.

[0195] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference can be made to the description of the method part for the relevant parts.

[0196] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A remote sensing image segmentation method based on Mamba architecture, characterized in that: include: Acquire multiple ship remote sensing images and perform segmentation and optimization processing to obtain an optimized image set; Training the segmentation model based on the optimized image set to obtain a trained segmentation model; Obtaining the image to be processed and inputting it into the trained segmentation model; The image to be processed is input into a partitioning layer to obtain a plurality of image blocks; Based on the multiple image blocks, input into a first convolution unit to obtain a first feature sequence; The first feature sequence is input into a downsampling unit to obtain a second feature sequence and a third feature sequence; The third feature sequence is input into the bottleneck unit to obtain a fourth feature sequence; The second feature sequence, the third feature sequence and the fourth feature sequence are input into an upsampling unit to obtain a fifth feature sequence; The first feature sequence and the fifth feature sequence are input into the second convolution unit to obtain a sixth feature sequence and linear projection is performed to obtain a ship segmentation result; The first convolution unit includes: a first depth convolution and a first processing unit; The down sampling unit comprises: a second processing unit and a third processing unit; The bottleneck unit includes: a fourth processing unit and a fifth processing unit; The up-sampling unit comprises: a sixth processing unit and a seventh processing unit; The second convolution unit includes: an eighth processing unit and a second depth convolution; All processing units have the same structure, consisting of two ResMambaBlock modules connected in series; The ResMambaBlock modules all include: a first normalization layer, a first linear transformation layer, a first standardization layer, an activation layer, a MambaLayer module, a second normalization layer, a third normalization layer and a second linear transformation layer; The MambaLayer module includes: a fourth normalization layer, a third standardization layer, an SSM model, a fifth normalization layer and a third linear transformation layer.

2. The remote sensing image segmentation method based on Mamba architecture according to claim 1, characterized in that: The segmentation specifically includes: Frame the ship remote sensing image based on the ship tag coordinates to obtain a framed area; Performing mask processing based on the framed area and cutting to generate a corresponding segmentation mask; The ship remote sensing image is segmented based on the segmentation mask to obtain a segmented image.

3. The remote sensing image segmentation method based on Mamba architecture according to claim 2, characterized in that: The optimization process specifically includes: Applying a morphological opening operation based on the segmented image to remove small white noise in the image to obtain a denoised image; Based on the denoised image, a morphological closing operation is applied to fill holes to obtain an optimized image; The optimized image set is obtained based on all optimized images.

4. The remote sensing image segmentation method based on Mamba architecture according to claim 1, characterized in that: The multiple image blocks are input into the first deep convolution to obtain multiple corresponding short sequences; Based on all the short sequences, splicing is performed to obtain a one-dimensional long sequence; The one-dimensional long sequence is input into the first processing unit to obtain the first feature sequence.

5. The remote sensing image segmentation method based on Mamba architecture according to claim 4, characterized in that: The first feature sequence is subjected to a patch merging operation and then input into the second processing unit to obtain the second feature sequence; The second feature sequence is subjected to a patch merging operation and then input into the third processing unit to obtain the third feature sequence; The third feature sequence is sequentially input to the fourth processing unit and the fifth processing unit to obtain the fourth feature sequence.

6. The remote sensing image segmentation method based on Mamba architecture according to claim 1, characterized in that: The fourth feature sequence is input into the sixth processing unit to obtain a first processing sequence; The first processing sequence and the third feature sequence are fused and then patch-extended to obtain a first extended sequence; The first extended sequence is input into the seventh processing unit to obtain a second processing sequence; The second processing sequence and the second feature sequence are fused and then patch-extended to obtain the fifth feature sequence.

7. The remote sensing image segmentation method based on Mamba architecture according to claim 6, characterized in that: The fifth feature sequence is input into the eighth processing unit to obtain a third processing sequence; The third processing sequence and the first feature sequence are fused and input into the second deep convolution to obtain the sixth feature sequence.

8. A remote sensing image segmentation method based on Mamba architecture according to claim 5 or 7, characterized in that: The input sequence is input into the first normalization layer to obtain a first intermediate sequence; The first intermediate sequence is sequentially input into the first linear transformation layer, the first normalization layer and the activation layer to obtain a second intermediate sequence; The second intermediate sequence is input into the MambaLayer module to obtain a third intermediate sequence; The third intermediate sequence is input into the second normalization layer to obtain a fourth intermediate sequence; The first intermediate sequence is input into the second linear transformation layer to obtain a fifth intermediate sequence; The fifth intermediate sequence and the fourth intermediate sequence are fused and input into the third normalization layer to obtain a sixth intermediate sequence; The sixth intermediate sequence and the input sequence are fused to obtain an output sequence.

9. The remote sensing image segmentation method based on Mamba architecture according to claim 8, characterized in that: The second intermediate sequence is sequentially input into the fourth normalization layer, the third standardization layer and the SSM model to obtain a process feature sequence; The process feature sequence is fused with the learnable parameters and then sequentially input into the fifth normalization layer and the third linear transformation layer to obtain the third intermediate sequence.

10. A remote sensing image segmentation system based on Mamba architecture, applied to a remote sensing image segmentation method based on Mamba architecture as claimed in any one of claims 1 to 9, characterized in that: include: Data acquisition module, model training module, image segmentation module, feature extraction module and result output module; The data acquisition module is used to acquire multiple ship remote sensing images and perform segmentation and optimization processing to obtain an optimized image set; The model training module is used to train the segmentation model based on the optimized image set to obtain a trained segmentation model; The image segmentation module is used to obtain the image to be processed and input it into the trained segmentation model; the image to be processed is input into the partition layer to obtain multiple image blocks; The feature extraction module is used to input the plurality of image blocks into a first convolution unit to obtain a first feature sequence; the first feature sequence is input into a downsampling unit to obtain a second feature sequence and a third feature sequence; the third feature sequence is input into a bottleneck unit to obtain a fourth feature sequence; the second feature sequence, the third feature sequence and the fourth feature sequence are input into an upsampling unit to obtain a fifth feature sequence; The result output module is used to input the first feature sequence and the fifth feature sequence into the second convolution unit to obtain the sixth feature sequence and perform linear projection to obtain the ship segmentation result.

Citation Information

Patent Citations

  • Optical remote sensing image segmentation method based on VMama model

    CN118365882A

  • Multi-modal fusion segmentation method and device based on prior information and Mama hybrid model

    CN118691820A