Land utilization remote sensing image semantic segmentation method based on improved RMSIN model
By improving the RMSIN model and the method of fusion multi-scale feature map, the problem of insufficient model complexity and robustness in the existing remote sensing image segmentation method is solved, and a more efficient and accurate semantic segmentation effect of remote sensing image is achieved.
Patent Information
- Application Number
- CN202510260339.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-03-06
AI Technical Summary
In the existing semantic segmentation methods of remote sensing images, the model structure is complex and the parameter adjustment is difficult, and the robustness is insufficient, especially when facing complex noise.
A fusion model based on the improved RMSIN model is adopted to achieve the fusion of multi-scale features through the stitching of multi-scale feature maps and upsampling of jump connections, and combined with the self-attention mechanism, the segmentation accuracy and efficiency of the model are improved.
The segmentation accuracy and efficiency of semantic segmentation of remote sensing images are significantly improved, especially when processing complex images and detailed areas, different types of land objects can be more accurately identified, and the processing ability of diverse land objects is improved.
Smart Images

Figure CN119942126A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of remote sensing image segmentation, and in particular to a land use remote sensing image semantic segmentation method based on an improved RMSIN model. Background Art
[0002] Information about the features of the Earth's surface obtained through remote sensing technology. This data mainly comes from sensors carried by satellites or aircraft to collect remote sensing land image data. Remote sensing land image data can be used for a variety of applications, including environmental monitoring, agricultural management, urban planning, etc., and is particularly important for managing dynamic changes in land information because it can provide continuous observations over a wide area. Remote sensing technology can be used to obtain valuable information about land changes and resource allocation, and to distinguish different types of land use, such as cities, farmland, forests, and water bodies.
[0003] In the prior art, CN119152205A discloses a multi-source remote sensing image semantic segmentation method based on Transformer, Mamba and diffusion model, using row-column pooling Transformer module instead of traditional Transformer module as the backbone, and both RGB information and depth information will be processed by multiple row-column pooling Transformer modules, while extracting detail texture information and large-scale structure information in the vertical and horizontal directions of the feature map. In this process, we improve the time step position encoding, combine Gaussian position encoding and trigonometric function position encoding, and embed the time step into the model. After being processed by a certain row-column pooling Transformer module, the RGB features and depth features will enter the cross-Mamba attention fusion module, which replaces the traditional Mamba module. This module constructs relevant parameters based on the HiPPO continuous time memory theory, and exchanges specific matrices through state space equation operations. Subsequently, the feature weighted mask attention fusion module is used to weight the RGB and depth features, highlighting the important feature areas and then cross-attention fusion. First, the weight mask is generated by deep separable convolution and Softmax activation function, and then reshaped, matrix multiplication and other operations are performed, and finally the cross attention mechanism is input to complete the fusion. The decoder module consists of a convolution layer, a linear layer, upsampling using bilinear interpolation and a Sigmoid activation function. The output features are obtained by deep separable convolution, linear transformation, upsampling and activation function operations. Finally, the model training and optimization process training phase parameters are performed. However, the segmentation method disclosed in CN119152205A has a complex model structure and is difficult to adjust the model parameters. There are many hyperparameters in the model, and the modules affect each other. Finding the optimal combination requires a lot of experiments and time. During the training process, noise samples are randomly extracted from the Gaussian distribution and added to the data, assuming that the noise in the actual application conforms to the Gaussian distribution. However, in real remote sensing scenes, the sources of noise are complex and diverse, and may not completely conform to the Gaussian distribution. This specific assumption about noise may cause the model to be insufficiently robust in the face of actual complex noise, and the segmentation effect is reduced.
[0004] In the prior art, CN114419449A discloses a remote sensing image semantic segmentation method with self-attention multi-scale feature fusion. The Swin-T network is pre-trained on the data set, used as an encoder, and a pyramid structure network with self-attention multi-scale feature fusion is used as a decoder. The feature encoder contains four Swin-T modules connected in sequence, which collect features of the input remote sensing image in sequence, generate four feature maps of different scales and pass them to the decoder. The decoder has three self-attention multi-scale feature fusion modules, into which the feature maps generated by the first, second and third modules of Swin-T are input. Each module performs upsampling, channel number adjustment and other operations on the feature maps generated by different modules according to the scale and number of channels at the current stage to achieve feature fusion. The calculation attention module in each feature fusion module performs global average pooling on the three adjusted feature maps, splices the pooled feature maps to calculate self-attention, and the split result obtains the correlation score, which is multiplied with the original feature map and spliced in the channel dimension and the number of channels is adjusted. The feature map generated by the fourth module of Swin-T is superimposed with the feature map output by each feature fusion module in sequence, all the superimposed feature maps and the feature map generated by the fourth module of Swin-T are upsampled to the original image size, each pixel is predicted and classified, and the prediction results of the four scales are fused to obtain the final segmentation result. In addition, a multi-scale strategy is adopted when performing remote sensing semantic segmentation prediction, and the input image is scaled according to different proportions for segmentation prediction, and finally the segmentation results at all scales are superimposed. However, the segmentation method disclosed in CN114419449A may overemphasize the correlation of distant pixels in some cases, especially when there is a lack of obvious contextual dependence in the image. This over-modeling may lead to a waste of computing resources, especially when processing local features, resulting in information redundancy. In addition, the differences in spatial size and semantic information of feature maps of different scales make multi-scale fusion complicated, especially in the alignment and fusion of pixel-level details. If not handled properly, it may affect the final segmentation effect. Summary of the invention
[0005] In view of the problems existing in the above-mentioned prior art, the present invention provides a fusion model, a semantic segmentation model of land use remote sensing images based on the fusion model and the improved RMSIN model, and a semantic segmentation method of land use remote sensing images, which significantly improve the segmentation accuracy and efficiency of the model.
[0006] In a first aspect, a fusion model for fusing multi-scale feature maps is provided, including:
[0007] A multi-scale feature map acquisition module is used to acquire feature maps of multiple different scales. Assume that the number of feature maps is N, and the feature maps are sorted from small to large in scale as a first-scale feature map, a second-scale feature map, ..., an N-th-scale feature map;
[0008] A first splicing module, used to splice the multiple different-scale feature maps acquired by the multi-scale feature map acquisition module according to channels and output a first splicing feature map;
[0009] A first fusion module is jump-connected to the first scale feature map, and is used to perform a first upsampling on the first concatenated feature map, and then concatenate and fuse the first scale feature map obtained by using the jump connection with the feature map output after the first upsampling, and then perform a convolution operation on the concatenated fused feature map and output it;
[0010] The second fusion module is jump-connected with the second scale feature map, and is used to perform a second upsampling on the feature map output by the first fusion module, and then concatenate and fuse the second scale feature map obtained by the jump connection with the feature map output after the second upsampling, and then perform a convolution operation on the concatenated fused feature map and output it;
[0011] The nth fusion module is jump-connected with the nth scale feature map, n=1, 2, ..., N, and is used to perform the nth upsampling on the feature map output by the n-1th fusion module, and then concatenate and fuse the nth scale feature map obtained by the jump connection with the feature map output after the nth upsampling, and then perform a convolution operation on the concatenated fused feature map and output it; the output of the Nth fusion module is taken as the fusion result.
[0012] In some implementations, the multi-scale feature map acquisition module includes:
[0013] Based on the outputs of layers of different depths in the same serial processing channel of the same deep neural network module, multiple feature maps of different scales are obtained;
[0014] or,
[0015] Based on the use of multiple encoding branches or feature extraction branches with different parallel processing channels for the input image, feature maps of different scales output by multiple parallel branches are obtained.
[0016] In a second aspect, a fusion method for fusing multi-scale feature maps is provided, comprising the following steps:
[0017] Acquire multiple feature maps of different scales, set the number of feature maps to be N, and sort them from small to large in scale as a first-scale feature map, a second-scale feature map, ..., an N-th-scale feature map;
[0018] Splicing multiple feature maps of different scales acquired by the multi-scale feature map acquisition module according to channels and outputting a first spliced feature map;
[0019] The first concatenated feature map is first upsampled, and then a first scale feature map is obtained by using a skip connection, and the first scale feature map is concatenated and fused with the feature map output after the first upsampling, and then a convolution operation is performed on the concatenated fused feature map and output, which is recorded as a first fused output;
[0020] The first fusion output is subjected to a second upsampling, and then a second scale feature map is obtained by using a skip connection, and the second scale feature map is concatenated and fused with the feature map output after the second upsampling, and then a convolution operation is performed on the concatenated fused feature map and output, which is recorded as the second fusion output;
[0021] By analogy, the n-1th fusion output is upsampled for the nth time, and then the nth scale feature map obtained by skip connection is concatenated with the feature map output after the nth upsampling, and then the concatenated fusion feature map is convolved and output, which is recorded as the nth fusion output, n = 1, 2, ..., N;
[0022] The Nth fusion output is used as the fusion result of the multi-scale feature map.
[0023] In some embodiments, the acquiring of multiple feature maps of different scales includes: obtaining multiple feature maps of different scales based on outputs of layers of different depths in a same serial processing channel of a same deep neural network module; or, obtaining feature maps of different scales output by multiple parallel branches based on encoding branches or feature extraction branches of multiple different parallel processing channels for an input image.
[0024] Thirdly, a semantic segmentation model of land use remote sensing images based on an improved RMSIN model is provided, including:
[0025] The first coding module is improved based on the multi-scale coding module of the RMSIN model, and the units on the natural speech data input and natural speech data processing path in each coding layer in the multi-scale coding module of the RMSIN model are discarded, and the units on the image input and image processing path in each coding layer are retained to form the first coding module;
[0026] A second encoding module, used for further processing the feature maps of multiple scales output by the first encoding module, wherein the second encoding module adopts a cross-scale interaction module of the RMSIN model;
[0027] The decoding module uses the fusion model as described in the first aspect or is used to execute the fusion method as described in the second aspect to process the feature maps of multiple scales output by the second encoding module to obtain the semantic segmentation result of the remote sensing image.
[0028] In some implementations, the first encoding module includes multiple encoding layers, each encoding layer outputs a feature map of a scale, and each encoding layer includes:
[0029] An image input unit is used to obtain a land use remote sensing image or a feature map output by a previous coding layer;
[0030] Transformer layer,
[0031] Multi-receptive field units,
[0032] CSMamba unit,
[0033] Vision door unit,
[0034] The image input by the image input unit is sequentially passed through the transformer layer, the multi-receptive field unit, the CSMamba unit, and the visual gate unit for image feature encoding.
[0035] In a fourth aspect, a training method for a land use remote sensing image semantic segmentation model based on an improved RMSIN model is provided, comprising the following steps:
[0036] Step B1: Collect remote sensing image data and divide the data set: divide the data set into training set, validation set, and test set in a ratio of 6:2:2;
[0037] Step B2: Data enhancement and data annotation: Apply rotations of 0°, 90°, 180°, and 270°, and zooms of 0.75, 1, and 1.25 to the images of the training set, validation set, and test set; use labelme annotation software to annotate the enhanced images to obtain a data set containing different label data, wherein the label data represents the land use type, and output a mask image in png format, where the label of each pixel is the category number;
[0038] Step B3: Constructing a semantic segmentation model of land use remote sensing images based on the improved RMSIN model as described in the third aspect:
[0039] Step B4: training the improved RMSIN model based on the data set to obtain a trained land use remote sensing image semantic segmentation model.
[0040] In a fifth aspect, a semantic segmentation method for land use remote sensing images based on an improved RMSIN model is provided, comprising the following steps:
[0041] Inputting the remote sensing image into a first encoding module for first encoding to obtain feature maps of multiple scales, wherein the first encoding module is obtained by improving the multi-scale encoding module of the RMSIN model, wherein the input of natural speech data and the units on the natural speech data processing path are discarded in each encoding layer in the multi-scale encoding module of the improved RMSIN model, and the units on the image input and image processing path in each encoding layer are retained;
[0042] The cross-scale interaction module of the RMSIN model is used to process the feature maps of multiple scales in sequence to obtain the feature maps of multiple scales after further processing;
[0043] The feature maps of multiple scales are processed using the fusion method described in the second aspect to obtain a fusion result as a semantic segmentation result of the remote sensing image.
[0044] In some implementations, the encoding method of the first encoding module includes:
[0045] Obtain input land use remote sensing images;
[0046] The input image is encoded in multiple encoding layers in sequence, and each encoding layer outputs a feature map of a scale; in each encoding layer, the image is encoded in sequence through the transformer layer, the multi-receptive field unit, the CSMamba unit, and the visual gate unit.
[0047] In a sixth aspect, a computer-readable storage medium is provided, on which computer instructions are stored. When the instructions are executed by a processor, the steps of the above-mentioned fusion method or the steps of the above-mentioned land use remote sensing image semantic segmentation method or the steps of the above-mentioned land use remote sensing image semantic segmentation model training method are implemented.
[0048] The present invention provides a land use remote sensing image semantic segmentation method based on an improved RMSIN model, which has the following features:
[0049] Beneficial effects:
[0050] 1. The fusion model of the present invention can realize multi-scale feature fusion by splicing multiple feature maps of different scales by channel, and up-sampling by combining jump connections, which can effectively restore image details and improve segmentation accuracy. The semantic information is extracted through convolution operation, which enhances the model's ability to understand complex images.
[0051] 2. The semantic segmentation model of land use remote sensing images based on the fusion model and the improved RMSIN model can more effectively handle the diversity and complexity in remote sensing images through multi-scale feature fusion and self-attention mechanism, and improve the segmentation accuracy, especially in the accurate prediction of detailed areas, such as the boundary between water and land. At the same time, the method can also accurately identify the segmentation of different types of objects, such as vast waters to scattered farmland, and improve the processing ability of diversified features. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is a structural diagram of the fusion model in the embodiment of the present application;
[0053] Figure 2 It is a structural schematic diagram of a land use remote sensing image semantic segmentation model based on an improved RMSIN model in an embodiment of the present application;
[0054] Figure 3 is a schematic diagram of the structure in each coding layer in the first coding module in an embodiment of the present application;
[0055] Figure 4 It is a schematic diagram of the structure of the CSMamba module in the embodiment of the present application;
[0056] Figure 5 It is a flow chart of the method for training a semantic segmentation model for land use remote sensing images in an embodiment of the present application. DETAILED DESCRIPTION
[0057] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0058] Exemplary embodiments of the present application are described below.
[0059] Embodiment 1
[0060] The present application embodiment provides a fusion model for fusing multi-scale feature maps, such as Figure 1 As shown, including:
[0061] A multi-scale feature map acquisition module 101 is used to acquire feature maps of multiple different scales. Assuming the number of feature maps is N, they are sorted from small to large in scale and recorded as a first-scale feature map, a second-scale feature map, ..., an N-th-scale feature map;
[0062] A first splicing module 102, configured to splice the multiple feature maps of different scales acquired by the multi-scale feature map acquisition module according to channels and output a first spliced feature map;
[0063] The first fusion module 103 is jump-connected to the first scale feature map, and is used to perform a first upsampling on the first spliced feature map, and then splice and fuse the first scale feature map obtained by the jump connection with the feature map output after the first upsampling, and then perform a convolution operation on the spliced and fused feature map and output it;
[0064] The second fusion module 104 is jump-connected to the second scale feature map, and is used to perform a second upsampling on the feature map output by the first fusion module, and then splice and fuse the second scale feature map obtained by the jump connection with the feature map output after the second upsampling, and then perform a convolution operation on the spliced and fused feature map and output it;
[0065] The nth fusion module is jump-connected with the nth scale feature map, n=1, 2, ..., N, and is used to perform nth upsampling on the feature map output by the n-1th fusion module, and then concatenate and fuse the nth scale feature map obtained by the jump connection with the feature map output after the nth upsampling, and then perform convolution operation on the concatenated fused feature map and output it;
[0066] The output of the Nth fusion module is used as the fusion result.
[0067] In the embodiment of the present application, the fusion model splices multiple feature maps of different scales by channel to achieve multi-scale feature fusion, and combines jump connections for upsampling, which can effectively restore image details, extract semantic information through convolution operations, and enhance the model's ability to understand complex images. The fusion model can obtain more detailed areas and multi-scale features during image processing, innovatively splices, upsamples and fuses features of different scales, effectively overcomes the limitations of single-scale feature utilization, and can fully explore and utilize multi-scale information. This enhances the generalization ability of the model, so that it can accurately capture key features when facing images of various scales. It can be understood that the fusion module provided in the embodiment of the present application can be used in target recognition tasks, in which feature maps of multiple scales are obtained through early feature extraction in target recognition tasks, and then the fusion module of the embodiment of the present application is used to process multi-scale feature maps, and target recognition classification is performed through a classification network. The fusion module provided in the embodiment of the present application can be used in image semantic segmentation tasks, and further, in semantic segmentation tasks for land use remote sensing images, feature maps of multiple scales are obtained through early feature extraction, and then image semantic segmentation is performed using the fusion module of the embodiment of the present application.
[0068] Embodiment 2
[0069] An embodiment of the present application provides a fusion model for fusing multi-scale feature maps, including a multi-scale feature map acquisition module 101, a first splicing module 102, a first fusion module 103, a second fusion module 104, ... an nth fusion module, ..., an Nth fusion module, with the output of the Nth fusion module as the fusion result.
[0070] In an embodiment of the present application, the multi-scale feature map acquisition module 101 is further described. Optionally, the multi-scale feature map acquisition module 101 includes: obtaining a plurality of feature maps of different scales based on the outputs of layers of different depths in the same serial processing channel of the same deep neural network module; or obtaining feature maps of different scales output by a plurality of parallel branches based on encoding branches or feature extraction branches of a plurality of different parallel processing channels respectively for the input image.
[0071] In this embodiment, the multi-scale feature map acquisition module can use multiple different encoding branches or feature extraction branches to obtain feature maps of different scales output by multiple branches for the input image, for example, the NECK network in the YOLOv5 network outputs three feature maps of different scales, and for another example, the input image data and voice data in the RMSIN model are respectively passed through the intra-scale interaction module (IIM) and the cross-scale interaction module (CIM) to obtain four feature maps of different scales. In another embodiment, the multi-scale feature map acquisition module can obtain multiple feature maps of different scales based on layers of different depths in the same deep neural network, such as the bottom-up process on the left side of the FPN feature pyramid, the ResNet network or the FasterRCNN backbone network, etc.
[0072] For other related technical solutions in the embodiments of the present application, please refer to the related description in Embodiment 1, and the same beneficial effects can be achieved. To avoid repetition, they will not be described here.
[0073] Embodiment 3
[0074] The embodiment of the present application provides a fusion method for fusing multi-scale feature maps, comprising the following steps:
[0075] Step A1, obtaining multiple feature maps of different scales, assuming that the number of feature maps is N, and sorting them from small to large in scale as a first-scale feature map, a second-scale feature map, ..., an N-th-scale feature map;
[0076] Step A2, splicing multiple feature maps of different scales acquired by the multi-scale feature map acquisition module according to channels and outputting a first spliced feature map;
[0077] Step A3, performing a first upsampling on the first concatenated feature map, then using a skip connection to obtain a first scale feature map, concatenating and fusing the first scale feature map with the feature map output after the first upsampling, and then performing a convolution operation on the concatenated fused feature map and outputting it, which is recorded as a first fused output;
[0078] Step A4, performing a second upsampling on the first fusion output, and then using a skip connection to obtain a second scale feature map, and concatenating and fusing the second scale feature map with the feature map output after the second upsampling, and then performing a convolution operation on the concatenated fused feature map and outputting it, which is recorded as the second fusion output;
[0079] Step A5, and so on, perform the n-1th fusion output upsampling, then concatenate and fuse the nth scale feature map obtained by skip connection with the feature map output after the nth upsampling, then perform convolution operation on the concatenated fusion feature map and output it, recorded as the nth fusion output, n = 1, 2, ..., N;
[0080] Step A6, taking the fusion result of the multi-scale feature map based on the Nth fusion output.
[0081] The fusion method for fusing multi-scale feature maps provided in the embodiment of the present application and the fusion model embodiment for fusing multi-scale feature maps provided in the above-mentioned embodiment 1 belong to the same concept. The specific implementation process is detailed in the above-mentioned embodiment 1 and will not be repeated here.
[0082] Embodiment 4
[0083] The embodiment of the present application provides a fusion method for fusing multi-scale feature maps, including the method steps of step A1 to step A6 of embodiment 3. In the embodiment of the present application, the above step A1 is further described. In the above step A1, multiple feature maps of different scales are obtained, including: obtaining multiple feature maps of different scales based on the output of layers of different depths in the same serial processing channel of the same deep neural network module; or, obtaining feature maps of different scales output by multiple parallel branches based on the encoding branches or feature extraction branches of multiple different parallel processing channels for the input image.
[0084] The fusion method for fusing multi-scale feature maps provided in the embodiment of the present application and the fusion model embodiment for fusing multi-scale feature maps provided in the above-mentioned embodiment 2 belong to the same concept. The specific implementation process is detailed in the above-mentioned embodiment 2 and will not be repeated here.
[0085] Embodiment 5
[0086] See also Figure 2 The embodiment of the present application provides a land use remote sensing image semantic segmentation model based on an improved RMSIN model, including:
[0087] The first encoding module is improved based on the multi-scale encoding module of the RMSIN model (i.e., the part of the RMSIN model that includes 4 Encoder Layers, each Encoder Layer outputs a scale feature map), and the units on the natural speech data input and natural speech data processing path in each encoding layer in the multi-scale encoding module of the RMSIN model are discarded, and the units on the image input and image processing path in each encoding layer are retained to form the first encoding module;
[0088] The second encoding module uses a cross-scale interaction module of the RMSIN model, namely a CIM module, to further process the feature maps of multiple scales output by the first encoding module;
[0089] The decoding module uses the fusion model described in the first embodiment or the fusion method described in the third embodiment to process the feature maps of multiple scales output by the second encoding module to obtain the semantic segmentation result of the remote sensing image.
[0090] In the embodiment of the present application, the first encoding module borrows the intra-scale interaction module (IIM module) of RMSIN, but improves the IIM module of RMSIN. In the IIM (Intra-scale Interaction Module) module, the multi-receptive field feature extraction within the original single size is retained; the language modality-related part is removed, that is, the cross-modal alignment branch and the language gate are deleted, and it is simplified to single-modal visual feature processing.
[0091] The main task of the modified IIM module is to extract rich feature information from all scales of the image and enhance the spatial perception ability of the model. By effectively enhancing the key information in the image, the IIM module helps the model better capture the details of the object and improve the prediction accuracy of different categories, especially in complex scenes.
[0092] The adjusted RMSIN decoder can effectively restore image details and improve segmentation accuracy by splicing the outputs of the four stages of CSIE by channel, realizing multi-scale feature fusion, and combining jump connections for upsampling. Semantic information is extracted through convolution operations to enhance the model's ability to understand complex remote sensing images. This change makes the decoder more suitable for semantic segmentation tasks, especially when capturing detailed areas and multi-scale features, which can significantly improve the segmentation accuracy and efficiency of the model.
[0093] Optionally, the first encoding module includes multiple encoding layers, each encoding layer outputs a feature map of a scale, such as Figure 3 As shown, each coding layer includes:
[0094] An image input unit is used to obtain a land use remote sensing image or a feature map output by a previous coding layer;
[0095] Transformer Layer,
[0096] Various Receptive Units,
[0097] CSMamba (Channel and Spatial Mamba Block) unit,
[0098] Vision Gate Unit,
[0099] The image input by the image input unit is sequentially passed through the transformer layer, the multi-receptive field unit, the CSMamba unit, and the visual gate unit for image feature encoding.
[0100] In an embodiment of the present application, in the first encoding module, the IIM module is further improved, and a CSMamba unit is added to the retained image input and image processing path, that is, a CSMamba unit is added between the Transformer Layer and the multi-receptive field unit. Therefore, the first encoding module in the embodiment of the present application includes: image input unit, Transformer Layer, Various Receptive, CSMamba unit, and vision gate unit in sequence. The first encoding module in the embodiment of the present application adds CSMamba between Various Receptive and Vision Gate in the improved RMSIN. The CSMamba module combines technologies such as deep convolution, channel and spatial attention, and 2D selection scanning module, and can effectively capture long-distance contextual information with linear time complexity when processing remote sensing images.
[0101] Specifically, the first encoding module includes multiple encoder layers (for example, four encoding layers) connected in sequence, and each encoding layer outputs a feature map of a certain scale. i Indicates that the output features of each encoding layer are expressed as Indicates that Removed language feature F l , only the visual features of the previous encoding layer are retained as input.
[0102] Multi-receptive field part:
[0103] (1) For the input visual features Use multiple branches, and the convolution kernel settings of each branch are different (such as size and step length) to obtain features of different receptive fields.
[0104] (2) Calculate the weight of each pixel using the formula:
[0105]
[0106] σ: Sigmoid function, used to normalize weights; The convolution kernel of the jth convolution branch; Input features.
[0107] (3) Weight w i Used to enhance input features: in, Represents element-wise multiplication.
[0108] CSMamba unit: An efficient feature fusion module that can effectively capture long-range dependencies in remote sensing images while improving multi-scale information fusion and semantic segmentation accuracy by combining channel and spatial attention mechanisms with the linear time complexity of the Mamba framework.
[0109] Visual Gate Unit: further adjusts the enhanced features,
[0110]
[0111] Features are adjusted through activation functions (Tanh and ReLU) and linear projection to supplement local detail information.
[0112] Output: Visual features enhanced by multiple receptive field branches and adjusted by visual gates The output formula is adjusted to
[0113]
[0114] See also Figure 4 In the CSMamba module, the input feature F is processed by two parallel branches. The first branch mainly extracts richer feature information by expanding the number of feature channels and performing deep convolution operations. After expanding the number of channels through the linear layer, the features are extracted using deep convolution and SiLU activation functions, and then processed by the 2D Selective Scanning Module (2D-SSM), and finally normalized by LayerNorm. The purpose of this process is to capture long-range dependencies and effectively extract multi-scale feature information, laying the foundation for subsequent feature fusion and global context modeling. The second branch weights the input features through the channel and spatial attention mechanism (CS), uses the attention module to focus on important areas, and then applies the SiLU activation function to enhance the feature representation, thereby effectively integrating the global and local information of the image.
[0115] Input features: F∈R H×W×C Represents an image feature map with a height of H, a width of W, and a number of channels of C.
[0116] First branch:
[0117] (1) Channel expansion: Through a linear layer, the number of channels of the input feature is expanded to λC (where λ is a predefined expansion factor).
[0118] F'=Linear(F)
[0119] Where: F∈R H×W×C is the input feature map, H is the height, W is the width, and C is the number of channels. H×W×λC It is the expanded feature map, and the number of channels is expanded to λC.
[0120] (2) Depth-Wise Convolution (DWConv): Performs a deep convolution operation. Deep convolution performs convolution independently on each channel, reducing computational complexity and the number of parameters. Compared with standard convolution, the computational complexity of deep convolution is O(C·K 2 ·H·W), while the computational complexity of standard convolution is O(C 2 ·K 2 ·H·W), where K is the convolution kernel corresponding to the input channel F.
[0121] (3) Activation function: Use SiLU (Sigmoid Linear Unit) activation function to enhance feature expression.
[0122] SiLU(F i )=F i ·σ(F i )
[0123] Among them: F i is the feature map after deep convolution (i.e., the value input to the SiLU activation function). i ) is the Sigmoid activation function, that is:
[0124] (4) 2D Selective Scanning Module (2D-SSM): used to process the long-range dependencies of images. It flattens the image features into a 1D sequence and scans in four directions from top left to bottom right, from bottom right to top left, from top right to bottom left, and from bottom left to top right. The scanning in each direction uses the Selective State-Space Model to model long-range dependencies. In this way, it can capture the long-range dependencies in each direction. The scanning results in the four directions (different features of long-range dependencies) are merged through aggregation operations to restore the 2D structure of the image.
[0125] Second branch:
[0126] The channel attention and spatial attention modules (CS) are introduced to improve the representation ability of feature maps. The channel attention module strengthens the focus on important channels through global feature aggregation, while the spatial attention module improves the response to key spatial areas by focusing on the spatial dimension. The combination of the two significantly improves the semantic segmentation effect of remote sensing images, especially in the processing of complex scenes and multi-scale information. The activation function also uses the SiLU activation function for feature conversion.
[0127] The first and second branches merge:
[0128] The output features of the two branches are fused through the Hadamard product (element-wise multiplication). This approach effectively combines two different types of information (one is the global context obtained through 2D selection scanning, and the other is the features focused through channel and spatial attention).
[0129] The final fused features are mapped back to the original number of channels C through the linear layer to obtain the final output feature F out .
[0130] In summary, the overall process formula of CSMamba can be expressed as:
[0131] F1=LN(2D-SSM(SiLU(DWConv(Linear(F)))))
[0132] F2=SiLU(CS(F))
[0133] F out =Linear(F1⊙F2).
[0134] Adding CSMamba between Various Receptive and Vision Gate in the improved RMSIN can effectively enhance the long-distance dependency modeling and information fusion capabilities, and improve the model's ability to handle complex scenes and multi-scale features in remote sensing images. Through the channel and spatial attention mechanism provided by CSMamba, it is possible to accurately capture global context information while maintaining computational efficiency, and fuse features from different scales, thereby improving the accuracy of semantic segmentation, especially when processing detailed areas and complex boundaries.
[0135] Embodiment 6
[0136] See also Figure 5 The present application embodiment provides a method for training a semantic segmentation model of land use remote sensing images based on an improved RMSIN model, comprising the following steps:
[0137] Step B1: Collect remote sensing image data and divide the data set: divide the data set into training set, validation set, and test set in a ratio of 6:2:2;
[0138] Step B2: Data enhancement and data annotation: rotations of 0°, 90°, 180°, and 270°, and scalings of 0.75, 1, and 1.25 were applied to the images of the training set, validation set, and test set; thereby increasing the data set by twelve times; the enhanced images were annotated using labelme annotation software to obtain a data set containing different label data, wherein the label data characterizes the land use type, including four types of labels: water area, residential area, farmland, and forestland; and a mask image in png format was output, wherein the label of each pixel is the category number (for example, 0 represents water area, 1 represents residential area, 2 represents farmland, and 3 represents forestland);
[0139] Step B3: Constructing a semantic segmentation model of land use remote sensing images based on the improved RMSIN model as described in Example 5:
[0140] Step B4: Train the improved RMSIN model based on the data set to obtain a trained land use remote sensing image semantic segmentation model. In this step B4, it includes: during the training process, through cross-validation, use the validation set to evaluate the model performance after each training cycle, and adjust the model parameters according to the evaluation results, and continuously optimize the model performance until the best effect is achieved.
[0141] Embodiment 7
[0142] The embodiment of the present application provides a land use remote sensing image semantic segmentation method based on an improved RMSIN model, comprising the following steps:
[0143] Step C1, inputting the remote sensing image into a first encoding module for first encoding to obtain feature maps of multiple scales, wherein the first encoding module is obtained by improving the multi-scale encoding module of the RMSIN model, wherein the input of natural speech data and the units on the natural speech data processing path are discarded in each encoding layer in the multi-scale encoding module of the improved RMSIN model, and the units on the image input and image processing path in each encoding layer are retained;
[0144] Step C2, using the cross-scale interaction module of the RMSIN model to process the feature maps of multiple scales in sequence, to obtain optimized feature maps of multiple scales;
[0145] Step C3, using the fusion model described in Example 1 to process the feature maps of multiple scales or using the fusion method described in Example 3 to process the feature maps of multiple scales to obtain a fusion result as a remote sensing image semantic segmentation result.
[0146] Furthermore, in step C1, the encoding method of the first encoding module includes:
[0147] Step C11, obtaining input land use remote sensing images;
[0148] Step C12, the input image is encoded in sequence through multiple encoding layers (for example, four encoding layers can be used), and each encoding layer outputs a feature map of a scale; the image is encoded through the transformer layer, the multi-receptive field unit, the CSMamba unit, and the visual gate unit in each encoding layer.
[0149] The land use remote sensing image semantic segmentation method based on the improved RMSIN model in the embodiment of the present application is the same as the concept of the segmentation model described in the above-mentioned embodiment 5, and will not be repeated here.
[0150] According to an embodiment of the present application, the present application also provides a computer-readable storage medium having computer instructions stored thereon, wherein when the instructions are executed by a processor, the steps of the fusion method for fusion of multi-scale feature maps as described in Embodiment 1 or Embodiment 2 are implemented, or, when the instructions are executed by a processor, the steps of the land use remote sensing image semantic segmentation method based on the improved RMSIN model as described in Embodiment 7 are implemented, or the steps of the training method of the land use remote sensing image semantic segmentation model based on the improved RMSIN model as described in Embodiment 6 are implemented. The "computer-readable medium" refers to any computer program product, device, and / or apparatus (e.g., a disk, an optical disk, a memory, a programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor.
[0151] The present invention is not limited to the above-mentioned specific implementation modes. Various changes made by ordinary technicians in this field based on the above-mentioned concepts without creative work are all within the protection scope of the present invention.
Claims
1. A fusion model for fusing multi-scale feature maps, characterized in that: include: A multi-scale feature map acquisition module is used to acquire feature maps of multiple different scales. Assume that the number of feature maps is N, and the feature maps are sorted from small to large in scale as a first-scale feature map, a second-scale feature map, ..., an N-th-scale feature map; A first splicing module, used to splice the multiple different-scale feature maps acquired by the multi-scale feature map acquisition module according to channels and output a first splicing feature map; A first fusion module is jump-connected to the first scale feature map, and is used to perform a first upsampling on the first concatenated feature map, and then concatenate and fuse the first scale feature map obtained by using the jump connection with the feature map output after the first upsampling, and then perform a convolution operation on the concatenated fused feature map and output it; The second fusion module is jump-connected with the second scale feature map, and is used to perform a second upsampling on the feature map output by the first fusion module, and then concatenate and fuse the second scale feature map obtained by the jump connection with the feature map output after the second upsampling, and then perform a convolution operation on the concatenated fused feature map and output it; The nth fusion module is jump-connected with the nth scale feature map, n=1, 2, ..., N, and is used to perform nth upsampling on the feature map output by the n-1th fusion module, and then concatenate and fuse the nth scale feature map obtained by the jump connection with the feature map output after the nth upsampling, and then perform convolution operation on the concatenated fused feature map and output it; The output of the Nth fusion module is used as the fusion result.
2. A fusion model for fusing multi-scale feature maps according to claim 1, characterized in that: include: The multi-scale feature map acquisition module includes: Based on the outputs of layers of different depths in the same serial processing channel of the same deep neural network module, multiple feature maps of different scales are obtained; or, Based on the use of multiple encoding branches or feature extraction branches of different parallel processing channels for the input image, feature maps of different scales output by multiple parallel branches are obtained.
3. A fusion method for fusing multi-scale feature maps, characterized in that: include: Acquire multiple feature maps of different scales, set the number of feature maps to be N, and sort them from small to large in scale as a first-scale feature map, a second-scale feature map, ..., an N-th-scale feature map; Splicing multiple feature maps of different scales acquired by the multi-scale feature map acquisition module according to channels and outputting a first spliced feature map; The first concatenated feature map is first upsampled, and then a first scale feature map is obtained by using a skip connection, and the first scale feature map is concatenated and fused with the feature map output after the first upsampling, and then a convolution operation is performed on the concatenated fused feature map and output, which is recorded as a first fused output; The first fusion output is subjected to a second upsampling, and then a second scale feature map is obtained by using a skip connection, and the second scale feature map is concatenated and fused with the feature map output after the second upsampling, and then a convolution operation is performed on the concatenated fused feature map and output, which is recorded as the second fusion output; By analogy, the n-1th fusion output is upsampled for the nth time, and then the nth scale feature map obtained by skip connection is concatenated with the feature map output after the nth upsampling, and then the concatenated fusion feature map is convolved and output, which is recorded as the nth fusion output, n = 1, 2, ..., N; The Nth fusion output is used as the fusion result of the multi-scale feature map.
4. A fusion method for fusing multi-scale feature maps according to claim 3, characterized in that: The acquiring of multiple feature maps of different scales includes: Based on the outputs of layers of different depths in the same serial processing channel of the same deep neural network module, multiple feature maps of different scales are obtained; or, Based on the use of multiple encoding branches or feature extraction branches of different parallel processing channels for the input image, feature maps of different scales output by multiple parallel branches are obtained.
5. A semantic segmentation model for land use remote sensing images based on an improved RMSIN model, characterized in that: include: The first coding module is improved based on the multi-scale coding module of the RMSIN model, and the units on the natural speech data input and natural speech data processing path in each coding layer in the multi-scale coding module of the RMSIN model are discarded, and the units on the image input and image processing path in each coding layer are retained to form the first coding module; A second encoding module, used for further processing the feature maps of multiple scales output by the first encoding module, wherein the second encoding module adopts a cross-scale interaction module of the RMSIN model; The decoding module uses the fusion model as described in claim 1 or is used to execute the fusion method as described in claim 3 to process the feature maps of multiple scales output by the second encoding module to obtain the semantic segmentation result of the remote sensing image.
6. The land use remote sensing image semantic segmentation model based on the improved RMSIN model according to claim 5, characterized in that: The first encoding module includes multiple encoding layers, each encoding layer outputs a feature map of a scale, and each encoding layer includes: An image input unit is used to obtain a land use remote sensing image or a feature map output by a previous coding layer; Transformer layer, Multi-receptive field units, CSMamba unit, Vision door unit, The image input by the image input unit is sequentially passed through the transformer layer, the multi-receptive field unit, the CSMamba unit, and the visual gate unit for image feature encoding.
7. A training method for a semantic segmentation model of land use remote sensing images based on an improved RMSIN model, characterized in that: include: Step B1: Collect remote sensing image data and divide the data set: divide the data set into training set, validation set, and test set in a ratio of 6:2:2; Step B2: Data enhancement and data annotation: Apply rotations of 0°, 90°, 180°, and 270°, and zooms of 0.75, 1, and 1.25 to the images of the training set, validation set, and test set; use labelme annotation software to annotate the enhanced images to obtain a data set containing different label data, wherein the label data represents the land use type, and output a mask image in png format, where the label of each pixel is the category number; Step B3: Constructing a semantic segmentation model of land use remote sensing images based on the improved RMSIN model as described in any one of claims 5-6: Step B4: training the improved RMSIN model based on the data set to obtain a trained land use remote sensing image semantic segmentation model.
8. A land use remote sensing image semantic segmentation method based on an improved RMSIN model, characterized in that: include: Inputting the remote sensing image into a first encoding module for first encoding to obtain feature maps of multiple scales, wherein the first encoding module is obtained by improving the multi-scale encoding module of the RMSIN model, wherein the input of natural speech data and the units on the natural speech data processing path are discarded in each encoding layer in the multi-scale encoding module of the improved RMSIN model, and the units on the image input and image processing path in each encoding layer are retained; The cross-scale interaction module of the RMSIN model is used to process the feature maps of multiple scales in sequence to obtain the feature maps of multiple scales after further processing; The feature maps of multiple scales are processed using the fusion method as described in claim 3 to obtain a fusion result as a semantic segmentation result of the remote sensing image.
9. The land use remote sensing image semantic segmentation method based on the improved RMSIN model according to claim 8, characterized in that: The encoding method of the first encoding module includes: Obtain input land use remote sensing images; The input image is encoded in multiple encoding layers in sequence, and each encoding layer outputs a feature map of a scale; in each encoding layer, the image is encoded in sequence through the transformer layer, the multi-receptive field unit, the CSMamba unit, and the visual gate unit.
10. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the instructions are executed by the processor, the steps of the method according to any one of claims 3 to 4 are implemented, or the steps of the method according to any one of claims 8 to 9 are implemented, or the steps of the method according to claim 7 are implemented.
Citation Information
Patent Citations
Remote sensing image semantic segmentation method based on self-attention multi-scale feature fusion
CN114419449A
Multisource remote sensing image semantic segmentation method based on Transform, Mama and diffusion model
CN119152205A
Remote sensing image semantic segmentation method combining Unet and Transform
CN116091929A
Deep learning-based multi-modal fusion remote sensing image semantic segmentation method
CN117456182A
Remote sensing image semantic segmentation method based on CNN and ViT mixed structure neural network
CN119181094A
Cited By
Enhanced voice recovery method and device based on multi-modal motion sensor fusion
CN122290615A