Remote sensing image building extraction method and system based on visual Mama model
By introducing multi-directional scanning and directional importance modeling modules into the encoder of the visual Mamba model, the problems of insufficient spatial structure perception and directional modeling for building extraction in remote sensing images are solved, and efficient building extraction and segmentation are achieved, which is suitable for map updating and urban planning.
Patent Information
- Application Number
- CN202510922244.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-09-05
AI Technical Summary
Existing building extraction methods from remote sensing images lack spatial structure perception and directional modeling capabilities, as well as low efficiency in large-format image processing. They are unable to effectively capture the complex spatial dependencies between buildings and their surroundings, resulting in deficiencies in the structural integrity and boundary continuity of the extraction results.
A multi-directional scanning module and a directional importance modeling module are introduced into the encoder of the visual Mamba model. Deep features are extracted by the multi-directional scanning module and weighted fusion is performed through the directional importance modeling module to enhance the modeling ability of building boundaries and structures. Building masks are output using an end-to-end training method.
It improves the accuracy and stability of building extraction in remote sensing images, and can achieve high-quality building segmentation and extraction in large-scale remote sensing images, which is suitable for remote sensing application scenarios such as map updates and urban planning.
Smart Images

Figure CN120599504A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing image processing and computer vision, and in particular relates to a method and system for extracting buildings from remote sensing images based on a visual Mamba model. Background Art
[0002] As the most prominent and central man-made features in urban environments, the spatial structure, distribution density, and morphological information of buildings have important application value in many fields, including urban planning and management, demographics, disaster response, and sustainable development. The task of building extraction can essentially be categorized as a semantic segmentation problem, namely, extracting building pixels from non-building pixels in remote sensing images. In recent years, convolutional neural networks (CNNs) have been widely used for this task, leveraging their advantages in local texture modeling, parameter sharing, and spatial invariance to achieve excellent edge and object extraction results. However, due to the inherent local receptive field limitations of convolution operations, CNNs struggle to model the complex spatial dependencies between buildings and their surroundings in remote sensing images, resulting in significant deficiencies in the extraction results in terms of structural integrity and boundary continuity.
[0003] To improve building extraction accuracy, the visual Transformer has been introduced into remote sensing image analysis tasks. The Transformer uses a self-attention mechanism to model long-range dependencies, demonstrating superior global modeling capabilities in complex background scenes. However, its computational complexity grows quadratically with input size, requiring image cropping or block-based input when processing large-format remote sensing imagery. This not only imposes computational and memory burdens but also easily destroys the spatial consistency of the original image, weakening spatial contextual information, and affecting the stability and generalization of model performance. To address these issues, the recently proposed Mamba architecture, an efficient implementation of the visual Mamba model (SSM), demonstrates Transformer-level global modeling capabilities and near-linear computational efficiency of convolutional architectures. By introducing time-varying parameters and hardware-aware algorithms, Mamba overcomes the static transformation limitations of traditional SSMs, offering enhanced contextual modeling capabilities and scalability. It has become a key research direction in long sequence modeling. However, the Mamba architecture, based on a sequential computation method involving unidirectional, one-dimensional scanning, struggles to effectively capture the complex, multi-directional structure of an image in two-dimensional space.
[0004] To enhance the model's two-dimensional directional modeling capabilities, existing technologies have proposed a variety of scanning expansion methods. For example, horizontal forward and reverse bidirectional scanning is introduced, and it is expanded to four directions: up, down, left, and right. These methods effectively improve the model's receptive field and directional feature modeling capabilities, and achieve good performance in remote sensing image analysis tasks. However, existing methods generally have the following shortcomings: First, there is a lack of modeling of the actual importance of different scanning directions in specific tasks. Features in multiple directions are often simply added together, ignoring the significant characteristics of the directional distribution of buildings in remote sensing images; second, in practical applications for building extraction, the inherent laws between the arrangement direction of buildings and regional structure are not fully utilized, making it impossible to effectively suppress redundant directional interference, which limits further improvement in model extraction accuracy. Summary of the Invention
[0005] The purpose of the present invention is to address the problems of existing remote sensing image building extraction methods in terms of insufficient spatial structure perception and directional modeling capabilities, as well as low large-format image processing efficiency. A remote sensing image building extraction method based on a visual Mamba model is provided. An expandable multi-directional scanning module and a directional importance modeling module are introduced into an encoder based on the visual Mamba model, thereby enhancing the modeling capabilities of building boundaries and structures, and improving the accuracy and stability of building extraction in remote sensing images. The method is suitable for remote sensing application scenarios such as map updating and urban planning, and achieves full modeling and high-quality segmentation of diverse building forms in remote sensing images.
[0006] According to one aspect of the present specification, a method for extracting buildings from remote sensing images based on a visual Mamba model is provided, comprising:
[0007] Acquire remote sensing images to be processed;
[0008] Inputting the remote sensing image to be processed into the trained visual Mamba model to obtain a building extraction result of the remote sensing image; wherein the visual Mamba model training includes:
[0009] Construct a training dataset of remote sensing images to be processed;
[0010] Build a visual Mamba model, including: a preprocessing layer for preprocessing the input remote sensing image training dataset; an encoder, including a multi-directional scanning module for extracting multi-directional depth features of remote sensing images and a directional importance modeling module for weighted fusion of multi-directional depth features of remote sensing images; a decoder for reconstructing the spatial resolution and edge details of remote sensing images and outputting building masks;
[0011] The visual Mamba model is trained on the constructed remote sensing image training dataset to be processed, and a cross entropy loss function is constructed to supervise the building mask, and the trained visual Mamba model is output.
[0012] Furthermore, the working process of the multi-directional scanning module includes:
[0013] Model the multi-directional state of buildings in remote sensing images and configure multi-directional state propagation paths;
[0014] Based on multi-directional state modeling and multi-directional state propagation paths, multi-directional deep features are extracted.
[0015] Furthermore, the working process of the direction importance modeling module includes:
[0016] A trainable attention mechanism is used to model multi-directional deep features, learn the importance weights of deep features in each direction in building extraction, and perform weighted fusion.
[0017] Furthermore, the decoder includes multiple layers of upsampling units, convolutions, and skip connections:
[0018] The layer upsampling unit and convolution are used to improve the spatial resolution of the feature map layer by layer; the jump connection is used to fuse shallow high-resolution features with deep semantic features.
[0019] Furthermore, the working process of the preprocessing layer includes:
[0020] Perform channel rearrangement, normalization, size alignment and padding operations on the remote sensing image to be processed to obtain a remote sensing image in tensor format;
[0021] Perform image segmentation on the remote sensing image in tensor format to obtain multiple image blocks;
[0022] Each image patch is mapped to a low-dimensional embedding vector via linear projection.
[0023] Furthermore, the construction of the visual Mamba model also includes:
[0024] The segmentation head is used to convert the building mask output by the decoder into a pixel-level binary building mask.
[0025] Furthermore, the method further includes:
[0026] The multi-directional scanning module supports multiple scanning strategies, each of which is combined with an image transformation operation to generate multiple sequence inputs in different directions, thereby expanding the spatial coverage of state modeling.
[0027] According to one aspect of the present specification, a system for extracting buildings from remote sensing images based on a visual Mamba model is provided, comprising:
[0028] A data acquisition module is used to acquire remote sensing images to be processed;
[0029] The remote sensing image building extraction module is used to input the remote sensing image to be processed into a trained visual Mamba model to obtain a remote sensing image building extraction result; wherein the visual Mamba model training includes:
[0030] Construct a training dataset of remote sensing images to be processed;
[0031] Build a visual Mamba model, including: a preprocessing layer for preprocessing the input remote sensing image training dataset; an encoder, including a multi-directional scanning module for extracting multi-directional depth features of remote sensing images and a directional importance modeling module for weighted fusion of multi-directional depth features of remote sensing images; a decoder for reconstructing the spatial resolution and edge details of remote sensing images and outputting building masks;
[0032] The visual Mamba model is trained on the constructed remote sensing image training dataset to be processed, and a cross entropy loss function is constructed to supervise the building mask, and the trained visual Mamba model is output.
[0033] According to one aspect of the present specification, an electronic device is provided, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method for extracting buildings from remote sensing images based on a visual Mamba model when executing the computer program.
[0034] According to one aspect of the present specification, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for extracting buildings from remote sensing images based on the visual Mamba model are implemented.
[0035] Compared with the prior art, the present invention has the following beneficial effects:
[0036] 1. The present invention introduces an extensible multi-directional scanning module into the encoder based on the visual Mamba model, supports flexible configuration of scanning modes and image transformation strategies, and enhances the building directional structure modeling capability;
[0037] 2. The present invention introduces a directional importance modeling module into the encoder based on the visual Mamba model, which has automatic direction perception capability and can adaptively select scanning directions that are highly relevant to building structures, thereby improving the model's responsiveness to features such as building boundaries and directional arrangements.
[0038] 3. The model proposed in this invention has a compact structure, high modeling efficiency, and end-to-end training and reasoning capabilities. It does not require additional post-processing modules, can enhance the expression of building boundaries and structures, improve the accuracy and geometric consistency of building extraction in large-scale remote sensing images, is easy to deploy and migrate, and is suitable for automatic building extraction and structural analysis tasks in large-scale remote sensing images. It has good practical value and promotion prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0040] Figure 1 Schematic diagram of a method flow in an embodiment of the present invention;
[0041] Figure 2 This is a schematic diagram of the visual Mamba model training according to an embodiment of the present invention;
[0042] Figure 3 This is a schematic diagram of the structure of a visual Mamba model according to an embodiment of the present invention;
[0043] Figure 4 This is a schematic diagram of the building extraction effect according to an embodiment of the present invention. DETAILED DESCRIPTION
[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0045] like Figure 1-2As shown, an embodiment of the present invention provides a method for extracting buildings from remote sensing images based on a visual Mamba model, comprising: obtaining a remote sensing image to be processed; inputting the remote sensing image to be processed into a trained visual Mamba model to obtain a building extraction result from the remote sensing image; wherein the visual Mamba model training comprises: constructing a training dataset of the remote sensing image to be processed; building a visual Mamba model comprising: a preprocessing layer for preprocessing the input training dataset of the remote sensing image to be processed; an encoder for performing state modeling on the preprocessed remote sensing image; a decoder for reconstructing the spatial resolution and edge details of the remote sensing image; and a cross-entropy loss function for supervising a building mask; training the visual Mamba model on the constructed training dataset of the remote sensing image to be processed, constructing a cross-entropy loss function to supervise the building mask, and outputting the trained visual Mamba model.
[0046] Specifically, the present invention acquires and preprocesses remote sensing images. The acquired remote sensing images are high-resolution remote sensing image data, which can be three-channel RGB or expanded to four-channel images including the near-infrared band. First, through standardization processing such as channel reordering (e.g., BGR to RGB), normalization, and size cropping (padding operation), a uniformly formatted tensor input is generated, with the shape:
[0047] (1)
[0048] in, is the image, b is the batch size, c is the number of channels, H and W are the height and width of the image respectively. The input tensor is used for sequence modeling in the encoder based on the visual Mamba model.
[0049] Specifically, the embodiment of the present invention constructs an encoder, such as Figure 3 As shown in the figure, a continuous state space modeling approach is used to transform the input image sequence into a latent state propagation process, and global modeling of large-format images is achieved through efficient linear dynamic scanning. Compared with traditional convolutional and self-attention structures, this structure has lower computational complexity while maintaining modeling capabilities. The preprocessed image is divided into fixed-size image blocks, each of size P×P. Each image block is embedded into a vector of dimension d through a linear embedding module, converting the two-dimensional image structure into a one-dimensional vector sequence, as shown below:
[0050] (2)
[0051] in, After linear embedding, each image block is represented as a d-dimensional real vector, which is used to represent the semantic features of the image block in the embedding space. As the input sequence of the visual Mamba model, it is input into the visual Mamba encoding module.
[0052] Specifically, an embodiment of the present invention introduces a multi-directional scanning module in an encoder based on the visual Mamba model, which is used to perform state modeling of building structures in remote sensing images along multiple spatial directions. By configuring multi-directional state propagation paths, deep feature representations with direction perception capabilities are extracted.
[0053] Specifically, the core encoder adopts the Mamba structure based on state space modeling and establishes the state space equation as follows:
[0054] (3)
[0055] (4)
[0056] in, is hidden state, For input, is the output, and A, B, C, and D are the model's learnable parameters. To achieve efficient calculation, discretization is introduced, converting the above continuous equation into:
[0057] (5)
[0058] in, The proposed method combines a parameterized matrix with a HiPPO projection kernel and a convolution kernel parameter mapping strategy to achieve efficient one-dimensional state propagation. The traditional Mamba architecture only supports one-way state sequence modeling and cannot effectively adapt to the multi-directional arrangement, varying scales, and complex structures of buildings in remote sensing images. This embodiment of the present invention introduces a multi-directional scanning module that, by designing multiple sets of state propagation paths, enables sequence encoding and feature modeling of images in different directions.
[0059] Specifically, for the multi-directional scanning module, the embodiment of the present invention supports setting scanning paths in four basic directions: horizontal, vertical, main diagonal, and sub-diagonal. It also has good scalability: through simple transformations such as image transposition, rotation, and flipping, a single scanning mode (such as Sweep) can be expanded to eight different scanning sequences to simulate state propagation under multiple perspectives. In addition, other classic scanning strategies such as Z-order curves can be configured to achieve free switching from global modeling to local fine-grained modeling. Each direction corresponds to an independent Mamba block, which outputs a direction-aware deep feature tensor. ,as follows:
[0060] (6)
[0061] Specifically, in order to improve the discrimination ability of multi-directional feature fusion, the embodiment of the present invention introduces a directional importance modeling module in the encoder based on the visual Mamba model. The multiple directional depth features obtained by scanning are input into the directional importance modeling module, and a trainable directional weight mechanism is used to adaptively learn the importance of each directional channel. , weighted summation of multi-channel features is performed to generate a fusion context representation with directional structure perception capability, namely fusion features. The expression is:
[0062] (7)
[0063] The final output fusion feature expression is:
[0064] (8)
[0065] in, is the number of scanning directions, For the Feature expression in each direction.
[0066] Specifically, the directional importance modeling module includes a directional weight normalization unit and a weighted fusion unit, which supports automatic learning of the response intensity in each direction according to the arrangement rules of buildings in different remote sensing images, enhances the feature expression consistent with the structural direction, and suppresses redundant directional interference.
[0067] Specifically, this embodiment of the present invention inputs the fused features into a decoder, uses a pixel-level semantic segmentation structure to predict building areas, and outputs a probability map of the same size as the input image. Finally, the output map is thresholded (binarized), with values greater than 0.5 set to 1 and otherwise set to 0, resulting in the final building mask. The decoder employs a symmetrical upsampling structure to restore spatial resolution and preserve boundary information. The entire model supports end-to-end training and inference, and the results can be directly used for building outline extraction, vector conversion, and downstream building modeling applications.
[0068] Furthermore, the decoder includes multi-layer convolution and upsampling units, which are combined with a skip connection structure to restore spatial dimensions and enhance edge responsiveness. The output building mask is a single-channel binary image that can be used for direct spatial projection or vectorization processing.
[0069] Specifically, an embodiment of the present invention further provides a segmentation head after the decoder, which is composed of a set of convolutional layers and upsampling operations, and further processes the high-order semantic features output by the decoder into a pixel-level binary building mask. Specifically, the segmentation head first performs spatial refinement and channel compression on the fused decoded features through multiple convolutional layers, and then restores the feature map to the same spatial resolution as the original input image through an upsampling operation (bilinear interpolation). Finally, convolution is used to generate a single-channel feature map, and combined with the Sigmoid activation function, the probability value of each pixel belonging to a building is output to form the final building extraction result map. The size is consistent with the input image, which is convenient for superposition with the original image, subsequent analysis or vectorization processing.
[0070] Specifically, semantic mapping is completed on the basis of maintaining the consistency of spatial size. The segmentation head converts the feature expression output by the decoder into pixel-level prediction results of buildings. It is a key output module for achieving accurate positioning and extraction of buildings.
[0071] Specifically, the embodiment of the present invention is deployed in a mainstream deep learning framework for training and inference. The loss function includes a cross-entropy loss function, which supports multi-scale supervision and direction-selective training optimization strategies to further improve the boundary integrity and geometric consistency of the building extraction results. The embodiment of the present invention adopts an end-to-end training method to construct a training pipeline under the Pytorch framework. The loss function uses a standard binary cross-entropy loss function to supervise the building mask. During the training process, all directional modeling and directional weighting are jointly optimized as part of the overall network, without the need to introduce additional post-processing or directional rule guidance modules. By fully learning the semantic contribution of each directional feature on the building directional data, better building extraction accuracy and generalization capabilities can be achieved.
[0072] Specifically, in order to verify the effectiveness and advancement of the present invention in extracting buildings from remote sensing images, the WHU building dataset is selected as the experimental benchmark data and compared with other advanced deep learning building extraction methods, including the U-Net segmentation model proposed by Ronneberger et al., document [1]: Ronneberger O, Fischer P, Brox TU-Net: Convolutional Networks for Biomedical Image Segmentation [C] / / Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015. Springer, 2015: 234–241; the deep convolutional network ResNet based on residual connection proposed by He et al., document [2]: He K, Zhang X, Ren S, et al. Identity Mappings in Deep Residual Networks [EB / OL]. arXiv: 1603.05027, published on April 12, 2016; the Swin-Unet model proposed by Cao et al., combined with Swin Transformer and U-Net structures achieve efficient image modeling. Reference [3]: Cao H, Wang Y, Chen J, et al. Swin-Unet: Unet-Like Pure Transformer for Medical Image Segmentation [C] / / Computer Vision – ECCV 2022 Workshops. Springer, published on May 12, 2021; Zhu et al. proposed the Samba method, which introduced the visual Mamba model into the remote sensing semantic segmentation task for the first time. Reference [4]: Zhu Q, Zhao Y, He L, et al. Samba: Semantic segmentation of remotely sensed images with state space model [J]. Heliyon, 2024, 10(19): e38495. published on September 26, 2024. All methods were trained and tested under the same hardware environment. The experimental environment was a high-performance computing platform equipped with NVIDIA RTX A6000 GPU. The experiment uses four typical indicators for quantitative evaluation, including precision, recall, F1 value and intersection-over-union ratio.Among them, the F1 value comprehensively reflects the balance between precision and recall, while the IoU indicator is used to measure the spatial overlap between the predicted mask and the true mask, and is an important reference for evaluating performance in building extraction tasks.
[0073] Specifically, the quantitative evaluation results of the method proposed in the embodiment of the present invention and the existing methods on the WHU building dataset are shown in Table 1. From the experimental results, it can be seen that the method of the present invention has achieved significantly better performance than other methods in all indicators. Among them, the improvement in the two key indicators of F1 value and IoU is the most significant, indicating that the method has stronger expression ability in boundary recognition and area coverage of building outlines. In addition, Figure 4 Figure 2 shows the building extraction results of the method proposed in this embodiment of the present invention and other methods. As can be seen from the above comparison, the multi-directional adaptive remote sensing Mamba method proposed in this invention can effectively capture the directional structural information of buildings in remote sensing images. While maintaining high-precision extraction capabilities, it also has good model robustness and practical value. It is particularly suitable for intelligent building extraction tasks in large-scale and diverse terrains.
[0074] Table 1 Comparison of building extraction effects between the method of the present invention and other methods (%)
[0075] method Accuracy Recall F1 value Intersection and Union References[1] 94.50 90.88 92.65 86.31 References[2] 94.21 94.76 94.48 89.54 References[3] 94.39 93.72 94.06 88.79 References[4] 94.88 93.31 94.09 88.84 Method of the present invention 95.37 95.36 95.37 91.15
[0076] The implementation basis of each embodiment of the present invention is achieved through programmed processing by a device with processor functionality. Therefore, in practical engineering, the technical solutions and functions of each embodiment of the present invention are encapsulated into various modules. Based on this reality, and in addition to the above-mentioned embodiments, an embodiment of the present invention provides a system for extracting buildings from remote sensing images based on a visual Mamba model. This system is used to implement a method for extracting buildings from remote sensing images based on a visual Mamba model as described in the above-mentioned method embodiments.
[0077] The system includes: a data acquisition module for acquiring a remote sensing image to be processed; a remote sensing image building extraction module for inputting the remote sensing image to be processed into a trained visual Mamba model to obtain a remote sensing image building extraction result; wherein the visual Mamba model training includes: constructing a training dataset of the remote sensing image to be processed; building a visual Mamba model including: a preprocessing layer for preprocessing the input training dataset of the remote sensing image to be processed; an encoder including a multi-directional scanning module for extracting multi-directional depth features of the remote sensing image and a directional importance modeling module for weightedly fusing the multi-directional depth features of the remote sensing image; a decoder for reconstructing the spatial resolution and edge details of the remote sensing image and outputting a building mask; training the visual Mamba model on the constructed training dataset of the remote sensing image to be processed, constructing a cross entropy loss function to supervise the building mask, and outputting the trained visual Mamba model.
[0078] The remote sensing image building extraction system based on the visual Mamba model provided by the embodiments of the present invention addresses the problems of existing remote sensing image building extraction methods, such as insufficient spatial structure perception and directional modeling capabilities, and low large-format image processing efficiency. By adopting several modules, an expandable multi-directional scanning module and a directional importance modeling module are introduced into an encoder based on the visual Mamba model. This enhances the modeling capabilities of building boundaries and structures, and improves the accuracy and stability of building extraction in remote sensing images. The system is suitable for remote sensing application scenarios such as map updates and urban planning, and achieves comprehensive modeling and high-quality segmentation of diverse building forms in remote sensing images.
[0079] Based on the same inventive concept as the aforementioned embodiment, an embodiment of the present invention further provides an electronic device, including a memory and a processor, wherein the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement a remote sensing image building extraction method based on a visual Mamba model as proposed in the aforementioned embodiment.
[0080] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program. When executed by a processor, this program overcomes existing methods' shortcomings, such as insufficient spatial structure perception and directional modeling capabilities, as well as low processing efficiency for large-format images. It enhances the modeling capabilities of building boundaries and structures, and improves the accuracy and stability of building extraction from remote sensing images. The program is suitable for remote sensing applications such as map updates and urban planning, achieving comprehensive modeling and high-quality segmentation of diverse building forms in remote sensing images.
[0081] The storage medium can be any non-volatile storage device such as a hard disk, solid-state drive, flash drive, optical disk, etc., which is used to store computer program code and necessary data files. The stored computer program includes: a data acquisition module and a remote sensing image building extraction module.
[0082] In summary, the present invention provides a method and system for extracting buildings from remote sensing images based on the visual Mamba model. This method preprocesses and segments remote sensing images, constructs an encoder using the visual Mamba model, introduces a multi-directional scanning module to extract directional features, and fuses this multi-directional information through a directional importance adaptive mechanism to ultimately generate building masks. This method possesses end-to-end modeling capabilities, enhances the representation of building boundaries and structures, and improves the accuracy and geometric consistency of building extraction in large-scale remote sensing images, demonstrating its practical value.
[0083] Finally, it should be noted that the above specific embodiments are merely representative examples of the present invention. Obviously, the present invention is not limited to the above specific embodiments and is susceptible to numerous variations. Any simple modifications, equivalent variations, and modifications to the above specific embodiments based on the technical essence of the present invention shall be deemed to fall within the scope of protection of the present invention.
Claims
1. A method for extracting buildings from remote sensing images based on a visual Mamba model, characterized in that: include: Acquire remote sensing images to be processed; Inputting the remote sensing image to be processed into the trained visual Mamba model to obtain a building extraction result of the remote sensing image; wherein the visual Mamba model training includes: Construct a training dataset of remote sensing images to be processed; Build a visual Mamba model, including: a preprocessing layer for preprocessing the input remote sensing image training dataset; an encoder, including a multi-directional scanning module for extracting multi-directional depth features of remote sensing images and a directional importance modeling module for weighted fusion of multi-directional depth features of remote sensing images; a decoder for reconstructing the spatial resolution and edge details of remote sensing images and outputting building masks; The visual Mamba model is trained on the constructed remote sensing image training dataset to be processed, and a cross entropy loss function is constructed to supervise the building mask, and the trained visual Mamba model is output.
2. The method for extracting buildings from remote sensing images based on a visual Mamba model according to claim 1, wherein: The working process of the multi-directional scanning module includes: Model the multi-directional state of buildings in remote sensing images and configure multi-directional state propagation paths; Based on multi-directional state modeling and multi-directional state propagation paths, multi-directional deep features are extracted.
3. The method for extracting buildings from remote sensing images based on a visual Mamba model according to claim 1, wherein: The working process of the directional importance modeling module includes: A trainable attention mechanism is used to model multi-directional deep features, learn the importance weights of deep features in each direction in building extraction, and perform weighted fusion.
4. The method for extracting buildings from remote sensing images based on a visual Mamba model according to claim 1, wherein: The decoder includes multiple layers of upsampling units, convolutions, and skip connections: The layer upsampling unit and convolution are used to improve the spatial resolution of the feature map layer by layer; the jump connection is used to fuse shallow high-resolution features with deep semantic features.
5. The method for extracting buildings from remote sensing images based on a visual Mamba model according to claim 1, wherein: The working process of the preprocessing layer includes: Perform channel rearrangement, normalization, size alignment and padding operations on the remote sensing image to be processed to obtain a remote sensing image in tensor format; Perform image segmentation on the remote sensing image in tensor format to obtain multiple image blocks; Each image patch is mapped to a low-dimensional embedding vector via linear projection.
6. The method for extracting buildings from remote sensing images based on a visual Mamba model according to claim 1, wherein: The construction of the visual Mamba model also includes: The segmentation head is used to convert the building mask output by the decoder into a pixel-level binary building mask.
7. The method for extracting buildings from remote sensing images based on a visual Mamba model according to claim 1, wherein: The method further comprises: The multi-directional scanning module supports multiple scanning strategies, each of which is combined with an image transformation operation to generate multiple sequence inputs in different directions, thereby expanding the spatial coverage of state modeling.
8. A remote sensing image building extraction system based on the visual Mamba model, characterized in that: include: A data acquisition module is used to acquire remote sensing images to be processed; The remote sensing image building extraction module is used to input the remote sensing image to be processed into a trained visual Mamba model to obtain a remote sensing image building extraction result; wherein the visual Mamba model training includes: Construct a training dataset of remote sensing images to be processed; Build a visual Mamba model, including: a preprocessing layer for preprocessing the input remote sensing image training dataset; an encoder, including a multi-directional scanning module for extracting multi-directional depth features of remote sensing images and a directional importance modeling module for weighted fusion of multi-directional depth features of remote sensing images; a decoder for reconstructing the spatial resolution and edge details of remote sensing images and outputting building masks; The visual Mamba model is trained on the constructed remote sensing image training dataset to be processed, and a cross entropy loss function is constructed to supervise the building mask, and the trained visual Mamba model is output.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the remote sensing image building extraction method based on the visual Mamba model according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for extracting buildings from remote sensing images based on the visual Mamba model according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Multisource remote sensing image semantic segmentation method based on Transform, Mama and diffusion model
CN119152205A
Instance segmentation method and device of remote sensing image and electronic equipment
CN119273700A
Intelligent building extraction method based on remote sensing image
CN119445380A
Remote sensing image semantic change detection method and device based on Mamba model
CN119580258A
Remote sensing reference image segmentation method based on comparative learning and cross attention
CN119785348A