Remote Sensing Image Building Extraction Method Based on Multi-Scale Interaction and Cross Decoding
The MGCC network addresses the challenge of multi-scale feature fusion and semantic alignment in remote sensing images by using MSAG and CLCC modules to improve building extraction accuracy in complex scenarios.
Patent Information
- Application Number
- CN202411654545.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-19
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-11-19
AI Technical Summary
The existing remote sensing image building extraction method based on encoder-decoder architecture has information redundancy and semantic gaps in the multi-scale feature fusion and decoding process, resulting in insufficient building extraction accuracy, especially in complex scenarios, it is difficult to effectively distinguish between buildings and backgrounds.
The multi-scale interaction and cross-decoding method is adopted to shorten the semantic gap between multi-level features through the multi-scale interaction guidance module (MSAG), and optimize the context fusion of the decoder through the cross-layer cross-decoding module (CLCC), enhancing the semantic information capture capability of the model to the building.
It improves the accuracy and accuracy of building extraction in remote sensing images, especially in complex scenarios, which can better identify building boundaries and shapes, suppress background interference information, and achieve higher precision building extraction.
Smart Images

Figure CN119540767B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of remote sensing image information extraction, and relates to a method for extracting buildings from remote sensing images based on multi-scale interaction and cross-decoding. Background Technique
[0002] With the development of remote sensing sensors and earth observation technologies, it is now possible to quickly obtain large-scale high-quality image data. Compared with traditional low-resolution images, very high-resolution (VHR) remote sensing images contain more details of ground objects, spatial textures, and semantic information, which are convenient for accurately identifying ground objects such as buildings and roads. Therefore, they are widely used in fields such as urban management, smart cities, and virtual reality. However, building extraction based on VHR images faces many challenges. In addition to buildings, VHR images also contain a large amount of redundant information such as trees and roads, which affects the extraction accuracy. In addition, the roof materials, shapes, sizes, and spectral characteristics of buildings are diverse and are easily confused with background ground objects, further increasing the extraction difficulty.
[0003] In recent years, significant progress has been made in the application of deep learning methods in building extraction. Building extraction based on deep learning belongs to the semantic segmentation task, in which the encoder-decoder architecture (such as FCN) provides a basis for image segmentation and can effectively restore the spatio-temporal information of images. As typical encoder-decoder networks, U-Net and SegNet have remarkable effects when processing VHR remote sensing images, but they are still difficult to handle target extraction in complex scenes. For this reason, the attention mechanism is introduced to improve the extraction ability of the model in complex environments. For example, SER-UNet combines the attention mechanism and skip connections, effectively improving the extraction accuracy of building edges; MSSDMPA-Net accurately extracts building and road footprints by introducing a dynamic attention module. In addition, multi-scale feature learning helps to improve the sensitivity of the model to objects of different sizes and enhances its adaptability in complex environments. MAP-Net optimizes feature fusion by learning and retaining multi-scale features through parallel paths; MDCGA-Net improves the effects of global and local feature extraction by introducing directional attention. The utilization of global context information further enhances the semantic understanding ability of the model. DPENet combines the dual-path feature extraction of CNN and Transformer, enhancing the semantic information extraction ability. Based on the above technical progress, the method for extracting buildings from remote sensing images based on deep learning is constantly evolving, and the accuracy and efficiency of building extraction are improved through the combination of multiple mechanisms.
[0004] However, the building extraction framework based on the encoder-decoder architecture still has some limitations in multi-scale feature fusion at the decoder stage. Specifically, on the one hand, due to the lack of interaction between different levels, not just the interaction between the high and low levels, directly fusing high-level and low-level features may lead to information redundancy and conflicts, affecting the accuracy of building extraction. On the other hand, in the context fusion process of high-level and low-level features, previous building extraction methods use the attention mechanism to focus on building features, but easily ignore their semantic similarity, which will cause the model to fail to fully focus on the key semantic features of buildings, thus reducing the accuracy of the extraction results. In summary, there is an urgent need for an optimization method to bridge the semantic gap between these levels, enable more reasonable fusion of multi-scale information during the decoding process, and optimize the semantic capture ability of decoding. Especially for the need of building extraction in remote sensing images, it is necessary to design modules that can gradually refine and strengthen building features, which can not only retain semantic consistency but also suppress interference information, and finally achieve higher-precision extraction of buildings. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a method for extracting buildings from remote sensing images based on multi-scale interaction and cross-decoding, so as to improve the accuracy of building extraction from remote sensing images.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A method for extracting buildings from remote sensing images based on multi-scale interaction and cross-decoding specifically includes the following steps:
[0008] S1: Obtain a remote sensing image and label it, divide it into a training set, a validation set, and a test set, and perform image data augmentation when the image is input.
[0009] S2: Construct an MGCC model, including an MSAG module and a CLCC module, where MGCC represents a building extraction network based on multi-scale interaction and cross-decoding, MSAG represents a multi-scale interaction guidance module, and CLCC represents a cross-layer cross-decoding module;
[0010] S3: Train the MGCC model: Input the training set into the MGCC model, calculate the corresponding loss function, and then use the AdamW optimizer and the cosine restart strategy to iteratively train the MGCC model until the model converges;
[0011] S4: Input the test set into the trained MGCC model to obtain an image with only building features.
[0012] Further, in step S1, the image enhancement includes random flipping and random rotation.
[0013] Furthermore, in step S2, the constructed MGCC model adopts an encoder-decoder architecture. The encoder part uses the BFB Vit proposed in BuildFormer to capture global multi-scale information. At the same time, ASPP (Atrous Spatial Pyramid Pooling) is added to the last layer of feature extraction to form the Backbone of the model. The output features from top to bottom are F1, F2, F3, F as ; The encoder process is expressed as:
[0014] F1, F2, F3, F as = Backbone(X) (1)
[0015] The decoder part includes an MSAG module and a CLCC module.
[0016] Furthermore, in step S2, the MSAG module is set after the encoder performs feature extraction, and is used to shorten the semantic gap between multi-level features and embed multi-scale information into the features; the MSAG module includes three parts: Multi-Scale Fusion, Position Interactive, and Channel Interactive;
[0017] Specifically, the multi-scale fusion is as follows: for the extracted feature C represents the channel dimension number of the feature map, H represents the height of the feature map, and W represents the width of the feature map. They are concatenated to obtain represents the concatenated feature of the i-th layer. Subsequently, The channels are compressed to C through a 1×1 convolution i (i = 1, 2, 3), and the compressed features are fed into 3 parallel depthwise separable convolutions with kernel sizes of 3×3, 5×5, and 7×7 respectively. The results of the parallel outputs are added to obtain the multi-scale fusion result The above process is expressed as:
[0018] FMS i = B(Conv 3×3 (F ci )) + B(Conv 5×5 (F ci )) + B(Conv 7×7 (F ci )) (2)
[0019] where B(·) represents BatchNorm; Conv k×kDenote the convolution operation with a convolution kernel of k, and perform this operation on the first three layers of the encoder to obtain three multi-scale embedded features FMS1, FMS2, and FMS3;
[0020] The specific position interaction is as follows: After multi-scale fusion, subsequent spatial position refinement is performed. First, along the horizontal and vertical directions respectively, for the input feature FMS i Apply global average pooling and concatenate the output feature maps to generate The above process is expressed as follows:
[0021] F hwi = [GAP h (FMS i ), GAP w (FMS i )] (3)
[0022] Among them, GAP h (·) represents global average pooling along the horizontal direction, and GAP w (·) represents global average pooling along the vertical direction; [,] represents the concatenation operation, and then apply a 1×1 convolution for channel reduction to obtain Capture the information from the horizontal and vertical coordinates. r is the multiple of channel reduction. The above process is expressed as:
[0023] F rhwi = BR(Conv(F hwi )) (4)
[0024] Among them, BR(·) represents the BatchNorm and ReLU operations, and Conv represents a 1×1 convolution; then split F rhwi along the horizontal and vertical directions (Split) to obtain F hi and F wi , and apply a 1×1 convolution to F hi and F wi to restore the channels to C i Then activate with the Sigmod function to obtain the direction weights W hi and W wi along the width and height dimensions. Finally, multiply this weight with FMS i to obtain the attention feature F pi with spatial position information interaction. The above process is expressed as:
[0025] F pi = FMS i σ(Conv(F hi )) * σ(Conv(F wi )) (5)
[0026] Among them, σ represents the Sigmoid activation function;
[0027] The specific channel interaction is as follows: In the first step of multi-scale fusion, 1×1 is used to compress the channels, which results in a certain degree of loss of information from each layer. To make up for this shortcoming through channel interaction, first use global average pooling, and perform 1×1 convolution and ReLU activation to generate channel attention features, adaptively filter out redundant channel features, and then expand the channels to the output size through 1×1 convolution and multiply with the attention feature F with spatial position information interaction to obtain the complete multi-scale interaction output F pi The above process is expressed as: mpci ,
[0028] F mpci =F pi *Conv(ReLU(Conv(GAP(FMS i )))) (6)
[0029] Therefore, the expression of MSAG is defined as:
[0030] X = MSAG(X) (7)
[0031] Furthermore, in step S2, the structure of the CLCC module is as follows: The features from the low layer and the high layer are sent into the multi-scale strip convolution block for multi-scale alignment and generate the relevant Q, K, V features, and then cross-sent into the multi-head self-attention to capture the similar semantic information of the building and perform global long-range dependence perception. At the same time, in order to enhance the local detail attention ability, the C-MLP module is added, where C-MLP represents the convolutional multi-layer perceptron;
[0032] The specific calculation process is as follows: Denote the feature from the low layer as F l , Denote the feature from the high layer as F h , First, perform upsampling on F h and achieve channel alignment through 1×1 convolution to match F l , and the above process is expressed as:
[0033] F h =Conv(Up(F hwi )) (8)
[0034] Subsequently, send F l and F h into the multi-scale strip convolution block MSSC respectively. For the input feature X, X ∈ RC ×H×W , define MSSC as follows:
[0035] Q, K, V = MSSC(X) = QKV(SC7(X)+SC11 (X) + SC 21 (X)) (9)
[0036] Among them, QKV(X) represents obtaining three outputs by performing three 1×1 convolutions on X, and SC k represents a set of serial 1×k and k×1 strip convolutions; for F l and F h The outputs obtained after feeding them into MSSC are expressed as follows:
[0037] Q l , K l , V l = MSSC(BN(F l )) (10)
[0038] Q h , K h , V h = MSSC(BN(F h )) (11)
[0039] After obtaining the Q, K, and V matrices from the low-level features and high-level features, the Q matrix is cross-fed into the self-attention module to obtain SA l and SA h , and then a residual connection is performed to obtain the output F sa , where the self-attention module is expressed as:
[0040]
[0041] The above process is expressed as:
[0042] F sa = SA(Q h , K l , V l ) + SA(Q l , K h , V h ) + F l + F h (13)
[0043] Finally, F sa is fed into C-MLP to obtain the final output D, and the above process is expressed as:
[0044] C_MLP(X) = Conv(DwConv 3×3 (ReLU(Conv(X)))) (14)
[0045] D = F sa + C MLP (BN(F sa )) (15)
[0046] Among them, DwConv 3×3 represents depthwise separable convolution with a 3×3 convolution kernel;
[0047] Therefore, the expression of CLCC is defined as:
[0048] X = CLCC(F l , F h ) (16)
[0049] Furthermore, in step S2, after the MGCC model refines the results layer by layer, an output is obtained, and the network includes a multi-scale output for deep supervision; the last layer decoding SegHead is expressed as:
[0050] SegHead(X) = Conv(Up(Dropout(CBR(X)))) (17)
[0051] Among them, CBR represents Conv 3×3 , BatchNorm, ReLU, Up represents upsampling, Conv represents 1×1 convolution, and its output channels are the number of predicted classes. For the input image X, the overall network flow is expressed as:
[0052] F1, F2, F3, F as = Backbone(X) (18)
[0053] F c1 , F c2 , F c3 = Concat(F1, F2, F3) (19)
[0054] F1, F2, F3 = MASG(F c1 ), MSAG(F c2 ), MSAG(F c3 ) (20)
[0055] D3 = CLCC(F3, F as ) (21)
[0056] D2 = CLCC(F2, D3) (22)
[0057] D1 = CLCC(F1, D2) (23)
[0058] The output is expressed as:
[0059] R, R1, R2, R3 = SegHead(D1), Conv(D1), Conv(D2), Conv(D3) (24)
[0060] Furthermore, in step S3, the loss function combines the cross-entropy loss function and the dice loss function to supervise and train R, R1, R2, and R3 during training.
[0061] The beneficial effects of the present invention are as follows:
[0062] In response to the need for remote sensing image building extraction, by optimizing the multi-scale feature fusion and decoding process, the building semantic information capture ability of the model is improved, specifically including the following innovative modules and their optimization effects:
[0063] (1) Multi-scale Interaction Guidance Module (MSAG): This module aims to shorten the semantic gap between multi-level features while embedding multi-scale information into the features. In the remote sensing image scenario, buildings have diversity in scale and morphology. The MSAG module effectively fuses information of different scales through cross-scale feature interaction and position and spatial interaction, shortens the semantic information at different levels, enhances the model's recognition ability of building boundaries and shapes, and thus improves the extraction accuracy of buildings.
[0064] (2) Cross-layer Cross Decoding Module (CLCC): To optimize the context fusion of the decoder, the present invention designs the CLCC module, which performs context fusion through cross-layer cross-attention, enabling the model to effectively distinguish the semantic information between buildings and the background. The CLCC module gradually transmits and refines semantic features during the decoding process, enabling the model to focus on the semantic features of similar buildings while suppressing background information with large semantic differences from buildings. The layer-by-layer cross-decoding mechanism of this module ensures the maintenance of semantic consistency, making the building extraction result more refined, especially suitable for complex remote sensing image scenarios.
[0065] Other advantages, objectives, and features of the present invention will be elaborated to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. Description of the Drawings
[0066] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:
[0067] Figure 1 is the overall architecture diagram of the MGCC model;
[0068] Figure 2 is the architecture diagram of the MSAG;
[0069] Figure 3 is the detailed diagram of the MSAG;
[0070] Figure 4 It is the CLCC architecture diagram;
[0071] Figure 5 It is the MSSC architecture diagram;
[0072] Figure 6 It is the prediction result diagram of the Massachusetts dataset and the WHU dataset. Specific implementation manners
[0073] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only schematically illustrate the basic concept of the present invention. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0074] Please refer to Figures 1 to 6 , the present invention proposes a network structure (MGCC-Net) based on multi-scale interaction and cross-decoding. The overall architecture is as Figure 1 shown, including a multi-scale interaction guidance module (MSAG) and a cross-layer cross-decoding module (CLCC). This network adopts an encoder-decoder architecture. The encoder part uses BFBVit proposed in BuildFormer to capture global multi-scale information. At the same time, ASPP (Atrous Spatial Pyramid Pooling) is added to the last layer of feature extraction to form the Backbone of the present invention. The output features from top to bottom are F1, F2, F3, F as . The expression is as follows:
[0075] F1, F2, F3, F as = Backbone(X) (1)
[0076] (1) MSAG (multi-scale interaction guidance module)
[0077] After feature extraction by the encoder, the multi-scale interaction guidance module (MSAG) is used to shorten the semantic gap between multi-level features and embed multi-scale information into the features. This module mainly includes three parts: ① Multi-Scale Fusion; ② Position Interactive; ③ Channel Interactive. The overall architecture is as Figure 2 , and the detailed implementation is asFigure 3 。
[0078] ① Multi-scale fusion
[0079] For the extracted features Concatenate them to obtain Denote the concatenated features of the i-th layer. Subsequently Compress the channels to C through 1×1 convolution i (i = 1, 2, 3), and feed the compressed features into 3 parallel depthwise separable convolutions with kernel sizes of 3×3, 5×5, and 7×7 respectively. The results of the parallel outputs are added together to obtain the multi-scale fusion result This process is expressed as:
[0080] FMS i = B(Conv 3×3 (F ci )) + B(Conv 5×5 (F ci )) + B(Conv 7×7 (F ci )) (2)
[0081] Where B(·) represents BatchNorm. Conv k×k Represents the convolution operation with a kernel of k. Perform this operation on the first three layers of the encoder to obtain three multi-scale embedded features FMS1, FMS2, and FMS3.
[0082] ② Position Interactive
[0083] After multi-scale fusion, subsequent spatial position refinement is performed. First, apply global average pooling to the input feature FMS i along the horizontal and vertical directions respectively, and connect the output feature maps to generate This process is expressed as follows:
[0084] F hwi = [GAP h (FMS i ), GAP w (FMS i )] (3)
[0085] Where [,] represents the concatenation operation. Subsequently, apply 1×1 convolution for channel reduction to obtain Capture the information from the horizontal and vertical coordinates. r is the multiple of channel reduction. This process is expressed as:
[0086]
[0087] Among them, BR(·) represents BatchNorm and ReLU operations, and Conv represents 1×1 convolution. Then, for F rhwi Split along the horizontal and vertical directions to obtain F hi and F wi , and apply 1×1 convolution to F hi and F wi to restore the number of channels to C i Then, activate through the Sigmod function to obtain the directional weights W hi and W wi along the width and height dimensions. Finally, multiply this weight by FMS i to obtain the attention feature F pi with spatial position information interaction. This process can be expressed as:
[0088] F pi = FMS i *σ(Conv(F hi )) * σ(Conv(F wi )) (5)
[0089] where σ represents the Sigmoid activation function.
[0090] ③ Channel Interactive
[0091] In the first step of multi-scale fusion, 1×1 is used to compress the channels, resulting in a certain degree of loss of information from each layer. To make up for this shortcoming through channel interaction, first use global average pooling, and then perform 1×1 convolution and ReLU activation to generate channel attention features, adaptively filtering out redundant channel features. Subsequently, expand the channels to the output size through 1×1 convolution and multiply with the attention feature F pi with spatial position information interaction to obtain the complete multi-scale interaction output F mpci . This process can be expressed as:
[0092] F mpci = F pi * Conv(ReLU(Conv(GAP(FMS i )))) (6)
[0093] Therefore, define an MSAG, which is expressed as follows:
[0094] X = MSAG(X) (7)
[0095] (2) CLCC (Cross-Layer Cross Decoding Module)
[0096] The CLCC architecture, as shown in Figure 4As shown, features from the low layer and the high layer are fed into the multi-scale strip convolution block for multi-scale alignment and to generate relevant Q, K, V features. Subsequently, they are cross-fed into the multi-head self-attention to capture the similar semantic information of the building and perform global long-range dependence perception. At the same time, to enhance the ability to focus on local details, the C-MLP module is added.
[0097] The specific process is as follows:
[0098] Denote the feature from the low layer as F l , and the feature from the high layer as F h . First, upsample F h and perform channel alignment through a 1×1 convolution to match F l . This process is expressed as:
[0099] F h = Conv(Up(F hwi )) (8)
[0100] Subsequently, send F l and F h into the multi-scale strip convolution block (MSSC as Figure 5 ). For the input feature X, X ∈ R C ×H×W , define MSSC as follows:
[0101] Q, K, V = MSSC(X) = QKV(SC7(X) + SC 11 (X) + SC 21 (X)) (9)
[0102] Among them, QKV(X) represents performing three 1×1 convolutions on X to obtain three outputs, and SC k represents a group of serial 1×k and k×1 strip convolutions. The outputs obtained after sending F l and F h into MSSC are expressed as follows:
[0103] Q l , K l , V l = MSSC(BN(F l )) (10)
[0104] Q h , K h , V h = MSSC(BN(F h )) (11)
[0105] After obtaining the Q, K, and V matrices from the low-level features and high-level features, the Q matrix is cross-fed into the self-attention module to obtain SA l and SA h , and then a residual connection as shown in Figure 4 is performed to obtain the output F sa , where the self-attention module is expressed as:
[0106]
[0107] This process is expressed as:
[0108] F sa = SA(Q h , K l , V l ) + SA(Q l , K h , V h ) + F l + F h (13)
[0109] Finally, F sa is fed into C-MLP to obtain the final output D. This process is expressed as:
[0110] C_MLP(X) = Conv(DwConv 3×3 (ReLU(Conv(X)))) (14)
[0111] D = F sa + C MLP (BN(F sa )) (15)
[0112] where DwConv 3×3 represents a depthwise separable convolution with a 3×3 convolutional kernel.
[0113] For this reason, a CLCC is defined as follows:
[0114] X = CLCC(F l , F h ) (16)
[0115] (3) Overall output
[0116] After refining the results through layer-by-layer decoding, the output is obtained. At the same time, the network includes a multi-scale output for deep supervision.
[0117] The last layer of decoding, SegHead, is expressed as:
[0118] SegHead(X) = Conv(Up(Dropout(CBR(X)))) (17)
[0119] Among them, CBR represents Conv 3×3 , BatchNorm, ReLU, Up represents upsampling, Conv represents 1×1 convolution, and its output channels are the number of predicted categories. For the input image X, the overall network flow is expressed as:
[0120] F1,F2,F3,F as = Sackbone(X) (18)
[0121] F c1 ,F c2 ,F c3 = Concat(F1,F2,F3) (19)
[0122] F1,F2,F3 = MSAG(F c1 ),MSAG(F c2 ),MSAG(F c3 ) (20)
[0123] D3 = CLCC(F3,F as ) (21)
[0124] D2 = CLCC(F2,D3) (22)
[0125] D1 = CLCC(F1,D2) (23)
[0126] The output is expressed as:
[0127] R,R1,R2,R3 = SegHead(D1),Conv(D1),Conv(D2),Conv(D3) (24)
[0128] During training, R, R1, R2, and R3 are supervised and trained.
[0129] The network model training is implemented as follows:
[0130] (1) Loss function: The cross-entropy loss function and the dice loss function are combined, and the initial weight coefficients of different levels are all 1, and deep supervision is implemented for R, R1, R2, and R3.
[0131] (2) On two datasets, the AdamW optimizer and the cosine restart strategy are used to train the model for 240 rounds to ensure model convergence.
[0132] (3) On the Massachusetts dataset, we crop it to a size of 512×512, the cropping overlap rate is 0.1, and the parts containing blanks in the data are removed.
[0133] (4) During the training input, a data augmentation strategy is adopted, including random flipping and random rotation to increase the diversity of the data.
[0134] Verification experiment:
[0135] To evaluate the performance of the MGCC-Net network model, this experiment compares it with two existing cutting-edge methods, including: (1) BCTNet: a Transformer extraction method for dual-branch cross-fusion; (2) SDSNet: a building extraction method based on cross-layer feature information interaction filtering, which has achieved good results.
[0136] The F1, IOU, Precision, and Recall of the method of the present invention on the WHU dataset and the Massachusetts dataset reached 0.9556, 0.9151, 0.9564, 0.955 and 0.8609, 0.7558, 0.8778, 0.8446 respectively. The prediction effect is shown in Figure 6 .
[0137] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A method for extracting buildings from remote sensing images based on multi-scale interaction and cross-decoding, characterized in that The method specifically includes the following steps: S1: Obtain a remote sensing image, annotate labels, divide it into a training set, a validation set, and a test set. When the image is input, perform image data augmentation on it; S2: Construct the MGCC model, including the MSAG module and the CLCC module, where MGCC represents the building extraction network based on multi-scale interaction and cross-decoding, MSAG represents the multi-scale interaction guidance module, and CLCC represents the cross-layer cross-decoding module; the constructed MGCC model adopts an encoder-decoder architecture. The encoder part uses BFB Vit proposed in BuildFormer to capture global multi-scale information. At the same time, ASPP is added to the last layer of feature extraction to form the Backbone of the model, where ASPP represents the Atrous Spatial Pyramid Pooling. The output features are from top to bottom as ; The encoder process is described as: The decoder part includes an MSAG module and a CLCC module; The MSAG module is set after the encoder performs feature extraction, and is used to shorten the semantic gap between multi-level features and embed multi-scale information into the features; the MSAG module includes three parts: multi-scale fusion, position interaction, and channel interaction; The structure of the CLCC module is: send the features from the low layer and the high layer into the multi-scale strip convolution block for multi-scale alignment, and generate relevant Q, K, V features, and then cross-send them into the multi-head self-attention to capture the similar semantic information of the building and perform global long-range dependence perception. At the same time, in order to enhance the ability to focus on local details, a C-MLP module is added, where C-MLP represents a convolutional multi-layer perceptron; S3: Train the MGCC model: input the training set into the MGCC model, calculate the corresponding loss function, and then use the AdamW optimizer and the cosine restart strategy to iteratively train the MGCC model until the model converges; S4: Input the test set into the trained MGCC model to obtain an image with only building features.
2. The method for extracting buildings from remote sensing images according to claim 1, wherein In step S1, the image augmentation includes random flipping and random rotation.
3. The method for extracting buildings from remote sensing images according to claim 1, wherein In step S2, the multi-scale fusion is specifically as follows: for the extracted features , , i = 1, 2, 3, C represents the number of channels of the feature map, H represents the height of the feature map, W represents the width of the feature map, and they are concatenated to obtain , denotes the feature concatenated at the i th layer. Subsequently, is compressed to through a 1×1 convolution, and the compressed features are fed into 3 parallel depthwise separable convolutions. The results of the parallel outputs are added to obtain the multi-scale fusion result ; the above process is expressed as: Among them, represents BatchNorm; represents a convolution operation with a convolution kernel of k Performing this operation on the first three layers of the encoder yields three multi-scale embedded features FMS 1, FMS 2, FMS 3; The specific position interaction is as follows: After multi-scale fusion, subsequent spatial position refinement is performed. First, the input features are respectively processed along the horizontal and vertical directions by applying global average pooling and connecting the output feature maps to generate , and the above process is described as follows: Among them, represents global average pooling along the horizontal direction, represents global average pooling along the vertical direction; [,] represents the concatenation operation, and then a 1×1 convolution is applied for channel reduction to obtain ; capturing information from the horizontal and vertical coordinates, r is the multiple of channel reduction, and the above process is expressed as: Among them, represents the BatchNorm and ReLU operations, and Conv represents the 1×1 convolution; then is split along the horizontal and vertical directions to obtain and , and and are applied with 1×1 convolution to restore the channels to and then activated by the Sigmod function to obtain the directional weights along the width and height dimensions and . Finally, the weights are multiplied by to obtain the attention features with spatial position information interaction . The above process is described as: Among them, represents the Sigmoid activation function; The specific channel interaction is as follows: First, global average pooling is used, followed by 1×1 convolution and ReLU activation to generate channel attention features, adaptively filtering out redundant channel features. Subsequently, the channels are expanded to the output size through 1×1 convolution and multiplied by the attention features with spatial position information interaction to obtain the complete multi-scale interaction output. The above process is expressed as: Therefore, define the expression of MSAG as: 。 4. The method for extracting buildings from remote sensing images according to claim 3, wherein In step S2, the specific calculation process of the CLCC module is as follows: Denote the features from the lower layer as , , and denote the features from the upper layer as , ; First, perform upsampling on and achieve channel alignment through 1×1 convolution to match F l . The above process is expressed as: Subsequently, and are respectively fed into the multi-scale strip convolution block MSSC to process the input features , , and the MSSC is defined as follows: Among them, QKV ( X ) represents obtaining three outputs through three 1×1 convolutions on X , and represents a set of serial 1×k and k×1 strip convolutions; the outputs obtained after sending and into MSSC are expressed as follows: After obtaining the Q, K, and V matrices from the low-level features and high-level features, the Q matrix is cross-fed into the self-attention module to obtain and , and then a residual connection is performed to obtain the output , where the self-attention module is expressed as: The above process is expressed as: Finally, is fed into the C-MLP convolutional multi-layer perceptron to obtain the final output D , and the above process is expressed as: Among them, represents depthwise separable convolution with a 3×3 convolution kernel; Therefore, define the expression of CLCC as: 。 5. The method for extracting buildings from remote sensing images according to claim 4, wherein In step S2, after the MGCC model obtains the output after layer-by-layer decoding and refinement, the network includes a multi-scale output for deep supervision; the last layer of decoding SegHead is expressed as: Among them, CBR represents , Up represents upsampling, Conv represents a 1×1 convolution, and its output channels are the number of predicted classes. For the input image X , the overall network flow is expressed as: The output is expressed as: 。 6. The method for extracting buildings from remote sensing images according to claim 5, wherein In step S3, the loss function combines the cross-entropy loss function and the dice loss function, and performs supervised training on during training.
Citation Information
Patent Citations
Shielding target detection method and related equipment
CN115331194A
High-precision regression infrared small target tracking method
CN117218378A