A method for extracting buildings from remote sensing images based on deep learning

By improving the Deeplabv3+ model, using MobilenetV2 and DenseASPP, and adding the ECA attention module, the problem of loss of boundary information during building extraction in high-resolution remote sensing images is solved, and the segmentation accuracy and calculation efficiency are achieved.

CN116912708BActive Publication Date: 2025-07-01CHONGQING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310894293.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-20
Publication Date
2025-07-01
Estimated Expiration
2043-07-20

AI Technical Summary

Technical Problem

In the prior art, when extracting buildings in high-resolution remote sensing images, there are problems such as loss of boundary information, poor segmentation results, and high model complexity.

Method used

Improve the Deeplabv3+ model, use the lightweight MobilenetV2 as the backbone network, and replace the ASPP module with DenseASPP, and add the ECA attention module to improve the model's ability to extract small-scale information.

Benefits of technology

It achieves higher segmentation accuracy, more detailed edge information extraction, reduced model parameters, and improved calculation efficiency, and is suitable for semantic segmentation of complex remote sensing images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116912708B_ABST
    Figure CN116912708B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for extracting buildings from remote sensing images based on deep learning, belonging to the field of remote sensing images. The method includes: First, based on the processed public dataset, the Deeplabv3+ network model is improved by replacing its backbone network with the lightweight MobilenetV2; the ASPP module of the model is replaced with DenseASPP, and a set of dilated convolutions are connected in a denser way to generate multi-scale features covering a larger scale range. Aiming at the problem that the model has insufficient representation ability for small-scale feature information, a channel attention mechanism is added after the DenseASPP module to improve the model performance by strengthening the channel features more interested in small buildings; considering that the shallow features contain more original information, two layers of shallow features are first fused in the decoding area to provide more refined spatial information and enhance robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of remote sensing images, and relates to a method for extracting buildings from remote sensing images based on deep learning. Background Art

[0002] At the present stage, the methods for extracting remote sensing image information are mainly carried out pixel by pixel. Due to the rich details and large amount of data of high-resolution images, traditional pixel-based methods are not suitable for processing high-resolution image data. At present, there are mainly two methods for extracting building features from high-resolution remote sensing images: 1) Traditional machine learning classifiers (such as support vector machine (SVM), random forest) are used to extract building features, and corresponding post-processing steps are usually adopted to refine the segmentation results; 2) Based on traditional computer vision methods, artificial features such as vegetation indices, textures and color features are used. However, these methods not only have high model complexity, but are also limited by artificial knowledge and experience. For example, convolutional neural network CNN (Convolution Neural Network) or fully convolutional network FCN (Fully Convolutional Networks) based on the encoder-decoder architecture has been successfully applied in this field and is superior to traditional computer vision methods. However, due to the use of pooling layers for downsampling, although increasing the receptive field of convolutional kernels can extract global features of images, high-frequency details in the images are lost at the same time, resulting in the lack of boundary information in the segmentation results, and it is easy to segment building features into many "round spots", making it difficult to extract the complete boundaries of building features. Therefore, in the face of a large amount of high-resolution remote sensing image data, how to quickly and effectively extract the urban building information we need from the large amount of data is the difficulty and focus of high-resolution remote sensing image interpretation; secondly, due to the diversity of building features in different regions, such as factors like color, shape, size, etc., and the similarity between building features and the background or other objects, the research on methods for extracting building features from remote sensing images is extremely challenging.

[0003] The recognition and detection of buildings will directly affect the automation level of feature mapping. Therefore, timely mastering the information of urban buildings is of great significance for better urban development planning and building a digital city.

[0004] As early as in the 1970s of the last century, people first used computers to analyze urban buildings in remote sensing images. The main methods can be divided into texture-based methods, edge-based methods and shadow-based methods.

[0005] Stephen Levitt proposed a texture-based building detection method. Since there are differences in textures between artificial buildings and natural areas, by measuring textures, different textures can be used to distinguish building areas from other areas. Ye Sheng utilized amplitude spectrum information to extract multi-layer building textures and edge features, and combined the two to achieve the extraction of multi-layer building information, which has practical application value. However, this method cannot avoid the interference of the "same frequency, different object" phenomenon in high-resolution images on the extraction accuracy. Zhang Hao et al. proposed a method for extracting LiDAR point cloud buildings based on gray-level co-occurrence matrix texture to automatically extract buildings. To reduce the computational amount, the authors compressed the gray levels during the calculation of the gray-level co-occurrence matrix, which would cause some textures to be lost and lead to misclassification. Generally speaking, texture-based extraction methods are very effective for processing medium- and low-resolution images. However, high-resolution images have complex textures, and it is difficult to extract buildings using this method. C. Lin and Nevitia et al. proposed the perceptual grouping theory. They first used edge detection to extract the outlines of buildings; then grouped the extracted outlines according to spatial relationships, and finally searched for parallel lines to obtain rectangular outlines, and then extracted the external positions of buildings. Chen et al. proposed the edge regularity index and the shadow line index as new features for building candidates obtained from the segmentation method to refine the boundaries of the detection results.

[0006] In recent years, deep learning has developed rapidly in the field of artificial intelligence, and the research methods in various fields have gradually shifted from traditional research methods to deep neural networks. Since the convolutional neural network was proposed, deep learning represented by the convolutional neural network has achieved amazing results in the field of image learning. As an emerging research hotspot, deep learning has achieved good results in fields such as speech recognition, computer vision, and natural language processing. Compared with traditional neural networks, deep neural networks have made significant improvements. Through the method of "layer-by-layer learning", it effectively reduces the difficulty of data training problems. By learning deep non-linear network structures, it can approximate numerous complex functions. Its method can layer-by-layer mine data features from a large number of labeled data and learn the essence of their features, and demonstrate considerable learning ability in unlabeled datasets. Therefore, deep learning has been widely applied in fields such as image classification, object detection, and image segmentation. In the field of image segmentation, semantic segmentation of high-altitude remote sensing images is helpful for fields such as urban road planning, geological exploration, and national defense construction. Bischke et al. introduced a new cascaded multi-task loss in the deep network structure, fused the boundary information of buildings, overcame the generation of "speckle" segmentation phenomena, and solved the problem of retaining semantic segmentation boundaries in high-resolution satellite images. The mean intersection over union index of this method reached 73%. Zhou et al. proposed a D-LinkNet semantic segmentation network for road extraction in high-resolution satellite images. This network adopts a link network structure and expands the convolutional layer in the central part, and learns and fuses multi-scale information of high-level semantic features through the central part. The mloU index in the CVPR 2018 Deep Globe Road Extraction Challenge reached 64.66%. Liu Hao et al. improved the Unet network and used a loss function composed of the Dice coefficient and the cross-entropy function for training, achieving good results in extracting irregular buildings.

[0007] High-resolution remote sensing images can provide detailed building information. Deep learning can better learn the data features of high spatial resolution and can efficiently extract building features. The Deeplab series of networks incorporates dilated convolutions to expand the receptive field and proposes the Atrous Spatial Pyramid Pooling (ASPP) layer, using different dilation rates to extract multi-scale features of images and improve the accuracy of remote sensing image segmentation. However, the Deeplab series still has problems such as slow fitting speed, relatively rough extraction of edge information, blurred segmentation of small-scale targets, and holes in the segmentation of large-scale targets when extracting buildings from remote sensing images. Therefore, how to select a suitable and efficient deep learning network to extract buildings has always been a hot topic of concern among scholars. Summary of the Invention

[0008] In view of this, the purpose of the present invention is to provide a method for extracting buildings from remote sensing images based on deep learning.

[0009] To achieve the above object, the present invention provides the following technical solutions:

[0010] A method for extracting buildings from remote sensing images based on deep learning, the method comprising the following steps:

[0011] S1: Select the publicly available remote sensing building dataset lnria Aerial lmage Labeling Dataset as the original data, and perform data preprocessing steps including cropping and data augmentation;

[0012] S2: Improve the traditional Deeplabv3+ model, replace its backbone network with the lightweight MobilenetV2 to effectively reduce the number of parameters and improve the model speed; replace the main module ASPP of the model with DenseASPP to connect a group of dilated convolutions in a more dense manner; add an attention mechanism ECA module, which directly generates a channel attention map through two 1×1 convolutional layers, recalibrates the channel weights of the feature map, selects more important feature channels, and improves the model's ability to extract small-scale information, to obtain the DAEC-Deeplabv3+ network;

[0013] S3: Input the processed remote sensing image dataset into the DAEC-DeepLabv3+ network for training to obtain a trained building detection model;

[0014] S4: Detect the trained building detection model in the test set of remote sensing images to obtain the remote sensing building image segmentation result.

[0015] Optionally, the S1 specifically includes:

[0016] S11: Download the required remote sensing building dataset from the dataset website. The lnria Aerial lmage LabelingDataset dataset includes 180 color images of 5000×5000 pixels and their corresponding 180 binary grayscale images of 5000×5000 pixels. The test images include 180 color images of pixels; this dataset covers urban residential areas in Europe and the United States, and the scope covers five major cities of Austin, Chicago, Kitsap, Tyrol, and Vienna, with 36 images corresponding to each city; all the annotated images classify these buildings into two categories: buildings and non-buildings. The pixel value of the building area is 255, and the pixel value of the non-building area is 0; crop the images in the dataset into 500 pixels * 500 pixels, and perform dataset augmentation using rotation transformation, mirror transformation, or brightness transformation;

[0017] S12: Randomly divide the data in the dataset into training data, validation data, and test data according to the ratio of 7:2:1. The divided files are stored in sub-files, namely train.txt, val.txt, and test.txt respectively.

[0018] Optionally, the specific content of S2 includes:

[0019] S21: Construct an improved DAEC_Deeplabv3+ network model, and use the lightweight network MobileNetV2 as the backbone network;

[0020] S22: In the Encoder of the coding area, replace the original ASPP module with DenseASPP, and apply the dense connection idea in DenseNet to ASPP; it cascades multiple dilated convolutional layers and conveys the output of each dilated convolutional layer to all subsequent unvisited dilated convolutional layers in a dense connection manner;

[0021] DenseASPP is represented by formula (1):

[0022] y l = H K,dl ([y l-1 ,y l-2 ,...,y0]) (1)

[0023] d l represents the dilation rate of layer l, [...] represents the concatenation operation, and [y l-1 ,...,y0] represents the features generated after connecting the outputs of all previous layers;

[0024] S23: After feature extraction is completed by DensaASPP, the features will be stacked and thickened, and the ECA module is added. This module removes the fully connected layer by directly using a 1×1 convolutional layer after the global average pooling layer, recalibrates the channel weights of the feature map, selects more important feature channels, and improves the model's ability to extract small-scale information; it efficiently realizes local cross-channel interaction with 1D convolution and extracts the dependency relationship between channels; the specific steps are: first perform a global average pooling operation on the input feature map, then perform a 1D convolution operation with a convolutional kernel size of k, and obtain the weights w of each channel through the Sigmoid activation function, as shown in formula (2):

[0025] ω = σ(C1D K (y)) (2)

[0026] Multiply the weights with the corresponding elements of the original input feature map to obtain the final output feature map, and then use a 1×1 convolution operation on this deep feature with higher semantic information to adjust the number of channels; where C1D represents one-dimensional convolution;

[0027] S24: In the decoder of the decoding area, extract the shallow features of the 4th and 7th layers from the backbone network, perform a multi-scale feature fusion MSFF operation, add them to the deep feature layer after 2-fold upsampling, and then perform another 2-fold upsampling to restore the size of the feature layer to achieve semantic segmentation; then connect with the original shallow features obtained from the backbone network to increase the number of channels, and then use a 3×3 convolution for feature extraction, and finally adjust the output image to the same size as the input image; the MSFF operation among them is to perform multi-scale feature fusion on the two feature layers.

[0028] Optionally, the specific steps of S3 are as follows:

[0029] Use the PyCharm programming software to train in the PyTorch deep learning framework; input the processed remote sensing image dataset into the DAEC_DeepLabv3+ network for training, obtain the trained building detection model, and save the pre-trained model parameters for different segmentation objects.

[0030] Optionally, the specific steps of S4 are as follows:

[0031] S41: Input the test images in the test set into the trained improved deeplabv3+ network model, select the segmentation object as a building or background, and after outputting the corresponding segmentation result, save the result image of the model;

[0032] S42: Select cross-entropy as the loss of the algorithm, and the mean pixel accuracy mPA and the mean intersection over union mIoU as evaluation indicators. Evaluate the training results of the model from two perspectives: the proportion of the correctly predicted pixels in the union of the predicted pixels and the true pixels, and the proportion of the correct pixels in the total pixels. The higher the MIoU value, the better the image segmentation effect; the formula for cross-entropy is:

[0033]

[0034] where, y i is the true value of a certain pixel, and the true value is 0 or 1 in the binary classification task; is the predicted value of a certain pixel; n is the number of samples for each calculation of the loss;

[0035] The calculation formulas for the mean pixel accuracy mpA and the mean intersection over union mIoU are respectively:

[0036]

[0037]

[0038] Among them, TP represents that the model prediction is correct, that is, both the model prediction and the actual are positive examples; FP represents that the model prediction is incorrect, that is, the model predicts the category as a positive example, but the actual category is a negative example; FN represents that the model prediction is incorrect, that is, the model predicts the category as a negative example, but the actual category is a positive example; TN represents that the model prediction is correct, indicating that both the model prediction and the actual are negative examples.

[0039] The beneficial effects of the present invention are as follows:

[0040] First, the backbone network is changed to the lightweight Mobilenetv2, which greatly reduces the number of model parameters and the amount of computation.

[0041] Second, the original ASPP module of the model is changed to DenseASPP. The convolutional layers with four different dilation rates (3, 6, 12, 18) are connected in the connection form of DenseNet to form a dense feature pyramid. Each layer in it will fuse multiple different scale features of its previous layers in parallel, so as to generate fused features with multi-scale large receptive fields. Such a large receptive field provides more context information for the segmentation of large targets in high-resolution images.

[0042] Third, the ECA attention module added in the present invention directly generates a channel attention map through two 1×1 convolutional layers, avoiding the traditional attention matrix multiplication, effectively improving the computational efficiency, enhancing the model's ability to extract small-scale information, and the MIoU index of the segmented remote sensing image is improved, and the segmentation accuracy is higher;

[0043] Fourth, considering that the shallow features contain more original information, two shallow feature layers are first fused in the decoding area to retain richer spatial information and enhance the robustness. In addition, the method of the present invention is easy to understand and operate as a whole, and can be applied to the semantic segmentation of other complex remote sensing images.

[0044] Other advantages, objectives and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:

[0046] Figure 1This invention provides a preferred embodiment, which is a framework diagram of a remote sensing image building extraction model based on an improved DAEC_Deeplabv3+;

[0047] Figure 2 It is the DenseASPP module in the encoding network;

[0048] Figure 3 It is the attention mechanism ECA added to the encoding network;

[0049] Figure 4 It is the multi-scale feature fusion module MSFF. Specific implementation mode

[0050] The following uses specific specific examples to illustrate the implementation mode of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation modes. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0051] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and cannot be understood as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged or reduced, and do not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0052] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, it is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and cannot be understood as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0053] Please refer to Figures 1 to 4, S1 is the source of the input data of this model and the data preprocessing method. The lnriaAerial lmage Labeling Dataset includes 180 color images of 5000×5000 pixels and their corresponding 180 binary grayscale images of 5000×5000 pixels. The test images include 180 color images of pixels. All the labeled images classify these buildings into two categories: buildings and non-buildings. The pixel value of the building area is 255 (white), and the pixel value of the non-building area is 0 (black). The images in the dataset are cropped into 500 pixels * 500 pixels, and geometric and non-geometric methods such as rotation transformation, mirror transformation, and brightness transformation are used for dataset augmentation. After image augmentation, the entire dataset is divided into a training set, a test set, and a validation set according to the ratio of 7:2:1.

[0054] S2 is to construct an improved DAEC_Deeplabv3+ network model. Its backbone network is replaced with the lightweight MobilenetV2, which effectively reduces the number of parameters and improves the model speed; in the Encoder of the encoding area, DenseASPP is used to replace the original ASPP module, and the dense connection idea in DenseNet is applied to ASPP. ASPP can be expressed by formula (1):

[0055] y = H 3,6(x) +H 3,12(x) +H 3,18(x) +H 3,24(x) (1)

[0056] H K,d(x) is used to represent a dilated convolution, and y represents the fused feature.

[0057] DenseASPP cascades multiple dilated convolution layers and conveys the output of each dilated convolution layer to all subsequent unvisited dilated convolution layers in a dense connection manner. In DenseASPP, the dilated convolution layer makes full use of a reasonable dilation rate. Through a series of feature connections, all the neurons of the intermediate features can encode semantic information of different scales, and different intermediate features have different scale ranges. Therefore, the final DenseASPP feature output has a larger receptive field of more scales, and a denser and larger feature pyramid can be generated by using several dilated convolution layers.

[0058] DenseASPP can be expressed by formula (2):

[0059]

[0060] d l represents the dilation rate of the l-th layer, and [...] represents the concatenation operation, [yl-1 ,..., y0] represents the features generated after concatenating the outputs of all previous layers. Compared with ASPP, DenseASPP stacks all dilated convolutional layers in a dense connection manner, resulting in a denser feature pyramid and a larger receptive field. After feature extraction by DensaASPP, the features are stacked and thickened, and an ECA (External Channel Attention) module is added. This module can avoid dimensionality reduction and effectively capture cross-channel interaction information. It efficiently realizes local cross-channel interaction using 1D convolution and extracts the dependency relationships between channels. The specific steps are as follows: First, perform global average pooling on the input feature map, then perform 1D convolution with a kernel size of k, and obtain the weights w of each channel through the Sigmoid activation function, as shown in Equation (3):

[0061] ω = σ(C1D K (y)) (3)

[0062] Multiply the weights by the corresponding elements of the original input feature map to obtain the final output feature map, and then use a 1×1 convolution operation to adjust the number of channels for this deep feature with high semantic information. In the decoder Decoder, extract the shallow features of the 4th and 7th layers from the backbone network, perform MSFF operation and add them to the deep feature layer that has been upsampled by 2 times, then perform another upsampling by 2 times to restore the size of the feature layer to achieve semantic segmentation. Then concatenate and stack with the original shallow features obtained from the backbone network to increase the number of channels, and then use a 3×3 convolution for feature extraction. Finally, adjust the output image to the same size as the input image. The MSFF operation among them is to perform multi-scale feature fusion on the two feature layers.

[0063] Furthermore, S3 is trained using the PyCharm programming software in the PyTorch deep learning framework. As an implementable approach, select the professional version of PyCharm and use the deep learning framework PyTorch 1.11.0 to train the network structure. Input the processed remote sensing image dataset into the DAEC-DeepLabv3+ network for training, obtain a trained building detection model, and save the pre-trained model parameters for different segmentation objects.

[0064] Further, in S4, the test images in the test set are input into the trained improved DAEC_Deeplabv3+ network model. The selected segmentation objects are buildings or backgrounds. After the corresponding segmentation results are output, the result images of the model are saved. During the network training process, cross-entropy is selected as the loss of the algorithm, and the mean pixel accuracy (mPA) and the mean intersection over union (mIoU) are used as evaluation metrics. The training results of the model are evaluated from two aspects: the proportion of correctly predicted pixels in the union of the predicted pixels and the true pixels, and the proportion of correct pixels in the total pixels. The higher the mIoU value, the better the image segmentation effect. The formula for cross-entropy is:

[0065]

[0066] where y i is the true value of a certain pixel, and in a binary classification task, the true value is 0 or 1; is the predicted value of a certain pixel; n is the number of samples for each loss calculation;

[0067] The calculation formulas for the mean pixel accuracy (mpA) and the mean intersection over union (mIoU) are respectively:

[0068]

[0069]

[0070] In the above formula, TP represents that the model prediction is correct, that is, both the model prediction and the actual are positive examples; FP represents that the model prediction is incorrect, that is, the model predicts this category as a positive example, but actually this category is a negative example; FN represents that the model prediction is incorrect, that is, the model predicts this category as a negative example, but actually this category is a positive example; TN represents that the model prediction is correct, meaning that both the model prediction and the actual are negative examples;

[0071] The experimental results show that the present invention is based on the improved DAEC_Deeplabv3+ network model, with smaller training parameter quantities, higher segmentation accuracy, more detailed edge information extraction, and effectively improving problems such as holes in large-scale target segmentation. In addition, the method of the present invention has good robustness, is easy to understand as a whole, and is easy to operate, and can be applied to the semantic segmentation of other complex remote sensing images.

[0072] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A method for extracting buildings from remote sensing images based on deep learning, characterized in that: The method includes the following steps: S1: Select the publicly available remote sensing building dataset lnria Aerial lmage Labeling Dataset as the original data, and perform data preprocessing steps including cropping and data augmentation; S2: Improve the traditional Deeplabv3+ model by replacing its backbone network with the lightweight MobilenetV2, effectively reducing the number of parameters and improving the model speed; replace the main module ASPP of the model with DenseASPP to connect a group of dilated convolutions in a more dense manner; add the attention mechanism ECA module, which directly generates a channel attention map through two 1×1 convolutional layers, recalibrates the channel weights of the feature map, selects more important feature channels, and improves the model's ability to extract small-scale information, obtaining the DAEC-Deeplabv3+ network; The specific content of S2 includes: S21: Construct an improved DAEC_Deeplabv3+ network model, and use the lightweight network MobileNetV2 as the backbone network; S22: In the Encoder of the encoding area, replace the original ASPP module with DenseASPP (Dense Atrous Spatial Pyramid Pooling), and apply the dense connection idea in DenseNet to ASPP; it cascades multiple dilated convolutional layers and conveys the output of each dilated convolutional layer to all subsequent unvisited dilated convolutional layers in a dense connection manner; DenseASPP is represented by formula (1): d l represents the expansion rate of layer l, [...] represents the concatenation operation, [y l-1 ,..., y0] represents the feature generated after concatenating the outputs of all previous layers; S23: After feature extraction is completed by DensaASPP, the features will be stacked and thickened, and the ECA module is added. This module directly uses a 1×1 convolutional layer after the global average pooling layer, removes the fully connected layer, recalibrates the channel weights of the feature map, selects more important feature channels, and improves the model's ability to extract small-scale information; it efficiently realizes local cross-channel interaction with 1D convolution and extracts the dependency relationship between channels; the specific steps are as follows: first, perform global average pooling operation on the input feature map, then perform 1D convolution operation with a convolution kernel size of k, and obtain the weights w of each channel through the Sigmoid activation function, as shown in formula (2): ω=σ(C1D K (y)) (2) Multiply the weights by the corresponding elements of the original input feature map to obtain the final output feature map, and then use a 1×1 convolution operation on this deep feature with higher semantic information to adjust the number of channels; where C1D represents 1D convolution; S24: In the Decoder of the decoding area, extract the shallow features of the 4th and 7th layers from the backbone network, perform multi-scale feature fusion (MSFF) operation and add them to the deep feature layer that has been upsampled by 2 times, then perform another upsampling by 2 times to restore the size of the feature layer to achieve semantic segmentation; then connect with the original shallow features obtained from the backbone network to increase the number of channels, and then use a 3×3 convolution for feature extraction, and finally adjust the output image to the same size as the input image; the MSFF operation among them is to perform multi-scale feature fusion on two feature layers; S3: Input the processed remote sensing image dataset into the DAEC-DeepLabv3+ network for training to obtain a trained building detection model; S4: Detect the trained building detection model in the test set of remote sensing images to obtain the remote sensing building image segmentation result.

2. The method for extracting buildings from remote sensing images based on deep learning according to claim 1, characterized in that: The specific content of S1 includes: S11: Download the required remote sensing building dataset from the dataset website. The lnria Aerial lmage Labeling Dataset dataset includes 180 color images of 5000×5000 pixels and their corresponding 180 binary grayscale images of 5000×5000 pixels. The test images include 180 color images of pixels; this dataset covers urban residential areas in Europe and the United States, covering five major cities of Austin, Chicago, Kitsap, Tyrol, and Vienna, with 36 images corresponding to each city; all the annotated images classify these buildings into two categories: buildings and non-buildings. The pixel value of the building area is 255, and the pixel value of the non-building area is 0; crop the images in the dataset into 500 pixels * 500 pixels, and use rotation transformation, mirror transformation, or brightness transformation methods to enhance the dataset; S12: Randomly divide the data in the dataset into training data, validation data, and test data according to the ratio of 7:2:

1. The divided files are stored in sub-files, namely train.txt, val.txt, and test.txt.

3. A method for extracting buildings from remote sensing images based on deep learning according to claim 1, characterized in that: The specific content of S3 includes: Use the PyCharm programming software to conduct training in the PyTorch deep learning framework; input the processed remote sensing image dataset into the DAEC_DeepLabv3+ network for training to obtain a trained building detection model and save the pre-trained model parameters for different segmentation objects.

4. The method for extracting buildings from remote sensing images based on deep learning according to claim 3, characterized in that: The specific content of S4 includes the following steps: S41: Input the test images in the test set into the trained improved deeplabv3+ network model, select the segmentation object as a building or background, and after outputting the corresponding segmentation result, save the result image of the model; S42: Select cross-entropy as the loss of the algorithm, and the mean pixel accuracy mPA and the mean intersection over union mIoU as evaluation indicators. Evaluate the training result of the model from two perspectives: the proportion of the correctly predicted pixels in the union of the predicted pixels and the true pixels, and the proportion of the correct pixels in the total pixels. The higher the MIoU value, the better the image segmentation effect; the formula for cross-entropy is: where y i is the true value of a certain pixel, and in a binary classification task, the true value is 0 or 1; is the predicted value of a certain pixel; n is the number of samples for each loss calculation; The calculation formulas for the mean pixel accuracy mpA and the mean intersection over union mIoU are respectively: Among them, TP represents that the model prediction is correct, that is, both the model prediction and the actual are positive examples; FP represents that the model prediction is incorrect, that is, the model prediction category is a positive example, but the actual category is a negative example; FN represents that the model prediction is incorrect, that is, the model prediction category is a negative example, but the actual category is a positive example; TN represents that the model prediction is correct, indicating that both the model prediction and the actual are negative examples.

Citation Information

Patent Citations

  • Remote sensing image semantic segmentation method based on channel-space attention and DeeplabV3plus

    CN114596500A

  • Remote sensing image building extraction method based on improved DeepLabV3 +

    CN114663759A