High spatial resolution remote sensing image building extraction method
By combining multi-layer densely connected convolutional blocks and local and global attention units, the accuracy and complex scene adaptability problems of building extraction in high-spatial resolution remote sensing images are solved, and high-precision extraction and model simplification of buildings are achieved.
Patent Information
- Application Number
- CN202510341806.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-08-08
AI Technical Summary
In the existing technology, building extraction in high-spatial resolution remote sensing images, there are problems such as weak generalization ability, poor adaptability to complex scenes, blurred edges, and multi-scale adhesions. The semi-supervised method has insufficient edge refinement and high model complexity.
Multi-layer densely connected convolutional blocks (DenseNet) are used to combine local and global attention units to extract local features through local attention units, and global attention units are used to restore global information. Combined with pyramid pooling and pixel-by-pixel multiplication operations, the full utilization of features and precise extraction of buildings are achieved.
It improves the accuracy and performance of building extraction, and achieves a clear distinction between buildings and other land objects, especially the precise extraction of small-target buildings, with good generalization ability and low model complexity.
Smart Images

Figure CN120451766A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of building image processing, and in particular to a method for extracting buildings from remote sensing images with high spatial resolution. Background Art
[0002] With the rapid development of high spatial resolution remote sensing technology, building extraction has become a core demand in urban planning, disaster monitoring and other fields. However, traditional methods rely on manually designed spectral, texture and other features (such as edge density and shadow analysis), and have problems with weak generalization ability and poor adaptability to complex scenes. Although the convolutional neural network (CNN)-based method has significantly improved the accuracy through multi-scale feature fusion, it still faces challenges such as blurred building edges and multi-scale adhesion. Models such as DeeplabV3+ are prone to adhesion between adjacent buildings, while networks such as lightweight semantic segmentation (BiseNet) have the problem of incomplete contours. Although existing semi-supervised methods (such as pseudo-label iterative training) reduce the annotation cost, the edge refinement is insufficient and the model complexity is relatively high. Summary of the Invention
[0003] The present invention discloses a method for extracting buildings from remote sensing images with high spatial resolution. The specific method is as follows:
[0004] Get the image to be processed;
[0005] The image to be processed is input into a multi-layer densely connected convolution block. Starting from the second layer of multi-layer densely connected convolution blocks, the input of the current densely connected convolution block is composed of the outputs of all previous densely connected convolution blocks.
[0006] The outputs of the multi-layer densely connected convolutional blocks are processed by the corresponding local attention units to obtain several local attention outputs corresponding to the levels of the multi-layer densely connected convolutional blocks;
[0007] The highest-level local attention output is processed by the highest-level global attention unit and then concatenated with the next-level local attention output. The concatenated output is then processed by the next-level global attention unit until the first-level local attention output is completed. Finally, it is concatenated with the output of the first multi-layer densely connected convolutional block as the final building extraction image.
[0008] Furthermore, the size of the image to be processed is 256×256, the output image of the first layer of multi-layer densely connected convolution block is 256×256×112, the output image of the second layer of multi-layer densely connected convolution block is 128×128×192, the output image of the third layer of multi-layer densely connected convolution block is 64×64×304, the output image of the fourth layer of multi-layer densely connected convolution block is 32×32×464, and the output image of the fifth layer of multi-layer densely connected convolution block is 16×16×656.
[0009] Furthermore, the local attention unit includes a 1×1 convolutional layer, a 3×3 convolutional layer, a 5×5 convolutional layer, and a 7×7 convolutional layer;
[0010] The outputs of the multi-layer densely connected convolutional blocks are convolved through four convolutional layers respectively. After concatenating the outputs of the 3×3 convolutional layer, the 5×5 convolutional layer, and the 7×7 convolutional layer, they are dot-producted with the output of the 1×1 convolutional layer as the output of the local attention unit.
[0011] Furthermore, the global attention unit includes a first branch consisting of a global average pooling module, a 1×1 convolutional layer and a first deconvolutional layer, and a second branch consisting of a second deconvolutional layer. The outputs of the first branch and the second branch are concatenated as the output of the global attention unit.
[0012] Due to the adoption of the above technical solution, the present invention has the following beneficial effects:
[0013] 1. The multi-layer densely connected convolutional block adopted in this invention, namely DenseNet, performs well in image classification. It has the advantages of alleviating the gradient vanishing problem in deep neural networks, powerful feature extraction capabilities, promoting feature propagation in training and evaluation, and encouraging feature reuse in classification and segmentation tasks. At the same time, DenseNet can reduce the number of parameters, making it easy to train.
[0014] 2. The present invention adopts multi-layer densely connected convolution blocks to extract features and complete feature contraction; then, a feature class graph is generated based on the global attention unit as the contraction part. These two parts are connected by local attention units, making full use of the features of different stages and achieving good building extraction performance.
[0015] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings of the present invention are described below.
[0017] Figure 1 It is a schematic diagram of the overall architecture of the present invention.
[0018] Figure 2 Schematic diagram of the structure of the local attention unit. Figure 3 Design graph for global attention unit. Figure 4 A schematic diagram of the test results. DETAILED DESCRIPTION
[0019] The present invention will be further described below with reference to the accompanying drawings and examples.
[0020] A method for extracting buildings from remote sensing images with high spatial resolution, such as Figure 1 As shown, it specifically includes: obtaining an image to be processed;
[0021] The image to be processed is input into a multi-layer densely connected convolution block. Starting from the second layer of multi-layer densely connected convolution blocks, the input of the current densely connected convolution block is composed of the outputs of all previous densely connected convolution blocks.
[0022] The outputs of the multi-layer densely connected convolutional blocks are processed by the corresponding local attention units to obtain several local attention outputs corresponding to the levels of the multi-layer densely connected convolutional blocks;
[0023] The highest-level local attention output is processed by the highest-level global attention unit and then concatenated with the next-level local attention output. The concatenated output is then processed by the next-level global attention unit until the first-level local attention output is completed. Finally, it is concatenated with the output of the first multi-layer densely connected convolutional block as the final building extraction image.
[0024] Multi-layer densely connected convolutional block DenseNet:
[0025] Compared to other traditional deep convolutional neural network structures in this embodiment, for each feature extraction layer of the multi-layer densely connected convolution block, i.e. DenseNet, all previous layers are regarded as inputs to that layer. Therefore, there are L(L+1) / 2 connections in the L layer, instead of only L connections like in some traditional structures. In this way, DenseNet requires fewer parameters during training because redundant features do not need to be relearned. At the same time, each layer of the contraction part can access the gradient formed at the end of the entire model and the beginning of the structure, which improves the flow of information between layers and makes weights and biases easy to train.
[0026] Direct connection mode is used in dense blocks, where all layers are connected. This structure improves the flow of information between layers. To ensure that all layers in a dense block have the same size, a 3×3 convolution with padding is used in the blocks after the batch normalization and ReLU layers. This is defined as a nonlinear transformation Tl(). Any feature map xl can be calculated by Tl() through the previous layers, including x0, ...xl-1:
[0027] x l =T l ([x0,x1,…,x l-1 ])
[0028] Among them, [x0,x1,…,xl-1 ] is the concatenation operation of all output layers 0,...l-1.
[0029] As an important operation in deep neural networks, pooling changes the size of feature maps, helps extract information from different levels during training, and serves as a component between convolutional (Conv) layers. To facilitate downsampling in the DenseNet model and fully utilize the advantages of the architecture, dense blocks are connected through a pooling operation after 1×1 convolution.
[0030] The architecture of a deep neural convolutional network is symmetrical. The dilation phase is used to recover building information from the feature maps extracted by the contraction phase. Each dense block in the contraction phase corresponds to a local attention unit (LAU) and a global attention unit (GAU) in the dilation phase. The LAU is designed to extract local building information from remote sensing images during different neural network training stages. The GAUs are used to recover global building information from the deep feature maps. The extraction model is designed to be end-to-end, and the output size is the same as the input remote sensing image. Therefore, some upsampling operations are required in the dilation phase. During the dilation process, a LAU and a GAU are treated as a group, followed by a deconvolution layer to increase the size of the feature map. At the end of the model, a 1×1 convolution is used to map the feature map into two categories: building and non-building; the building extraction problem can be solved by binary classification. Training a deep neural convolutional network is a process of minimizing an energy function via gradient descent. Since the output of the softmax function can be used to represent the class distribution, the model's energy function is defined as the cross-entropy between the estimated class probabilities and the true distribution. In order to train the convolutional neural layer and the classifier coherently, the softmax layer and cross entropy are used in the network, where the softmax layer is defined to calculate the classification probability as follows:
[0031]
[0032] Among them, p i represents the probability that the pixel is predicted to belong to category i.
[0033] The image is divided into foreground and background. The number of categories K is set to 2 in the experiment. a is the output of the last layer of the model. The function can be defined as follows:
[0034]
[0035] Here, N represents the number of samples in the training dataset, y represents the expected output, and p is the probability mentioned above.
[0036] Local Attention Unit LAU:
[0037] The local attention unit (LAU) is designed to provide precise pixel-level attention to feature maps extracted from deep layers of a DenseNet. Thanks to its unique structure, pyramid pooling can extract information from feature maps of different scales; this approach also helps increase the receptive field and is widely used in semantic segmentation. However, the pyramid structure does not significantly address global contextual information, and the channel-level attention vectors used in it are limited to extracting pixel-level information.
[0038] In order to extract local pixel-level information of building networks from remote sensing images, LAU is designed to fuse feature maps of different scales and extract pixel-level information from the deep layer of DenseNet. In order to improve the performance of LAU in extracting information from feature maps of different scales, four different convolution operations are used with kernel sizes of 1×1, 3×3, 5×5 and 7×7. These features are gradually integrated by LAU from bottom to top, such as Figure 2 As shown in Figure 3, this allows for accurate incorporation of contextual information from neighboring scales. At the top of the LAU, a 1×1 convolution is designed to perform pixel-by-pixel multiplication with the feature information extracted by the bottom convolution operation. The pyramid structure is used to fuse information from different scales, while the pixel-by-pixel multiplication allows for better extraction of local pixel-level information required for building extraction.
[0039] Global Attention Unit GAU:
[0040] Global information is crucial for extracting buildings from remote sensing images, and some semantic segmentation models have designed methods for this purpose that directly generate results using bilinear upsampling. One-step decoders are limited to recovering positional information and, due to the lack of low-level feature maps at different scales, lose some global building features.
[0041] The contraction part is combined with the pyramid structure to improve the performance of semantic segmentation, and the global attention unit GAU is designed in the expansion part. The GAU introduces the global average pooling GAP into the unit to extract global building information, which is connected to the result of deconvolution and used as a guide to restore information, such as Figure 3 As shown in the figure. Specifically, 1×1 convolution and dense blocks are first applied, corresponding to the blocks in the contracting part, to operate on high-level feature maps. Then, the features are operated on in two ways: one using a deconvolution layer and the other using a GAP operation followed by 1×1 convolution and deconvolution. Finally, the features from these two methods are summed as the output of the GAU. This proposed unit effectively considers both low-level and high-level feature maps and provides global information to guide feature recovery in the dilated part of the network.
[0042] Experimental simulation:
[0043] The high-resolution remote sensing image building extraction method of the present invention achieves good results in overall accuracy, recall, F1-score, mean intersection over union (MIU), and Kappa coefficient (KappaCoefficient) accuracy evaluation, as shown in the following table.
[0044]
[0045] The overall accuracy of the method of the present invention is 94.65%, the recall rate, F1 score and MIoU are 0.9352, 0.9312 and 0.8891 respectively, and the Kappa coefficient is 0.92, indicating that the accuracy indicators of the method of the present invention are relatively ideal, and it can be applied to the actual extraction of buildings from high-resolution remote sensing images, and has very important practical application value.
[0046] In order to intuitively demonstrate the building extraction effect, the building results on the test set samples were obtained under the deep learning high-resolution remote sensing image building extraction method of the present invention, such as Figure 4 As shown, a representative building extraction image is compared and analyzed with an actual building image. The underlying surface of the image contains various features, including buildings, roads, vegetation, and water bodies. This demonstrates that the deep learning high-resolution remote sensing image building extraction method of the present invention can clearly distinguish buildings from other features and accurately extract building features. This application can also accurately extract small buildings.
[0047] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.
Claims
1. A method for extracting buildings from remote sensing images with high spatial resolution, characterized in that: The specific method is as follows: Get the image to be processed; The image to be processed is input into a multi-layer densely connected convolution block. Starting from the second layer of multi-layer densely connected convolution blocks, the input of the current densely connected convolution block is composed of the outputs of all previous densely connected convolution blocks. The outputs of the multi-layer densely connected convolutional blocks are processed by the corresponding local attention units to obtain several local attention outputs corresponding to the levels of the multi-layer densely connected convolutional blocks; The highest-level local attention output is processed by the highest-level global attention unit and then concatenated with the next-level local attention output. The concatenated output is then processed by the next-level global attention unit until the first-level local attention output is completed. Finally, it is concatenated with the output of the first multi-layer densely connected convolutional block as the final building extraction image.
2. The high spatial resolution remote sensing image building extraction method according to claim 1, wherein: The size of the image to be processed is 256×256, the output image of the first layer of multi-layer densely connected convolution block is 256×256×112, the output image of the second layer of multi-layer densely connected convolution block is 128×128×192, the output image of the third layer of multi-layer densely connected convolution block is 64×64×304, the output image of the fourth layer of multi-layer densely connected convolution block is 32×32×464, and the output image of the fifth layer of multi-layer densely connected convolution block is 16×16×656.
3. The high spatial resolution remote sensing image building extraction method according to claim 1, wherein: The local attention unit includes 1×1 convolution layer, 3×3 convolution layer, 5×5 convolution layer and 7×7 convolution layer; The outputs of the multi-layer densely connected convolutional blocks are convolved through four convolutional layers respectively. After concatenating the outputs of the 3×3 convolutional layer, the 5×5 convolutional layer, and the 7×7 convolutional layer, they are dot-producted with the output of the 1×1 convolutional layer as the output of the local attention unit.
4. The method for extracting buildings from remote sensing images with high spatial resolution according to claim 1, wherein: The global attention unit includes a first branch consisting of a global average pooling module, a 1×1 convolutional layer and a first deconvolutional layer, and a second branch consisting of a second deconvolutional layer. The outputs of the first branch and the second branch are concatenated as the output of the global attention unit.
Citation Information
Cited By
Pavement crack detection equipment based on double spectrums
CN121095246A