An algorithm to improve the accuracy of building edge segmentation using boundary information

By using the algorithm of boundary information under the multi-task learning framework, processing data sets and defining joint cost functions to train UNet networks, the problem of insufficient edge segmentation accuracy in the existing technology is solved, and higher edge segmentation accuracy and robustness are achieved.

CN114240977BActive Publication Date: 2025-05-16TIANDI INFORMATION NETWORK RES INST (ANHUI) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111575708.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-21
Publication Date
2025-05-16
Estimated Expiration
2041-12-21

AI Technical Summary

Technical Problem

The existing building extraction algorithms with high resolution remote sensing images have a contradiction between accuracy and robustness in edge segmentation. The traditional algorithms have low edge accuracy but low robustness, while the deep learning-based algorithms have blurred edges but high overall matching coefficient.

Method used

Under the multi-task learning framework, an algorithm that uses boundary information to improve the accuracy of building edge segmentation is used to generate benchmark data for building main body, boundary, separation area and method direction by processing the data set, and a joint cost function training UNet network is defined, building edge information is extracted, and multiple elements are fused through the watershed algorithm to improve edge segmentation accuracy.

Benefits of technology

It effectively improves the building edge segmentation accuracy based on high-resolution remote sensing images, combines the edge accuracy of traditional algorithms and the robustness of deep learning algorithms, and improves the overall matching coefficient.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114240977B_ABST
    Figure CN114240977B_ABST
Patent Text Reader

Abstract

The present invention discloses an algorithm for improving the accuracy of building edge segmentation by using boundary information. First, a data set is processed to generate four kinds of auxiliary training benchmark data, namely, building main body, boundary, separation area between buildings and edge normal direction, based on existing building marks; secondly, a joint cost function based on building extraction and edge extraction is defined, and a UNet-like network is trained using the processed data set; then, for any image to be segmented, the above four elements are extracted using the trained UNet-like network; finally, the four elements are fused to obtain an edge-enhanced building extraction result. The present invention converts building extraction into a multi-task joint learning problem, starting from constructing a joint cost function, strengthening the constraint on the edge during the building extraction process, and fusion of the inferred edge and building information during the fusion process to further improve the segmentation effect of the building edge.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of building extraction algorithms, and in particular to an algorithm for improving building edge segmentation accuracy by using boundary information. Background Art

[0002] Existing building extraction algorithms are mainly divided into two categories: traditional algorithms and deep learning-based algorithms.

[0003] The traditional algorithm directly detects the edges of building images. Wleite uses random forests to create two binary classifiers. One classifier is used to distinguish whether each pixel belongs to an edge, and the other classifier is used to distinguish whether each pixel is inside or outside the building. Then, based on edge detection and pixel classification, brute force matching is used to perform polygon screening, and each polygon candidate is screened using the IOU value.

[0004] Most deep learning-based algorithms do not use boundary information at all and obtain results directly from image learning; Audebert et al. proposed to use OpenStreetMap information joint learning to improve the segmentation accuracy of buildings, by using some rules to create more labeling constraints, and at the same time using the residential land, agricultural land, industrial land, water area, buildings and roads and other layer information in OpenStreetMap; Marmanis first extracts the edge information and puts it into the neural network as an input together with the image information for training.

[0005] In the building extraction algorithm based on high-resolution remote sensing images, the building edges segmented by the traditional algorithm are more accurate, but the robustness is not high, and the overall matching coefficient IOU is low; the overall matching coefficient of the building extraction algorithm based on deep learning is higher, but because the ordinary convolutional neural network is an inductive deduction of large areas, the extracted building edges are blurred. Summary of the invention

[0006] In response to the technical problems existing in the existing building extraction algorithms for high-resolution remote sensing images, the present invention provides an algorithm that uses boundary information to improve the accuracy of building edge segmentation under a multi-task learning framework, thereby effectively improving the accuracy of building edge segmentation based on high-resolution remote sensing images.

[0007] An algorithm for improving the accuracy of building edge segmentation using boundary information includes the following steps:

[0008] Step 1: Process the data set and generate four types of auxiliary training benchmark data, including building main body, boundary, separation area between buildings and edge direction, based on the existing building labels;

[0009] Step 2: define a joint cost function based on building extraction and edge extraction, and train a UNet-like network using the data set processed in step 1;

[0010] Step 3: For any image to be segmented, use the UNet-like network trained in step 2 to extract four elements: the building main body heat map, the boundary heat map, the separation area between buildings, and the edge normal direction;

[0011] Step 4: fuse the four elements obtained in step 3 to obtain the edge-enhanced building extraction result.

[0012] Furthermore, the building boundaries are obtained by using the Canny edge detection algorithm on the main body of the building.

[0013] Furthermore, the method for obtaining the separation area between buildings is as follows: mark the main body of the building into different connected areas, use the morphological dilation algorithm to dilate the area by 5-9 pixels, and the parts where the expansion of different areas overlap but are not in the original area are recorded as the building separation area.

[0014] Furthermore, the edge normal direction is obtained by: the boundary points are arranged in a counterclockwise direction, and any two adjacent points p i and p j The vector difference The tangent vector of the point is rotated 90 degrees counterclockwise to obtain That is, the direction of the normal vector of the point pointing inside the region.

[0015] Furthermore, we define the joint cost function L = L b +α e L e +α s L s +α n L n , where L b is the cost function of the building body, L e is the boundary cost function, α e Its weight coefficient, L s Separation region cost function, α s Its weight coefficient, L n Normal cost function, α n is its weight coefficient; L b , L e , L s are the interactive entropy cost functions of the predicted value and the benchmark value, L n is the absolute value loss function in the normal direction are two continuously changing sub-loss functions.

[0016] Furthermore, step 4 uses the watershed algorithm to fuse the four elements obtained in step 3. The height value of the watershed algorithm is h = m b *(1-m e ), the watershed seed area is s = h*(1-m b )*(1-m s )>0.75, the calculation area of ​​the watershed is mask=mask h ∪mask θ , represents the merging of two regions, where mask h =h>0.5 represents the mask obtained by the main body of the building, Represents the mask obtained by expanding the edge inward, where E = m e >0.5 indicates the point set of the valid edge, mask θ It is the set of points obtained by expanding all the points in E toward the inside of the edge by a length of l, where l ranges from 0 to 4.

[0017] The present invention transforms building extraction into a multi-task joint learning problem. It starts from constructing a joint cost function, strengthens the constraints on edges during building extraction, and fuses the inferred edges and building information in the post-processing process to further improve the segmentation accuracy of building edges. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 The four elements of training are obtained by processing the labeled benchmark data;

[0019] Figure 2 It is a schematic diagram of forward reasoning;

[0020] Figure 3 It is a schematic diagram of the fusion processing principle. DETAILED DESCRIPTION

[0021] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. The embodiments of the present invention are provided for the purpose of illustration and description, and are not intended to be exhaustive or to limit the present invention to the disclosed forms. Many modifications and variations will be apparent to those of ordinary skill in the art. The embodiments are selected and described in order to better illustrate the principles and practical applications of the present invention, and to enable those of ordinary skill in the art to understand the present invention and thereby design various embodiments with various modifications suitable for specific uses.

[0022] Since the extraction of buildings and the learning of building edges are highly related, the present invention proposes an algorithm that uses boundary information to improve the accuracy of building edge segmentation, converting building extraction into a multi-task joint learning problem.

[0023] The algorithm for improving the accuracy of building edge segmentation using boundary information mainly includes the following steps:

[0024] 1. Process the data set and generate four types of auxiliary training benchmark data based on the existing building labels: building body, boundary, separation area between buildings and edge direction, such as Figure 1 shown.

[0025] Among them, the building boundary is obtained by using the Canny edge detection algorithm on the main body of the building;

[0026] The method for obtaining the separation area between buildings is as follows: mark the main body of the building into different connected areas, use the morphological dilation algorithm to dilate the area by 5-9 pixels, and the parts of the different areas that overlap but are not in the original area are recorded as the building separation area;

[0027] The method for obtaining the edge normal direction is as follows: the boundary points are arranged in a counterclockwise direction, and any two adjacent points p i and p j The vector difference The tangent vector of the point is rotated 90 degrees counterclockwise to obtain That is, the direction of the normal vector of the point pointing inside the region.

[0028] 2. Define a joint cost function based on building extraction and edge extraction, and use the data set processed in step 1 to train the UNet network. During the training process, use the network forward reasoning, such as Figure 2 As shown, update the network parameters.

[0029] Define the joint cost function L = L b +α e L e +α s L s +α n L n , where L b is the cost function of the building body, L e is the boundary cost function, α e Its weight coefficient, L s Separation region cost function, α s Its weight coefficient, L n Normal cost function, α n is its weight coefficient; L b , L e , L s are the interactive entropy cost functions of the predicted value and the benchmark value, L n is the absolute value loss function in the normal direction are two continuously changing sub-loss functions.

[0030] 3. For any image to be segmented, use the UNet-like network trained in step 2 to extract four elements: the building main body heat map, boundary heat map, separation area between buildings, and edge normal direction.

[0031] 4. The four elements obtained in step 3 are integrated to obtain the edge-enhanced building extraction result.

[0032] The four elements obtained in step 3 are fused by the watershed algorithm, such as Figure 3 As shown, the height value of the watershed algorithm is h = m b *(1-m e ), the watershed seed area is s = h*(1-m b )*(1-m s )>0.75, the calculation area of ​​the watershed is mask=mask h ∪mask θ , represents the merging of two regions, where mask h =h>0.5 represents the mask obtained by the main body of the building, Represents the mask obtained by expanding the edge inward, where E = m e >0.5 indicates the point set of the valid edge, mask θ It is the set of points obtained by expanding all the points in E toward the inside of the edge by a length of l, where l ranges from 0 to 4.

[0033] Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field and related fields without creative work should fall within the scope of protection of the present invention.

Claims

1. An algorithm for improving the accuracy of building edge segmentation using boundary information, characterized in that: The following steps are involved: Step 1: Process the data set and generate four types of auxiliary training benchmark data, including building main body, boundary, separation area between buildings and edge direction, based on the existing building labels; Step 2: define a joint cost function based on building extraction and edge extraction, and train a UNet-like network using the data set processed in step 1; Step 3: For any image to be segmented, use the UNet-like network trained in step 2 to extract four elements: the building main body heat map, the boundary heat map, the separation area between buildings, and the edge normal direction; Step 4: fuse the four elements obtained in step 3 to obtain the edge-enhanced building extraction result.

2. The algorithm for improving the accuracy of building edge segmentation by using boundary information according to claim 1 is characterized in that: The building boundaries are obtained by using the Canny edge detection algorithm on the building body.

3. The algorithm for improving the accuracy of building edge segmentation by using boundary information according to claim 1 is characterized in that: The method for obtaining the separation area between buildings is as follows: mark the main body of the building into different connected areas, use the morphological dilation algorithm to dilate the area by 5-9 pixels, and the parts of different areas that overlap but are not in the original area are recorded as building separation areas.

4. The algorithm for improving the accuracy of building edge segmentation by using boundary information according to claim 1 is characterized in that: The method for obtaining the edge normal direction is as follows: the boundary points are arranged in a counterclockwise direction, and any two adjacent points p i and p j The vector difference The tangent vector of the point is rotated 90 degrees counterclockwise to obtain That is, the direction of the normal vector of the point pointing inside the region.

5. The algorithm for improving the accuracy of building edge segmentation by using boundary information according to claim 1 is characterized in that: Define the joint cost function L = L b +α e L e +α s L s +α n L n , where L b is the cost function of the building body, L e is the boundary cost function, α e Its weight coefficient, L s Separation region cost function, α s Its weight coefficient, L n Normal cost function, α n is its weight coefficient; L b , L e , L s are the interactive entropy cost functions of the predicted value and the benchmark value, L n Absolute value loss function in the normal direction are two continuously changing sub-loss functions.

6. The algorithm for improving the accuracy of building edge segmentation using boundary information according to any one of claims 1 to 5, characterized in that: Step 4 uses the watershed algorithm to fuse the four elements obtained in step 3. The height value of the watershed algorithm is h = m b *(1-m e ), the watershed seed area is s = h*(1-m b )*(1-m s )>0.75, the calculation area of ​​the watershed is mask=mask h ∪mask θ , represents the merging of two regions, where mask h =h>0.5 represents the mask obtained by the main body of the building, represents the mask obtained by expanding the edge inward, where E=m e >0.5 indicates the point set of the valid edge, mask θ It is the set of points obtained by expanding all the points in E toward the inside of the edge by a length of l, where l ranges from 0 to 4.

Citation Information

Patent Citations

  • High-resolution synthetic aperture radar image linear building detecting method based on marked watershed algorithm

    CN104715474A

  • Multi-scale-structure-learning-based method for high-precision farmland boundary extraction with satellite image

    CN108830870A