An algorithm to improve the accuracy of building edge segmentation using boundary information
By using the algorithm of boundary information under the multi-task learning framework, processing data sets and defining joint cost functions to train UNet networks, the problem of insufficient edge segmentation accuracy in the existing technology is solved, and higher edge segmentation accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202111575708.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-12-21
AI Technical Summary
The existing building extraction algorithms with high resolution remote sensing images have a contradiction between accuracy and robustness in edge segmentation. The traditional algorithms have low edge accuracy but low robustness, while the deep learning-based algorithms have blurred edges but high overall matching coefficient.
Under the multi-task learning framework, an algorithm that uses boundary information to improve the accuracy of building edge segmentation is used to generate benchmark data for building main body, boundary, separation area and method direction by processing the data set, and a joint cost function training UNet network is defined, building edge information is extracted, and multiple elements are fused through the watershed algorithm to improve edge segmentation accuracy.
It effectively improves the building edge segmentation accuracy based on high-resolution remote sensing images, combines the edge accuracy of traditional algorithms and the robustness of deep learning algorithms, and improves the overall matching coefficient.
Smart Images

Figure CN114240977B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of building extraction algorithms, and in particular to an algorithm for improving building edge segmentation accuracy by using boundary information. Background Art
[0002] Existing building extraction algorithms are mainly divided into two categories: traditional algorithms and deep learning-based algorithms.
[0003] The traditional algorithm directly detects the edges of building images. Wleite uses random forests to create two binary classifiers. One classifier is used to distinguish whether each pixel belongs to an edge, and the other classifier is used to distinguish whether each pixel is inside or outside the building. Then, based on edge detection and pixel classification, brute force matching is used to perform polygon screening, and each polygon candidate is screened using the IOU value.
[0004] Most deep learning-based algorithms do not use boundary information at all and obtain results directly from image learning; Audebert et al. proposed to use OpenStreetMap information joint learning to improve the segmentation accuracy of buildings, by using some rules to create more labeling constraints, and at the same time using the residential land, agricultural land, industrial land, water area, buildings and roads and other layer information in OpenStreetMap; Marmanis first extracts the edge information and puts it into the neural network as an input together with the image information for training.
[0005] In the building extraction algorithm based on high-resolution remote sensing images, the building edges segmented by the traditional algorithm are more accurate, but the robustness is not high, and the overall matching coefficient IOU is low; the overall matching coefficient of the building extraction algorithm based on deep learning is higher, but because the ordinary convolutional neural network is an inductive deduction of large areas, the extracted building edges are blurred. Summary of the invention
[0006] In response to the technical problems existing in the existing building extraction algorithms for high-resolution remote sensing images, the present invention provides an algorithm that uses boundary information to improve the accuracy of building edge segmentation under a multi-task learning framework, thereby effectively improving the accuracy of building edge segmentation based on high-resolution remote sensing images.
[0007] An algorithm for improving the accuracy of building edge segmentation using boundary information includes the following steps:
[0008] Step 1: Process the data set and generate four types of auxiliary training benchmark data, including building main body, boundary, separation area between buildings and edge direction, based on the existing building labels;
[0009] Step 2: define a joint cost function based on building extraction and edge extraction, and train a UNet-like network using the data set processed in step 1;
[0010] Step 3: For any image to be segmented, use the UNet-like network trained in step 2 to extract four elements: the building main body heat map, the boundary heat map, the separation area between buildings, and the edge normal direction;
[0011] Step 4: fuse the four elements obtained in step 3 to obtain the edge-enhanced building extraction result.
[0012] Furthermore, the building boundaries are obtained by using the Canny edge detection algorithm on the main body of the building.
[0013] Furthermore, the method for obtaining the separation area between buildings is as follows: mark the main body of the building into different connected areas, use the morphological dilation algorithm to dilate the area by 5-9 pixels, and the parts where the expansion of different areas overlap but are not in the original area are recorded as the building separation area.
[0014] Furthermore, the edge normal direction is obtained by: the boundary points are arranged in a counterclockwise direction, and any two adjacent points p i and p j The vector difference The tangent vector of the point is rotated 90 degrees counterclockwise to obtain That is, the direction of the normal vector of the point pointing inside the region.
[0015] Furthermore, we define the joint cost function L = L b +α e L e +α s L s +α n L n , where L b is the cost function of the building body, L e is the boundary cost function, α e Its weight coefficient, L s Separation region cost function, α s Its weight coefficient, L n Normal cost function, α n is its weight coefficient; L b , L e , L s are the interactive entropy cost functions of the predicted value and the benchmark value, L n is the absolute value loss function in the normal direction are two continuously changing sub-loss functions.
[0016] Furthermore, step 4 uses the watershed algorithm to fuse the four elements obtained in step 3. The height value of the watershed algorithm is h = m b *(1-m e ), the watershed seed area is s = h*(1-m b )*(1-m s )>0.75, the calculation area of the watershed is mask=mask h ∪mask θ , represents the merging of two regions, where mask h =h>0.5 represents the mask obtained by the main body of the building, Represents the mask obtained by expanding the edge inward, where E = m e >0.5 indicates the point set of the valid edge, mask θ It is the set of points obtained by expanding all the points in E toward the inside of the edge by a length of l, where l ranges from 0 to 4.
[0017] The present invention transforms building extraction into a multi-task joint learning problem. It starts from constructing a joint cost function, strengthens the constraints on edges during building extraction, and fuses the inferred edges and building information in the post-processing process to further improve the segmentation accuracy of building edges. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 The four elements of training are obtained by processing the labeled benchmark data;
[0019] Figure 2 It is a schematic diagram of forward reasoning;
[0020] Figure 3 It is a schematic diagram of the fusion processing principle. DETAILED DESCRIPTION
[0021] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. The embodiments of the present invention are provided for the purpose of illustration and description, and are not intended to be exhaustive or to limit the present invention to the disclosed forms. Many modifications and variations will be apparent to those of ordinary skill in the art. The embodiments are selected and described in order to better illustrate the principles and practical applications of the present invention, and to enable those of ordinary skill in the art to understand the present invention and thereby design various embodiments with various modifications suitable for specific uses.
[0022] Since the extraction of buildings and the learning of building edges are highly related, the present invention proposes an algorithm that uses boundary information to improve the accuracy of building edge segmentation, converting building extraction into a multi-task joint learning problem.
[0023] The algorithm for improving the accuracy of building edge segmentation using boundary information mainly includes the following steps:
[0024] 1. Process the data set and generate four types of auxiliary training benchmark data based on the existing building labels: building body, boundary, separation area between buildings and edge direction, such as Figure 1 shown.
[0025] Among them, the building boundary is obtained by using the Canny edge detection algorithm on the main body of the building;
[0026] The method for obtaining the separation area between buildings is as follows: mark the main body of the building into different connected areas, use the morphological dilation algorithm to dilate the area by 5-9 pixels, and the parts of the different areas that overlap but are not in the original area are recorded as the building separation area;
[0027] The method for obtaining the edge normal direction is as follows: the boundary points are arranged in a counterclockwise direction, and any two adjacent points p i and p j The vector difference The tangent vector of the point is rotated 90 degrees counterclockwise to obtain That is, the direction of the normal vector of the point pointing inside the region.
[0028] 2. Define a joint cost function based on building extraction and edge extraction, and use the data set processed in step 1 to train the UNet network. During the training process, use the network forward reasoning, such as Figure 2 As shown, update the network parameters.
[0029] Define the joint cost function L = L b +α e L e +α s L s +α n L n , where L b is the cost function of the building body, L e is the boundary cost function, α e Its weight coefficient, L s Separation region cost function, α s Its weight coefficient, L n Normal cost function, α n is its weight coefficient; L b , L e , L s are the interactive entropy cost functions of the predicted value and the benchmark value, L n is the absolute value loss function in the normal direction are two continuously changing sub-loss functions.
[0030] 3. For any image to be segmented, use the UNet-like network trained in step 2 to extract four elements: the building main body heat map, boundary heat map, separation area between buildings, and edge normal direction.
[0031] 4. The four elements obtained in step 3 are integrated to obtain the edge-enhanced building extraction result.
[0032] The four elements obtained in step 3 are fused by the watershed algorithm, such as Figure 3 As shown, the height value of the watershed algorithm is h = m b *(1-m e ), the watershed seed area is s = h*(1-m b )*(1-m s )>0.75, the calculation area of the watershed is mask=mask h ∪mask θ , represents the merging of two regions, where mask h =h>0.5 represents the mask obtained by the main body of the building, Represents the mask obtained by expanding the edge inward, where E = m e >0.5 indicates the point set of the valid edge, mask θ It is the set of points obtained by expanding all the points in E toward the inside of the edge by a length of l, where l ranges from 0 to 4.
[0033] Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field and related fields without creative work should fall within the scope of protection of the present invention.
Claims
1. An algorithm for improving the accuracy of building edge segmentation using boundary information, characterized in that: The following steps are involved: Step 1: Process the data set and generate four types of auxiliary training benchmark data, including building main body, boundary, separation area between buildings and edge direction, based on the existing building labels; Step 2: define a joint cost function based on building extraction and edge extraction, and train a UNet-like network using the data set processed in step 1; Step 3: For any image to be segmented, use the UNet-like network trained in step 2 to extract four elements: the building main body heat map, the boundary heat map, the separation area between buildings, and the edge normal direction; Step 4: fuse the four elements obtained in step 3 to obtain the edge-enhanced building extraction result.
2. The algorithm for improving the accuracy of building edge segmentation by using boundary information according to claim 1 is characterized in that: The building boundaries are obtained by using the Canny edge detection algorithm on the building body.
3. The algorithm for improving the accuracy of building edge segmentation by using boundary information according to claim 1 is characterized in that: The method for obtaining the separation area between buildings is as follows: mark the main body of the building into different connected areas, use the morphological dilation algorithm to dilate the area by 5-9 pixels, and the parts of different areas that overlap but are not in the original area are recorded as building separation areas.
4. The algorithm for improving the accuracy of building edge segmentation by using boundary information according to claim 1 is characterized in that: The method for obtaining the edge normal direction is as follows: the boundary points are arranged in a counterclockwise direction, and any two adjacent points p i and p j The vector difference The tangent vector of the point is rotated 90 degrees counterclockwise to obtain That is, the direction of the normal vector of the point pointing inside the region.
5. The algorithm for improving the accuracy of building edge segmentation by using boundary information according to claim 1 is characterized in that: Define the joint cost function L = L b +α e L e +α s L s +α n L n , where L b is the cost function of the building body, L e is the boundary cost function, α e Its weight coefficient, L s Separation region cost function, α s Its weight coefficient, L n Normal cost function, α n is its weight coefficient; L b , L e , L s are the interactive entropy cost functions of the predicted value and the benchmark value, L n Absolute value loss function in the normal direction are two continuously changing sub-loss functions.
6. The algorithm for improving the accuracy of building edge segmentation using boundary information according to any one of claims 1 to 5, characterized in that: Step 4 uses the watershed algorithm to fuse the four elements obtained in step 3. The height value of the watershed algorithm is h = m b *(1-m e ), the watershed seed area is s = h*(1-m b )*(1-m s )>0.75, the calculation area of the watershed is mask=mask h ∪mask θ , represents the merging of two regions, where mask h =h>0.5 represents the mask obtained by the main body of the building, represents the mask obtained by expanding the edge inward, where E=m e >0.5 indicates the point set of the valid edge, mask θ It is the set of points obtained by expanding all the points in E toward the inside of the edge by a length of l, where l ranges from 0 to 4.
Citation Information
Patent Citations
High-resolution synthetic aperture radar image linear building detecting method based on marked watershed algorithm
CN104715474A
Multi-scale-structure-learning-based method for high-precision farmland boundary extraction with satellite image
CN108830870A