Building automatic extraction method and device and electronic equipment
By using mask and frame field information in the building extraction model, combining the active skeleton model and graph cutting method, the problem of difficulty in achieving high automation and high edge accuracy in the prior art is solved, and efficient building extraction is achieved.
Patent Information
- Application Number
- CN202510149834.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-10
AI Technical Summary
Existing building extraction methods are difficult to achieve high automation and high edge accuracy at the same time.
By obtaining the target remote sensing image, input it into the building extraction model, the building is initially extracted and graphically optimized using the active skeleton model and graph cutting method, and model training is performed in combination with mask and frame field information.
It realizes high automation and high edge accuracy building extraction, and combines deep learning and map cutting technology to optimize the extraction effect of building map spots.
Smart Images

Figure CN120126019A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a method for automatically extracting buildings, a device for automatically extracting buildings, an electronic device, a machine-readable storage medium, and a computer program product. Background Art
[0002] Building extraction has important value both in scientific research and production applications. Building extraction methods can be divided into three categories according to the degree of automation, namely manual annotation, semi-automatic interpretation, and automatic interpretation. Among them, manual annotation has the highest accuracy, but it is extremely labor-intensive and difficult to carry out large-scale production applications; semi-automatic interpretation uses interactive semi-automatic technology to extract the initial building contour, and combined with regularization processing technology, relatively fast and high-precision building extraction can be achieved, but still requires a certain amount of human input; automatic interpretation uses automated technologies such as deep learning to extract the initial building contour and combines regularization processing technology to achieve interpretation, which has high efficiency, but the overall accuracy still needs to be improved.
[0003] That is, the existing building extraction methods are difficult to achieve building extraction with both high automation and high edge accuracy simultaneously. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a method, device, and electronic device for automatically extracting buildings, so as to solve the defect that the existing building extraction methods are difficult to achieve building extraction with both high automation and high edge accuracy simultaneously.
[0005] To achieve the above purpose, the embodiments of the present invention provide a method for automatically extracting buildings, including:
[0006] Obtain a target remote sensing image;
[0007] Input the target remote sensing image into a building extraction model to obtain a mask image and a frame field image of the target remote sensing image output by the building extraction model;
[0008] Use an active skeleton model to construct a preliminary building extraction result of the target remote sensing image based on the mask image and the frame field image;
[0009] Use a graph cut method to optimize the graph of the preliminary building extraction result to obtain a building extraction result;
[0010] Wherein, the building extraction model is trained based on a sample remote sensing image, a mask annotation of the sample remote sensing image, and a frame field annotation.
[0011] Optionally, the building extraction model is trained through the following steps:
[0012] Repeat the following steps until the set stop condition is reached:
[0013] Input the sample remote sensing image, the mask annotation of the sample remote sensing image, and the frame field annotation into the building extraction model to obtain the sample mask prediction image and the sample frame field prediction image output by the building extraction model;
[0014] Calculate the mask loss function of the building extraction model based on the mask annotation and the sample mask prediction image, and calculate the frame field loss function of the building extraction model based on the frame field annotation and the sample frame field prediction image;
[0015] Calculate the total loss function of the building extraction model based on the mask loss function and the frame field loss function;
[0016] Adjust the model parameters of the building extraction model based on the total loss function.
[0017] Optionally, calculating the mask loss function of the building extraction model based on the mask annotation and the sample mask prediction image includes:
[0018] Calculate the cross-entropy loss function and the Dice loss function of the building extraction model based on the mask annotation and the sample mask prediction image;
[0019] Calculate the mask loss function of the building extraction model based on the cross-entropy loss function and the Dice loss function.
[0020] Optionally, calculating the frame field loss function of the building extraction model based on the frame field annotation and the sample frame field prediction image includes:
[0021] Calculate the horizontal angle loss of the unsigned tangent vector of the polygon contour, the vertical angle loss of the unsigned tangent vector of the polygon contour, and the smoothing loss of the building extraction model based on the frame field annotation and the sample frame field prediction image;
[0022] Calculate the frame field loss function of the building extraction model based on the horizontal angle loss of the unsigned tangent vector of the polygon contour, the vertical angle loss of the unsigned tangent vector of the polygon contour, and the smoothing loss.
[0023] Optionally, using the graph cut method to perform graph optimization on the preliminary building extraction result to obtain the building extraction result includes:
[0024] Perform erosion operation and dilation operation on the preliminary building extraction result to obtain an image including multiple building patches;
[0025] Generate horizontal lines, vertical lines, acute - angled diagonals, and obtuse - angled diagonals as prior seed lines within each of the building patches respectively, and generate background lines outside each of the building patches respectively;
[0026] Perform graph cuts on the four prior seed lines within each of the building patches respectively to obtain the segmentation results of the four prior seed lines;
[0027] Perform average fusion on the segmentation results of the four prior seed lines within each of the building patches respectively to obtain the building extraction result.
[0028] Optionally, the performing graph cuts on the four prior seed lines within each of the building patches respectively to obtain the segmentation results of the four prior seed lines includes:
[0029] Determine the pixels of the first prior seed line within the first building patch as the point set of foreground targets, and randomly select the same number of pixels as the pixels of the first prior seed line from the background line corresponding to the first prior seed line as the point set of background targets;
[0030] Construct Gaussian mixture models respectively based on the point set of foreground targets and the point set of background targets;
[0031] Construct an energy function based on the Gaussian mixture model of foreground targets and the Gaussian mixture model of background targets;
[0032] Use the min - cut max - flow algorithm to minimize the energy function to obtain the segmentation result of the first prior seed line;
[0033] Wherein, the first building patch is any one of the multiple building patches, and the first prior seed line is any one of the four prior seed lines.
[0034] Optionally, after using the graph cut method to perform graphic optimization on the preliminary building extraction result to obtain the building extraction result, it further includes:
[0035] Perform at least one regularization process of right - angled polygon fitting, automatic complementary angles of right - angled polygons, and processing of gaps between adjacent buildings on the building extraction result.
[0036] Optionally, the performing right - angled polygon fitting on the building extraction result includes:
[0037] Rotate the first building image of the building extraction result at multiple angles within a set angle range respectively, and determine the fitting error between the polygon fitting result of the first building image at each angle and the fitting result of the unrotated first building image; wherein, the first building image is any one of the building images in the building extraction result;
[0038] Determine the rotation angle corresponding to the minimum fitting error as the optimal rotation angle, and perform rotation of the optimal rotation angle on the first building image;
[0039] Perform right-angled polygon fitting on the first building image rotated by the optimal rotation angle by using the right-angled polygon fitting method;
[0040] Perform reverse rotation of the optimal rotation angle on the first building image after right-angled polygon fitting, so that the first building image after right-angled polygon fitting is restored to the angle without performing rotation.
[0041] Optionally, the determining the fitting error between the polygon fitting result of the first building image at each of the angles and the first building image without rotation includes:
[0042] Based on the number of fitting line segments of the first building image, the fitting weight, and the sum of the distances between all the original line segments of the first building image and the fitting line segments, determine the fitting error between the polygon fitting result of the first building image at each of the angles and the first building image without rotation.
[0043] Optionally, the determining the fitting error between the polygon fitting result of the first building image at each of the angles and the first building image without rotation based on the number of fitting line segments of the first building image, the fitting weight, and the sum of the distances between all the original line segments of the first building image and the fitting line segments of the original line segments includes:
[0044]
[0045] where m is the number of fitting line segments of the first building image, represents the sum of the distances between the line segments on the j-th segment of the original contour line of the first building image and the fitting line segments after fitting; △ is the fitting weight, F is the fitting error, and {p} represents the vector nodes on the original contour line of the first building image.
[0046] On the other hand, an embodiment of the present invention further provides a building automatic extraction device, including:
[0047] An acquisition module, configured to acquire a target remote sensing image;
[0048] A prediction module, configured to input the target remote sensing image into a building extraction model, and obtain a mask image and a frame field image of the target remote sensing image output by the building extraction model;
[0049] A construction module for constructing a preliminary extraction result of a building in a target remote sensing image based on the mask image and the frame field image by using an active skeleton model;
[0050] An extraction module for optimizing the graph of the preliminary extraction result of the building by using a graph cut method to obtain an extraction result of the building;
[0051] Wherein, the building extraction model is trained based on a sample remote sensing image, a mask annotation of the sample remote sensing image, and a frame field annotation.
[0052] On the other hand, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-mentioned automatic building extraction method is implemented.
[0053] On the other hand, the present invention also provides a machine-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned automatic building extraction method is implemented.
[0054] On the other hand, the present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the above-mentioned automatic building extraction method is implemented.
[0055] Through the above technical solutions, the building extraction model in the embodiments of the present invention is trained based on a sample remote sensing image, a mask annotation of the sample remote sensing image, and a frame field annotation. Since both a mask representing the surface information of a building patch and a frame field representing the boundary and corner information of a relatively smooth building patch are used for model training, the effect of using an automated technique to extract building patches can be optimized. On this basis, the graph of the preliminary extraction result of the building is further optimized by a graph cut method based on semi-automatic extraction to finally obtain an extraction result of the building. The embodiments of the present invention effectively combine an automatic building extraction method based on deep learning and a semi-automatic building extraction method based on graph cut, realizing high-automation and high-edge-precision building extraction.
[0056] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The drawings are used to provide a further understanding of the embodiments of the present invention, and constitute a part of the specification. Together with the following specific implementation, they are used to explain the embodiments of the present invention, but do not constitute a limitation to the embodiments of the present invention. In the drawings:
[0058] Figure 1 is a schematic flowchart of the automatic building extraction method provided by the present invention;
[0059] Figure 2 It is a schematic structural diagram of the encoder-decoder network structure selected by the building extraction model provided by the present invention;
[0060] Figure 3 It is a schematic structural diagram of the MBConv module provided by the present invention;
[0061] Figure 4 It is a schematic structural diagram of the Fused-MBConv module provided by the present invention;
[0062] Figure 5 It is a schematic structural diagram of the mask annotation provided by the present invention;
[0063] Figure 6 It is a schematic structural diagram of the corners and edges of the grid provided by the present invention;
[0064] Figure 7 It is a schematic structural diagram of the frame field provided by the present invention;
[0065] Figure 8 It is a schematic flowchart of performing right-angled polygon fitting on the building extraction result provided by the present invention;
[0066] Figure 9 It is one of the schematic diagrams of the effect of performing right-angled polygon fitting on the building extraction result provided by the present invention;
[0067] Figure 10 It is another schematic diagram of the effect of performing right-angled polygon fitting on the building extraction result provided by the present invention;
[0068] Figure 11 It is the third schematic diagram of the effect of performing right-angled polygon fitting on the building extraction result provided by the present invention;
[0069] Figure 12 It is the fourth schematic diagram of the effect of performing right-angled polygon fitting on the building extraction result provided by the present invention;
[0070] Figure 13 It is the fifth schematic diagram of the effect of performing right-angled polygon fitting on the building extraction result provided by the present invention;
[0071] Figure 14 It is the sixth schematic diagram of the effect of performing right-angled polygon fitting on the building extraction result provided by the present invention;
[0072] Figure 15 It is the seventh schematic diagram of the effect of performing right-angled polygon fitting on the building extraction result provided by the present invention;
[0073] Figure 16It is a schematic flow chart of automatically supplementing right-angled polygons for the building extraction results provided by the present invention;
[0074] Figure 17 It is a schematic diagram of a flat angle provided by the present invention;
[0075] Figure 18 It is one of the schematic diagrams of the effect of automatically supplementing right-angled polygons provided by the present invention;
[0076] Figure 19 It is the second schematic diagram of the effect of automatically supplementing right-angled polygons provided by the present invention;
[0077] Figure 20 It is a schematic flow chart of the method for processing the gaps between adjacent buildings provided by the present invention;
[0078] Figure 21 It is a schematic diagram of one of the matching types of the method for processing the gaps between adjacent buildings provided by the present invention;
[0079] Figure 22 It is a schematic diagram of the second matching type of the method for processing the gaps between adjacent buildings provided by the present invention;
[0080] Figure 23 It is a schematic diagram of the third matching type of the method for processing the gaps between adjacent buildings provided by the present invention;
[0081] Figure 24 It is one of the schematic diagrams of the unilateral adsorption matching effect provided by the present invention;
[0082] Figure 25 It is the second schematic diagram of the unilateral adsorption matching effect provided by the present invention;
[0083] Figure 26 It is the third schematic diagram of the unilateral adsorption matching effect provided by the present invention;
[0084] Figure 27 It is a schematic diagram of the multi-sided adsorption matching effect provided by the present invention;
[0085] Figure 28 It is a schematic diagram of the deep learning automatic extraction result predicted by the building extraction model provided by the present invention;
[0086] Figure 29 It is a schematic diagram of the building extraction result provided by the present invention;
[0087] Figure 30 It is a schematic diagram of the building automatic extraction device provided by the present invention;
[0088] Figure 31 It is a schematic diagram of the electronic device provided by the present invention. Detailed implementation manners
[0089] The following details the specific implementation manners of the embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific implementation manners described herein are only for explaining and understanding the embodiments of the present invention, and are not used to limit the embodiments of the present invention.
[0090] Method embodiment
[0091] Please refer to Figure 1 , an automatic building extraction method is proposed in an embodiment of the present invention, which is characterized by including:
[0092] Step 100: Obtain a target remote sensing image.
[0093] An electronic device obtains a target remote sensing image of a target to be extracted. The target to be extracted may be a target with a regular shape such as a right angle, for example, a vehicle and a building, etc. In the embodiment of the present invention, the target to be extracted is mainly described by taking the automatic building extraction as an example.
[0094] Step 200: Input the target remote sensing image into a building extraction model to obtain a mask image and a frame field image of the target remote sensing image output by the building extraction model.
[0095] An electronic device inputs the target remote sensing image into a building extraction model to obtain a mask image and a frame field image of the target remote sensing image output by the building extraction model. The main difficulties in building extraction are difficult patch separation and low boundary accuracy. Therefore, in the embodiment of the present invention, building surface, corner point and boundary information are used at the same time, and the building extraction model is designed as a multi-task building extraction network with contour optimization to optimize the building patch extraction effect.
[0096] The building extraction model with contour optimization in the embodiment of the present invention adds additional boundary, corner point and related supervision information compared with the general model, but it is difficult to accurately learn directly using the boundary and corner points. In particular, the frame field can express boundary and corner point information relatively smoothly at the same time. In the frame field. The frame field is a vector field composed of elements representing the boundary point direction in a four-way vector. For each pixel, the frame field is represented as a set of vectors {u, -u, v, -v}. To avoid confusion of symbols and orders, the following formula is used for substitution.
[0097] c 0 = u 2 v 2 ;
[0098] c 2 = -(u 2 + v 2 );
[0099] Please refer to Figure 2, the building extraction model of the embodiment of the present invention can be constructed by using the network structure of an encoder-decoder. Among them, the encoder can adopt ResNeSt with a split attention mechanism to extract richer features. Among them, modules with attention mechanisms and residual characteristics are widely used. For example, Figure 3 the MBConv module of Figure 4 and the Fused-MBConv module of
[0100] use the SE module and residual connection, which can efficiently extract features. In the decoder part, multi-scale feature fusion prediction can be used to obtain more accurate and complete prediction results. For multi-scale feature fusion, first, the low-resolution scale features are upsampled to the features with 1 / 4 of the original image resolution, then multiple scales are directly concatenated or fused using attention convolution, and finally, the segmentation prediction module is used for the fused features to obtain the prediction results. Figures 5 to 7 The following Figure 5 is the building mask annotation, Figure 6 are the grid corner points and edges, Figure 7 is the frame field of the boundary. The process of frame field annotation conversion is as follows: obtain the corner points and edges of the annotated polygon in the grid according to the mask annotation, and calculate the frame field information of the points on the boundary through the direction of the edges.
[0101] In one embodiment, during model training, the inputs of the building extraction model include the sample remote sensing image, the mask annotation of the sample remote sensing image, and the frame field annotation of the sample remote sensing image. The outputs of the building extraction model include mask prediction and frame field prediction. In other embodiments, the inputs of the building extraction model include the sample remote sensing image, the mask annotation of the sample remote sensing image, the boundary annotation of the sample remote sensing image, and the frame field annotation of the sample remote sensing image. The outputs of the building extraction model include mask prediction, boundary prediction, and frame field prediction. Then, based on the prediction results and the model inputs, the loss function is calculated, and the model parameters of the building extraction model are adjusted based on the loss function. The trained building extraction model can perform building extraction on the input target remote sensing image to obtain the mask image and frame field image of the target remote sensing image.
[0102] Step 300, use the active skeleton model to construct the preliminary building extraction result of the target remote sensing image based on the mask image and the frame field image.
[0103] The building extraction model predicts and obtains a mask image and a frame field image simultaneously, and uses an Active Shape Model (ASM) to construct a preliminary building extraction result of the target remote sensing image based on the mask image and the frame field image. Among them, the establishment of the active shape model includes two stages: shape modeling and shape matching. In the shape modeling stage, the building key points in the training samples are selected and marked, and then statistical analysis is performed on these key points to obtain the average shape and covariance information of shape changes. In the shape matching stage, the active shape model is matched with the mask image and the frame field image to be detected. By iteratively adjusting the shape parameters, the model is made to match the building shape in the image as much as possible. The active shape model can be used to construct the boundary and vertices of the building polygon and simplify the vertices that are not corner points, so as to obtain the preliminary building extraction result.
[0104] Step 400: Use the graph cut method to optimize the graph of the preliminary building extraction result to obtain the building extraction result.
[0105] Although the deep learning model can provide a comprehensive building extraction result, limited by the quality of the labeled samples and the limitations of the image itself, in most cases, the boundary is not fine enough and further optimization is required. In the embodiment of the present invention, the graph cut method is used to optimize the graph of the preliminary building extraction result, and a segmentation result with high boundary quality can be obtained through the guidance of the prior knowledge of the graph cut method, which can be used as an effective means for boundary optimization. In one embodiment, Step 400: Use the graph cut method to optimize the graph of the preliminary building extraction result to obtain the building extraction result, including:
[0106] Step 410: Perform an erosion operation and a dilation operation on the preliminary building extraction result to obtain an image including multiple building patches.
[0107] The erosion operation can remove the small parts at the building edge, such as noise points or small protrusions, so that the building patches are smoother and more regular. First, a structuring element (usually a small matrix) is defined, and its shape and size determine the effect of the erosion operation. The structuring element is slid on the preliminary building extraction result, and for each position, the intersection of the structuring element and the building extraction result is calculated. If the intersection is not empty, the pixel at that position is retained; otherwise, the pixel at that position is set to the background. Thus, the noise points or small protrusions at the building edge can be removed through the erosion operation.
[0108] The dilation operation can fill the holes inside the building and expand the boundaries of the building, thus making the building patches more complete and continuous. First, a structuring element (also a small matrix) is defined, and its shape and size determine the effect of the dilation operation. The structuring element is slid over the preliminary building extraction result. For each position, the union of the structuring element and the building extraction result is calculated. If the union is not empty, the pixel at that position is marked as a building; otherwise, it remains as the background. Thus, through the dilation operation, the building patches can be made more complete and continuous.
[0109] Therefore, in the embodiment of the present invention, an erosion operation and a dilation operation are performed on the preliminary building extraction result to obtain an image including multiple building patches.
[0110] Step 420: Horizontallines, vertical lines, acute diagonal lines, and obtuse diagonal lines are respectively generated as prior seed lines within each of the building patches, and background lines are respectively generated outside each of the building patches.
[0111] The electronic device respectively generates horizontal lines, vertical lines, acute diagonal lines, and obtuse diagonal lines as prior seed lines within each of the building patches, and background lines are respectively generated outside each of the building patches. The prior seed lines refer to the lines used to segment the target within the patch, and the background lines refer to the lines in the outside (background) of the patch.
[0112] Step 430: Graph cuts are respectively performed on the four prior seed lines within each of the building patches to obtain the segmentation results of the four prior seed lines.
[0113] The electronic device respectively performs graph cuts on the four prior seed lines within each of the building patches to obtain the segmentation results of the four prior seed lines. In one embodiment, Step 430: Graph cuts are respectively performed on the four prior seed lines within each of the building patches to obtain the segmentation results of the four prior seed lines, including:
[0114] Step 431: Determine the pixels of the first prior seed line within the first building patch as the point set of the foreground target, and randomly select the same number of pixels as the pixels of the first prior seed line from the background line corresponding to the first prior seed line as the point set of the background target.
[0115] Step 432: Build a mixture Gaussian model respectively based on the point set of the foreground target and the point set of the background target.
[0116] Step 433: Build an energy function based on the mixture Gaussian model of the foreground target and the mixture Gaussian model of the background target.
[0117] Step 434: Use the minimum cut maximum flow algorithm to minimize the energy function to obtain the segmentation result of the first prior seed line.
[0118] Among them, the first building patch is any one of the multiple building patches, and the first prior seed line is any one of the four prior seed lines. Obtain the pixels of the first prior seed line within the first building patch as the point set of the foreground target, and randomly select an equal number of pixels from the background line corresponding to the first prior seed line to define as the point set of the background target; respectively construct corresponding Gaussian mixture models according to the point sets of the foreground target and the background target to describe the feature information of the corresponding regions. At this time, according to the Gaussian mixture models of the foreground and background targets, construct an energy function for graph cut:
[0119] E(α,k,θ,z)=U(α,k,θ,z)+V(α,z);
[0120] Among them, k represents the number of single Gaussian models in the Gaussian mixture model, α represents the foreground and background labels, z represents the image data, and θ represents the parameters of the Gaussian mixture model, where
[0121] θ={π(α,k),μ(α,k),Σ(α,k)};
[0122] Among them, π represents the weight of each component of the Gaussian mixture model, and μ and Σ respectively represent the mean and covariance matrix of the Gaussian function.
[0123] Define the data term as follows:
[0124]
[0125] Among them, n is the nth pixel number, and det is the determinant of the calculation matrix.
[0126] Define the boundary term as follows:
[0127]
[0128] Among them, γ and β are both weight coefficients, n is the nth pixel, m is the mth neighboring pixel of n, and dis(m,n) is the distance between pixels m and n.
[0129] By constructing the data term and the boundary term in the energy function as shown above, and then using the minimum cut maximum flow algorithm to minimize the energy function to obtain the segmentation result of the foreground and background.
[0130] Step 440: Respectively perform average fusion on the segmentation results of the four prior seed lines within each of the building patches to obtain the building extraction result.
[0131] The building extraction result is obtained by averaging and fusing the segmentation results of the four prior seed lines within each of the building patches. After the graph cut of the foreground and background targets is completed, the initially extracted building result of the input can be segmented into two regions, the foreground and the background. According to the foreground targets corresponding to the graph cut in the region where the prior seed lines are located, the initial segmentation of the building contour is realized, and the vector information of the building contour is obtained by calling the fast raster vectorization algorithm.
[0132] The building extraction model of the embodiment of the present invention is trained based on the sample remote sensing image, the mask annotation of the sample remote sensing image, and the frame field annotation. Since both the mask representing the surface information of the building patch and the frame field representing the boundary and corner information of the building patch relatively smoothly are used for model training, the effect of extracting the building patch using the automated technology can be optimized. On this basis, the initially extracted building result is graphically optimized by the graph cut method based on semi-automatic extraction, and finally the building extraction result is obtained. The embodiment of the present invention effectively combines the automatic building extraction method based on deep learning and the semi-automatic building extraction method based on graph cut, and realizes the building extraction with high automation and high edge accuracy.
[0133] In other aspects of the embodiment of the present invention, the building extraction model is obtained by training through the following steps:
[0134] Repeat the following steps until the set stop condition is reached:
[0135] Step 10: Input the sample remote sensing image, the mask annotation of the sample remote sensing image, and the frame field annotation into the building extraction model to obtain the sample mask prediction image and the sample frame field prediction image output by the building extraction model.
[0136] Input the sample remote sensing image, the mask annotation of the sample remote sensing image, and the frame field annotation into the building extraction model to obtain the sample mask prediction image and the sample frame field prediction image output by the building extraction model.
[0137] Step 20: Calculate the mask loss function of the building extraction model based on the mask annotation and the sample mask prediction image, and calculate the frame field loss function of the building extraction model based on the frame field annotation and the sample frame field prediction image.
[0138] In one embodiment, calculating the mask loss function of the building extraction model based on the mask annotation and the sample mask prediction image includes: calculating the cross-entropy loss function and the Dice loss function of the building extraction model based on the mask annotation and the sample mask prediction image; calculating the mask loss function of the building extraction model based on the cross-entropy loss function and the Dice loss function.
[0139] Since mask annotation and frame field annotation are input to the building extraction model, correspondingly, mask loss and frame field loss need to be introduced simultaneously in the optimization of the building extraction model. The mask loss adopts a linear combination of the cross-entropy loss function and the Dice loss function. That is, the embodiment of the present invention calculates the mask loss function of the building extraction model based on the sum of the cross-entropy loss function and the Dice loss function.
[0140] The calculation formulas of the cross-entropy loss function and the Dice loss function are respectively:
[0141] Cross-entropy loss function: Dice loss function:
[0142] where β = |Y - | / (|Y + | + |Y - |) and 1 - β = |Y + | / (|Y + | + |Y - |), |Y + | and |Y - | are the number of changed and unchanged pixels statistically based on mask annotation and sample mask prediction images; Pr(y j ) is the sigmoid output for pixel j. Pr(y j = 1) is the sigmoid output for pixel j = 1. Pr(y j = 0) is the sigmoid output for pixel j = 0. |Y| represents the number of pixels in the mask annotation, and |Y′| represents the number of pixels in the sample mask prediction image. |Y′∩Y| represents the number of pixels where the sample mask prediction image overlaps with the mask annotation. L bce represents the cross-entropy loss function, and L dice represents the Dice loss function.
[0143] Calculating the frame field loss function of the building extraction model based on the frame field annotation and the sample frame field prediction image includes:
[0144] Calculating the horizontal angle loss of the unsigned tangent vector of the polygon contour, the vertical angle loss of the unsigned tangent vector of the polygon contour, and the smoothing loss of the building extraction model based on the frame field annotation and the sample frame field prediction image;
[0145] Calculating the frame field loss function of the building extraction model based on the horizontal angle loss of the unsigned tangent vector of the polygon contour, the vertical angle loss of the unsigned tangent vector of the polygon contour, and the smoothing loss.
[0146] Among them, the horizontal angle loss of the unsigned tangent vector of the polygon contour, the vertical angle loss of the unsigned tangent vector of the polygon contour, and the smoothing loss are respectively expressed as:
[0147]
[0148] Among them, L align refers to the horizontal angle loss of the unsigned tangent vector of the polygon contour, L align90 refers to the vertical angle loss of the unsigned tangent vector of the polygon contour, L smooth refers to the smoothing loss, H and W are respectively the height and width of the image annotated by the frame field, x is the x-th pixel, I is the set of all pixels in the image, and y edge (x) is the prediction of whether the x-th pixel is a boundary, and are the outputs of the frame field, that is, the sample frame field prediction image, and are the moduli in the horizontal and vertical directions respectively, together characterizing the direction of the pixel point prediction frame field, and θ τ is the horizontal angle of the unsigned tangent vector of the polygon contour, that is, the ground truth, and its vector expression is (c 0 , c 2 ), θ τ⊥ is the vertical angle of the unsigned tangent vector of the polygon contour, and f is the distance between the frame field prediction vector of the sample frame field prediction image and the true angle, which is calculated as follows:
[0149]
[0150] In one embodiment, the frame field loss function is calculated based on the sum of the horizontal angle loss of the unsigned tangent vector of the polygon contour, the vertical angle loss of the unsigned tangent vector of the polygon contour, and the smoothing loss.
[0151] Step 30: Calculate the total loss function of the building extraction model based on the mask loss function and the frame field loss function.
[0152] The present invention calculates the total loss function of the building extraction model based on the sum of the mask loss function and the frame field loss function. In other embodiments, the total loss function of the building extraction model can also be calculated based on the weighted sum of the mask loss function and the frame field loss function.
[0153] Step 40: Adjust the model parameters of the building extraction model based on the total loss function.
[0154] The specific electronic device calculates the gradients of each weight and bias in the building extraction model using the backpropagation algorithm based on the total loss function. By repeatedly executing step 10 to step 40 until the building extraction model converges, at this time, the target remote sensing image is input into the building extraction model, and the mask image and frame field image of the target remote sensing image output by the building extraction model are obtained.
[0155] The building extraction model of the embodiment of the present invention is trained based on the sample remote sensing image, the mask annotation of the sample remote sensing image, and the frame field annotation. The present invention simultaneously uses the mask representing the surface information of the building patch and uses the relatively smooth frame field expressing the boundary and corner information of the building patch for model training, which can optimize the effect of extracting the building patch using automated technology.
[0156] In other aspects of the embodiment of the present invention, after step 400, using the graph cut method to optimize the graph of the preliminary building extraction result to obtain the building extraction result, it further includes:
[0157] Step 500, performing at least one regularization process of right-angled polygon fitting, automatic corner complementation of right-angled polygons, and adjacent building gap processing on the building extraction result.
[0158] The automatic building extraction technology based on deep learning realizes the extraction of the building graph. However, due to the influence of noise and pixelation in the image, the extraction result has obvious pixel sawtooth results; in practice, the building contour is usually a simple polygon with 90° corners. Therefore, the extracted graph result is input into the right-angled polygon fitting method of the building contour in the embodiment of the present invention to obtain an optimized semi-automatic building extraction vector result.
[0159] In one embodiment, the performing right-angled polygon fitting on the building extraction result includes:
[0160] Step 511, rotating the first building image of the building extraction result at multiple angles within a set angle range, and determining the fitting error between the polygon fitting result of the first building image at each angle and the fitting result of the unrotated first building image; wherein, the first building image is any one building image in the building extraction result.
[0161] Step 512, determining the rotation angle corresponding to the smallest fitting error as the optimal rotation angle, and rotating the first building image by the optimal rotation angle.
[0162] Step 513, using the right-angled polygon fitting method to perform right-angled polygon fitting on the first building image rotated by the optimal rotation angle.
[0163] Step 514: Perform reverse rotation of the optimal rotation angle on the first building image after right-angled polygon fitting, so that the first building image after right-angled polygon fitting is restored to the angle before rotation.
[0164] Buildings in remote sensing images usually have a certain angle with the horizontal direction of pixels. Therefore, please refer to Figure 8 , in the embodiment of the present invention, by rotating an arbitrary building image in the input building extraction result by an angle in the range of 0 to 90°, the fitting error between the polygon fitting result at the corresponding angle and the original building image without rotation is obtained, and the angle with the smallest fitting error is retained as the optimal rotation angle. Then, perform the right-angled polygon fitting method under the optimal rotation angle again to obtain an optimized building graphic. Not only are the corners of this graphic all right angles, but also the boundary corresponds to the corresponding building target contour. Perform reverse rotation of the optimal rotation angle on the first building image after right-angled polygon fitting, so that the first building image after right-angled polygon fitting is restored to the angle before rotation.
[0165] The implementation details of the method for performing right-angled polygon fitting on the building extraction result are as follows: Define the rotation angle α of the original building contour, α ∈ [0, 90], and rotate the building contour; p i (i = 0, 1,,, n) is the vector node on the building contour line, n is the number of nodes, and △ is the input fitting weight. Determining the fitting error between the polygon fitting result of the first building image at each angle and the first building image without rotation includes: Based on the number of fitting line segments of the first building image, the fitting weight, and the sum of the distances between all the original line segments and the fitting line segments of the first building image, determine the fitting error between the polygon fitting result of the first building image at each angle and the first building image without rotation. In one embodiment, based on the number of fitting line segments of the first building image, the fitting weight, and the sum of the distances between all the original line segments and the fitting line segments of the original line segments of the first building image, determining the fitting error between the polygon fitting result of the first building image at each angle and the first building image without rotation includes:
[0166]
[0167] where m is the number of fitting line segments of the first building image, It represents the sum of the distances between the line segments on the j-th segment of the original contour line of the first building image and the fitted line segments after fitting; △ is the fitting weight, where the fitting weight is a user-defined item. The larger the parameter, the simpler the fitted figure. Conversely, the smaller the parameter, the more sides the fitted figure has. F is the fitting error, and {p} represents the vector nodes on the original contour line of the first building image. Therefore, when the original contour line of the building image is divided into the optimal m line segments, F at this time has the minimum value, and the dynamic programming algorithm can be used to quickly solve the value of m and its corresponding F error value.
[0168] In addition, since the building contours are all composed of right angles, indicating that there is a perpendicular relationship between each adjacent line segment, the above formula can be extended to:
[0169]
[0170] where represents the residual when line segment 1 is a horizontal line segment, represents the residual when line segment 2 is a vertical line segment, and the last line segment is always perpendicular to the first line segment; the dynamic programming algorithm can continue to be used to quickly solve the value of m and the division of its line segments when the F error value is the smallest.
[0171] Given that each rotation angle does not affect each other, the embodiments of the present invention can calculate F α in parallel for each rotation angle α ∈ [0, 90], obtain the optimal rotation angle when the error term is the smallest, and then know the division of the input building contour. Use the least squares method to perform linear fitting on the nodes on the contour line segments. Since adjacent line segments have a perpendicular relationship, calculate the perpendicular intersection points of each adjacent line segment and rotate back to the original angle according to the optimal rotation angle to obtain the figure after fitting of the right-angled polygon. Thus, the embodiments of the present invention realize the right-angled polygon fitting for the building extraction result.
[0172] For example, please refer to Figures 9 to 15 , for Figure 9 the polygon to be fitted, perform a rotation at a certain angle (only taking the optimal angle as an example) to obtain Figure 10 the rotated polygon. Through dynamic programming, the contour edge division with m = 6 is obtained as shown in Figure 11 . Then perform least squares fitting on each contour segment. For example, Figure 12 the method for fitting the contour edge is: the observed data are the points on the contour, denoted as (x 1 , y 1 ), …, (x n , y n), where n is the number of points on the contour. Since the fitting target is a right-angled polygon, the target line is horizontal or vertical, and the illustrated contour edge is horizontal. The line fitting function is f(x) = ax + b, where for a horizontal line, a = 0. The goal of least squares optimization is to minimize the following error L:
[0173]
[0174] Fitting all the contour points and solving for the coefficient b when the minimum error is obtained gives the fitted line. The fitted line is as shown in Figure 13 shown. Fit the lines of other line segments in sequence. Finally, the closed figure formed by all the lines is the fitted right-angled polygon as shown in Figure 14 shown. Finally, rotate the fitted right-angled polygon back by the optimal rotation angle to obtain the right-angled polygon corresponding to the original polygon as shown in Figure 15 shown.
[0175] In other aspects of the embodiments of the present invention, although better extraction results can be obtained through the right-angled polygon fitting operation, there are limitations in some cases. For example, when the tree crown shadow blocks, there will be errors such as missing house corners and concave houses in the extracted results. In these cases, repairing with conventional editing tools involves moving and capturing multiple points, which is very time-consuming. Using the automatic corner filling technology for building contour right-angled polygons can very conveniently repair the house graphics.
[0176] As a semi-automatic house repair technology, the house corner filling technology requires manual visual recognition of the missing right angles or concave pits and uses customized repair tools for box selection to guide the algorithm to process the building target. Please refer to Figure 16 . The overall process of automatically filling corners of right-angled polygons for the building extraction results in the embodiments of the present invention is as follows: First, input the processed part of the building extraction results and perform a missing corner type judgment. Specifically, it includes the missing right angle type, smooth concave pit type, and invalid type. For the missing right angle type, search for the corners on both sides of the missing right angle, calculate the result displacement points based on the corners on both sides of the missing right angle, and update the house graphics based on the result displacement points. For the smooth concave pit type, identify the right-angled sides on both sides of the smooth concave pit, and update the house graphics by smoothing the concave pit and its joints based on the right-angled sides on both sides of the smooth concave pit. For the invalid type, no processing is performed.
[0177] The key points for automatically filling corners of right-angled polygons for the building extraction results in the embodiments of the present invention are as follows:
[0178] (1) There are certain constraints on the missing right angle. For the target concave angle, it must be a concave right angle; the corners on both sides of the concave right angle must be convex right angles. The judgment of convex and concave angles can be combined with the clockwise and counterclockwise attributes of the house point sequence and the directionality of the included angle for judgment.
[0179] (2) For the right-angled sides on both sides of the concave right angle, when identifying nodes, it is necessary to determine whether the angle at the node is a flat angle. If it is a flat angle, it is regarded as a part of the right-angled side and extended. If it is not a flat angle, it is necessary to determine whether it is a right-angled inflection point.
[0180] (3) When identifying a sunken flat pit and encountering a node, it is necessary to determine whether the angle at the node is a flat angle. If it is a flat angle, it is regarded as a part of the right-angled side and extended. If it is not a flat angle, it is necessary to determine whether it is a right-angled inflection point.
[0181] (4) The right-angled corners on both sides of the pit must be concave right angles. The judgment of concave and convex angles can be combined with the positive and negative rotation attributes of the house point sequence and the directivity of the included angle for judgment.
[0182] It should be noted that please refer to Figure 17 , a flat angle refers to Figure 17 the 180-degree angle outlined by the circular area in
[0183] The effects of the automatic supplementary angle of the right-angled polygon on the building extraction result in the embodiment of the present invention are as follows Figure 18 and Figure 19 shown as follows:
[0184] It should be noted that the missing right-angle supplementary angle method and the smooth sunken pit filling method can complement and cooperate with each other. In the actual operation and production process, the two algorithms will be used in series multiple times, and finally an ideal building shape will be formed.
[0185] In other aspects of the embodiment of the present invention, for a single building individual, the acquisition requirements have been met. However, from the overall perspective, there may be gaps and overlaps between the building individual and the surrounding individuals, which need to be regularized through the adjacent building gap processing method to solve the conflicts between buildings.
[0186] Please refer to Figure 20 , the flow chart of the adjacent building gap processing method is as follows: Obtain the building extraction result. When there are other buildings around the building, remove the flat-angle points, search for the adsorbable edges, and perform graphic transformation to complete the adsorption. When there are no other buildings around the building, end the method.
[0187] Since the regularization process is calculated based on the edge, the flat-angle points may split the originally belonging visual edge into two ends, resulting in an incorrect judgment of the adsorption type. Therefore, the first step is to remove the flat-angle points. A flat-angle point refers to the vertex of a 180-degree angle. Define the edge as the line segment connected by adjacent nodes in the building graph, and let s i (i = 0, 1,,, n) be the edges of the original building, where n is the number of edges of the original building, and t j(j = 0, 1,,, m) is the edge of the surrounding building, where m is the number of edges of the surrounding building, △ is the input adsorbable distance, and the specific details of the adjacent building gap treatment method are as follows:
[0188] The adjacent building gap treatment method determines whether it can be adsorbed through an edge matching algorithm:
[0189]
[0190] Among them, mLength(s i , t j ) is the calculated matching length, mArea(s i , t i ) is the calculated matching area. The calculation methods of the matching length and the matching area vary depending on the matching type. Please refer to Figures 21 to 23 , which is divided into the following three types:
[0191] Figure 21 The calculation methods of the matching length and the matching area of
[0192]
[0193] Figure 22 are as follows:
[0194]
[0195] Figure 23 The calculation methods of the matching length and the matching area of
[0196]
[0197] When Match(s i , t j , △) is true, it means a successful match, and t j is the adsorbable edge of s i . When the number of adsorbable edges of s i is 1, the changes in Figure 24 , Figure 25 and Figure 26 will occur; when the number of adsorbable edges of s i is multiple, only the head and the tail are offset as a whole, and only the matching part in the middle is offset. The specific multi-edge adsorption matching effect is as shown in Figure 27 . According to the above rules for graphic changes, when the edges of all buildings have changed, the adsorption is completed.
[0198] In summary, please refer to Figure 28 , Figure 28 represents the deep learning automatic extraction result predicted by the building extraction model of the embodiment of the present invention.Figure 29 It represents the building extraction result of the embodiment of the present invention. It can be seen that the embodiment of the present invention effectively combines methods such as deep learning automatic building extraction, graph cut semi-automatic extraction, and graphic regularization, realizing building extraction with high automation and high edge accuracy. This method can effectively support the requirements of automatic building extraction production in actual production applications.
[0199] Device Embodiment
[0200] Please refer to Figure 30 On the other hand, the embodiment of the present invention also provides a device for automatic building extraction, including:
[0201] An acquisition module 3001, configured to acquire a target remote sensing image;
[0202] A prediction module 3002, configured to input the target remote sensing image into a building extraction model to obtain a mask image and a frame field image of the target remote sensing image output by the building extraction model;
[0203] A construction module 3003, configured to construct a preliminary building extraction result of the target remote sensing image based on the mask image and the frame field image by using an active skeleton model;
[0204] An extraction module 3004, configured to perform graphic optimization on the preliminary building extraction result by using a graph cut method to obtain a building extraction result;
[0205] Wherein, the building extraction model is trained based on a sample remote sensing image, a mask annotation of the sample remote sensing image, and a frame field annotation.
[0206] Optionally, the building extraction model is trained through the following steps:
[0207] Repeat the following steps until a set stop condition is reached:
[0208] Input the sample remote sensing image, the mask annotation of the sample remote sensing image, and the frame field annotation into the building extraction model to obtain a sample mask prediction image and a sample frame field prediction image output by the building extraction model;
[0209] Calculate a mask loss function of the building extraction model based on the mask annotation and the sample mask prediction image, and calculate a frame field loss function of the building extraction model based on the frame field annotation and the sample frame field prediction image;
[0210] Calculate a total loss function of the building extraction model based on the mask loss function and the frame field loss function;
[0211] Adjust model parameters of the building extraction model based on the total loss function.
[0212] Optionally, calculating the mask loss function of the building extraction model based on the mask annotation and the sample mask prediction image includes:
[0213] Calculating the cross-entropy loss function and the Dice loss function of the building extraction model based on the mask annotation and the sample mask prediction image;
[0214] Calculating the mask loss function of the building extraction model based on the cross-entropy loss function and the Dice loss function.
[0215] Optionally, calculating the frame field loss function of the building extraction model based on the frame field annotation and the sample frame field prediction image includes:
[0216] Calculating the horizontal angle loss of the unsigned tangent vector of the polygon contour, the vertical angle loss of the unsigned tangent vector of the polygon contour, and the smooth loss of the building extraction model based on the frame field annotation and the sample frame field prediction image;
[0217] Calculating the frame field loss function of the building extraction model based on the horizontal angle loss of the unsigned tangent vector of the polygon contour, the vertical angle loss of the unsigned tangent vector of the polygon contour, and the smooth loss.
[0218] Optionally, using the graph cut method to optimize the graph of the preliminary building extraction result to obtain the building extraction result includes:
[0219] Performing an erosion operation and a dilation operation on the preliminary building extraction result to obtain an image including multiple building patches;
[0220] Generating horizontal lines, vertical lines, acute diagonals, and obtuse diagonals as prior seed lines within each of the building patches, and generating background lines outside each of the building patches;
[0221] Performing graph cut on the four prior seed lines within each of the building patches to obtain the segmentation results of the four prior seed lines;
[0222] Averaging and fusing the segmentation results of the four prior seed lines within each of the building patches to obtain the building extraction result.
[0223] Optionally, performing graph cut on the four prior seed lines within each of the building patches to obtain the segmentation results of the four prior seed lines includes:
[0224] Determine the pixels of the first prior seed line within the first building patch as the point set of the foreground target, and randomly select the same number of pixels as the pixels of the first prior seed line from the background line corresponding to the first prior seed line as the point set of the background target;
[0225] Construct a mixture Gaussian model based on the point set of the foreground target and the point set of the background target respectively;
[0226] Construct an energy function based on the mixture Gaussian model of the foreground target and the mixture Gaussian model of the background target;
[0227] Use the minimum cut maximum flow algorithm to minimize the energy function to obtain the segmentation result of the first prior seed line;
[0228] Wherein, the first building patch is any one of multiple building patches, and the first prior seed line is any one of four prior seed lines.
[0229] Optionally, after using the graph cut method to perform graphic optimization on the preliminary building extraction result to obtain the building extraction result, it further includes:
[0230] Perform at least one regularization process of right-angled polygon fitting, automatic complementary angle of right-angled polygon, and processing of gaps between adjacent buildings on the building extraction result.
[0231] Optionally, the performing right-angled polygon fitting on the building extraction result includes:
[0232] Rotate the first building image of the building extraction result at multiple angles within a set angle range, and determine the fitting error between the polygon fitting result of the first building image at each angle and the fitting result of the first building image without rotation; wherein, the first building image is any one of the building images in the building extraction result;
[0233] Determine the rotation angle corresponding to the minimum fitting error as the optimal rotation angle, and rotate the first building image by the optimal rotation angle;
[0234] Use the right-angled polygon fitting method to perform right-angled polygon fitting on the first building image rotated by the optimal rotation angle;
[0235] Perform reverse rotation of the optimal rotation angle on the first building image after right-angled polygon fitting to restore the first building image after right-angled polygon fitting to the angle without rotation.
[0236] Optionally, determining the fitting error between the polygon fitting result of the first building image at each of the angles and the first building image without rotation includes:
[0237] Based on the number of fitting line segments of the first building image, the fitting weights, and the sum of the distances between all the original line segments of the first building image and the fitting line segments, determining the fitting error between the polygon fitting result of the first building image at each of the angles and the first building image without rotation.
[0238] Optionally, the determining the fitting error between the polygon fitting result of the first building image at each of the angles and the first building image without rotation based on the number of fitting line segments of the first building image, the fitting weights, and the sum of the distances between all the original line segments of the first building image and the fitting line segments of the original line segments includes:
[0239]
[0240] where m is the number of fitting line segments of the first building image, represents the sum of the distances between the line segment on the j-th segment of the original contour line of the first building image and its fitting line segment after fitting; △ is the fitting weight, F is the fitting error, and {p} represents the vector nodes on the original contour line of the first building image.
[0241] The building automatic extraction device includes a processor and a memory. The above-mentioned acquisition module 3001, prediction module 3002, construction module 3003, extraction module 3004, etc. are all stored in the memory as program units, and the corresponding functions are implemented by the processor executing the above program units stored in the memory.
[0242] The processor contains a kernel, and the corresponding program unit is retrieved from the memory by the kernel. One or more kernels can be set.
[0243] The memory may include non-permanent memory in a computer-readable medium, forms such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.
[0244] Figure 31 Illustrates a schematic diagram of the physical structure of an electronic device, such as Figure 31As shown in the figure, the electronic device may include: a processor 3110, a communications interface 3120, a memory 3130, and a communication bus 3140. Among them, the processor 3110, the communications interface 3120, and the memory 3130 complete mutual communication through the communication bus 3140. The processor 3110 may call logical instructions in the memory 3130 to execute the building automatic extraction method, which includes: obtaining a target remote sensing image; inputting the target remote sensing image into a building extraction model to obtain a mask image and a frame field image of the target remote sensing image output by the building extraction model; using an active skeleton model to construct a preliminary building extraction result of the target remote sensing image based on the mask image and the frame field image; using a graph cut method to optimize the graph of the preliminary building extraction result to obtain a building extraction result; wherein, the building extraction model is trained based on a sample remote sensing image, a mask annotation of the sample remote sensing image, and a frame field annotation.
[0245] In addition, when the logical instructions in the above-mentioned memory 3130 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0246] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a machine-readable storage medium. When the computer program is executed by a processor, the computer can execute the building automatic extraction method, which includes: obtaining a target remote sensing image; inputting the target remote sensing image into a building extraction model to obtain a mask image and a frame field image of the target remote sensing image output by the building extraction model; using an active skeleton model to construct a preliminary building extraction result of the target remote sensing image based on the mask image and the frame field image; using a graph cut method to perform graphic optimization on the preliminary building extraction result to obtain a building extraction result; wherein, the building extraction model is trained based on a sample remote sensing image, a mask annotation of the sample remote sensing image, and a frame field annotation.
[0247] In another aspect, the present invention also provides a machine-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the execution of the building automatic extraction method, which includes: obtaining a target remote sensing image; inputting the target remote sensing image into a building extraction model to obtain a mask image and a frame field image of the target remote sensing image output by the building extraction model; using an active skeleton model to construct a preliminary building extraction result of the target remote sensing image based on the mask image and the frame field image; using a graph cut method to perform graphic optimization on the preliminary building extraction result to obtain a building extraction result; wherein, the building extraction model is trained based on a sample remote sensing image, a mask annotation of the sample remote sensing image, and a frame field annotation.
[0248] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0249] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0250] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for automatically extracting buildings, characterized in that: include: Acquire target remote sensing images; Inputting the target remote sensing image into a building extraction model to obtain a mask image and a frame field image of the target remote sensing image output by the building extraction model; Using an active skeleton model to construct a preliminary building extraction result of a target remote sensing image based on the mask image and the frame field image; Optimizing the preliminary building extraction result by using a graph cut method to obtain a building extraction result; The building extraction model is obtained based on sample remote sensing images, mask annotations of the sample remote sensing images, and frame field annotation training.
2. The automatic building extraction method according to claim 1, characterized in that: The building extraction model is trained by the following steps: Repeat the following steps until the set stop condition is reached: Inputting the sample remote sensing image, the mask annotation and the frame field annotation of the sample remote sensing image into a building extraction model to obtain a sample mask prediction image and a sample frame field prediction image output by the building extraction model; Calculating a mask loss function of the building extraction model based on the mask annotation and the sample mask prediction image, and calculating a frame field loss function of the building extraction model based on the frame field annotation and the sample frame field prediction image; Calculating a total loss function of the building extraction model based on the mask loss function and the frame field loss function; Model parameters of the building extraction model are adjusted based on the total loss function.
3. The automatic building extraction method according to claim 2, characterized in that: The calculating the mask loss function of the building extraction model based on the mask annotation and the sample mask prediction image includes: Calculating a cross entropy loss function and a Dice loss function of the building extraction model based on the mask annotation and the sample mask prediction image; A mask loss function of the building extraction model is calculated based on the cross entropy loss function and the Dice loss function.
4. The automatic building extraction method according to claim 2, characterized in that: The calculating the frame field loss function of the building extraction model based on the frame field annotation and the sample frame field prediction image comprises: Calculate the horizontal angle loss of the unsigned tangent vector of the polygonal contour of the building extraction model, the vertical angle loss of the unsigned tangent vector of the polygonal contour and the smoothing loss based on the frame field annotation and the sample frame field prediction image; A frame field loss function of the building extraction model is calculated based on the horizontal angle loss of the unsigned tangent vector of the polygon outline, the vertical angle loss of the unsigned tangent vector of the polygon outline, and the smoothness loss.
5. The automatic building extraction method according to claim 1, characterized in that: The method of using a graph cut method to perform graphical optimization on the preliminary building extraction result to obtain a building extraction result includes: Performing corrosion operations and dilation operations on the preliminary building extraction results to obtain an image including a plurality of building patches; Generating horizontal lines, vertical lines, acute-angle diagonal lines and obtuse-angle diagonal lines as prior seed lines in each of the building patches, and generating background lines outside each of the building patches; Performing graph cutting on the four prior seed lines in each of the building patches to obtain segmentation results of the four prior seed lines; The segmentation results of the four prior seed lines are respectively averaged and fused in each of the building patches to obtain the building extraction result.
6. The automatic building extraction method according to claim 5, characterized in that: The step of performing graph cutting on the four prior seed lines in each of the building patches to obtain segmentation results of the four prior seed lines includes: Determine the pixels of a first a priori seed line within the first building patch as a point set of a foreground target, and randomly select pixels of the same number as the pixels of the first a priori seed line from background lines corresponding to the first a priori seed line as a point set of a background target; Constructing a mixed Gaussian model based on the point set of the foreground target and the point set of the background target respectively; Constructing an energy function based on the mixed Gaussian model of the foreground target and the mixed Gaussian model of the background target; Minimizing the energy function using the minimum cut maximum flow algorithm to obtain a segmentation result of the first priori seed line; The first building patch is any one of a plurality of building patches, and the first a priori seed line is any one of four a priori seed lines.
7. The automatic building extraction method according to claim 1, characterized in that: After the graph optimization of the preliminary building extraction result by using the graph cut method to obtain the building extraction result, the method further includes: At least one regularization process of right-angle polygon fitting, right-angle polygon automatic angle filling, and adjacent building gap processing is performed on the building extraction result.
8. The automatic building extraction method according to claim 7, characterized in that: The performing rectangular polygon fitting on the building extraction result comprises: Rotating the first building image of the building extraction result at multiple angles within a set angle range, and determining a fitting error between a polygon fitting result of the first building image and the unrotated first building image at each angle; wherein the first building image is any one building image in the building extraction result; Determine a rotation angle corresponding to the minimum fitting error as an optimal rotation angle, and perform a rotation of the first building image at the optimal rotation angle; Performing rectangular polygon fitting on the first building image rotated by the optimal rotation angle using a rectangular polygon fitting method; The first building image after the right-angle polygon fitting is subjected to the reverse rotation of the optimal rotation angle, so that the first building image after the right-angle polygon fitting is restored to the angle before the rotation is performed.
9. The automatic building extraction method according to claim 8, characterized in that: The determining of the fitting error between the polygon fitting result of the first building image at each angle and the first building image that has not been rotated comprises: Based on the number of fitted line segments of the first building image, the fitting weight and the sum of the distances between all original line segments of the first building image and the fitted line segments, the fitting error between the polygon fitting result of the first building image at each angle and the unrotated first building image is determined.
10. The automatic building extraction method according to claim 9, characterized in that: The determining, based on the number of fitted line segments of the first building image, the fitting weight, and the sum of the distances between all original line segments of the first building image and the fitted line segments of the original line segments, a fitting error between a polygon fitting result of the first building image at each angle and the unrotated first building image comprises: Wherein, m is the number of fitted line segments of the first building image, represents the sum of the distances between the jth line segment on the original contour line of the first building image and the fitted line segment after fitting; △ is the fitting weight, F is the fitting error, and {p} represents the vector node on the original contour line of the first building image.
11. An automatic extraction device for a building, characterized in that: include: An acquisition module is used to acquire a target remote sensing image; A prediction module, used for inputting the target remote sensing image into a building extraction model to obtain a mask image and a frame field image of the target remote sensing image output by the building extraction model; A construction module, used for constructing a preliminary building extraction result of a target remote sensing image based on the mask image and the frame field image using an active skeleton model; An extraction module, used for performing graphical optimization on the preliminary building extraction result by using a graph cut method to obtain a building extraction result; The building extraction model is obtained based on sample remote sensing images, mask annotations of the sample remote sensing images, and frame field annotation training.
12. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method for automatically extracting buildings according to any one of claims 1 to 10 is implemented.
13. A machine-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for automatically extracting buildings according to any one of claims 1 to 10 is implemented.
14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for automatically extracting buildings according to any one of claims 1 to 10 is implemented.