A semi-supervised building instance extraction method based on high-resolution remote sensing images

By using semi-supervised learning and data augmentation technology in the building instance extraction method, the instance segmentation network is trained using and without label data, and through pseudo-label optimization, the problem of label acquisition difficulties and insufficient edge fineness in building instance extraction is solved, and higher extraction accuracy and edge fineness are achieved.

CN115861802BActive Publication Date: 2025-06-27CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211492219.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-25
Publication Date
2025-06-27
Estimated Expiration
2042-11-25

AI Technical Summary

Technical Problem

In the prior art, the building instance extraction method has problems such as difficulty in obtaining building pixel-level labels, high manual labeling costs, limited labeled data, weak generalization ability of the extraction method, and insufficient fine edges of the extracted building.

Method used

A semi-supervised building instance extraction method based on high-resolution remote sensing images is proposed. By obtaining image data with pixel-level labels and without pixel-level labels, data augmentation and instance segmentation network are trained, and the pseudo-label optimization model is used to improve the accuracy and edge fineness of building instance extraction.

Benefits of technology

It solves the problems of building label acquisition difficulties and insufficient edges in traditional methods, and improves the generalization ability of the model and the accuracy of building instance extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115861802B_ABST
    Figure CN115861802B_ABST
Patent Text Reader

Abstract

The present invention discloses a semi-supervised building instance extraction method based on high-resolution remote sensing images. The method comprises the following steps: cropping the high-resolution remote sensing images to obtain labeled data and unlabeled data; after augmenting the labeled data, training a first instance segmentation network; predicting pseudo-labels of building instances for the unlabeled data through the first instance segmentation network; jointly training a second instance segmentation network by combining the pseudo-labels with the remote sensing images to achieve semi-supervised automatic extraction of buildings based on high-resolution remote sensing images, wherein a multi-scale feature method and an edge operator extraction algorithm are used to refine the extraction of the edges of buildings. The present invention can be used for the extraction of building instances in high-resolution remote sensing images in scenarios with limited data labels, improving the generalization of the model while enhancing the extraction accuracy of buildings and refining the edges of buildings.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image target extraction, and particularly to a semi-supervised building instance extraction method based on high-resolution remote sensing images. Background Art

[0002] Building information, as an indispensable part of urban basic geographic information, plays an important supporting role in the operation, management, and planning of cities. In recent years, the continuous development of modern satellite remote sensing technology has provided sufficient data sources for building extraction. At the same time, the progress of artificial intelligence methods in related fields of computer vision has provided effective methods for automatic building extraction.

[0003] The purpose of building instance extraction is to detect all buildings and accurately segment each building. The building instance information extracted from high-resolution remote sensing images is of great significance in the research of urban planning, environmental management, urban change, and map updating. Traditional high-resolution remote sensing image building extraction methods are mainly based on the image analysis of geographical objects. Such methods are based on professional domain knowledge, and the extraction performance depends on the selection of manual features. With the implementation of deep learning methods in the remote sensing field, the building instance extraction method based on deep learning for high-resolution remote sensing images has achieved significant performance advantages. It assigns pixel-level instance labels to each building in the remote sensing image through the instance segmentation method to achieve automatic extraction of building instances.

[0004] For the instance segmentation method, it also follows the supervised machine learning paradigm and requires a large amount of data to drive model learning. The instance segmentation model needs to be trained using a large number of samples with pixel-level instance labels. However, it is difficult to obtain pixel-level instance labels of buildings in high-resolution remote sensing images, and the cost of manual annotation is relatively high. Therefore, it is often only possible to obtain partial labeled data. Therefore, it is particularly important to obtain a building extraction method with strong generalization ability and better extraction effect under the premise of limited pixel-level building instance labels.

[0005] Semi-supervised learning provides an effective approach for the building instance extraction task with limited pixel-level instance labels. Compared with traditional fully supervised learning methods, the semi-supervised learning building instance segmentation model can obtain richer building feature representations from unlabeled high-resolution remote sensing images while learning the building features in the remote sensing images with pixel-level instance labels, thereby obtaining better segmentation performance.

[0006] In addition to improving the segmentation performance, it is also particularly important to extract the fine degree of the contour of building instances from high-resolution remote sensing images. However, due to factors such as the scale of buildings and the diverse backgrounds of remote sensing images, the ordinary instance segmentation models have problems such as blurred contours and insufficiently fine edges in the extracted buildings.

[0007] In summary, the main challenges in building instance extraction based on high-resolution remote sensing images are as follows: it is difficult to obtain pixel-level instance labels, the labeled data is limited, and the edges of the extracted buildings are not fine enough. Summary of the Invention

[0008] The main purpose of the present invention is to solve the technical problems in the prior art, such as the difficulty in obtaining building pixel-level labels for building instance extraction methods, the high cost of manual annotation, the limited labeled data, the weak generalization ability of the extraction method, and the insufficiently fine edges of the extracted buildings, and to improve the accuracy of building instance extraction while enhancing the generality of the model.

[0009] A semi-supervised building instance extraction method based on high-resolution remote sensing images proposed by the present invention includes the following steps:

[0010] S1: Obtain the high-resolution remote sensing image I() and perform cropping. According to the building pixel-level annotation l, obtain the high-resolution remote sensing image data I with pixel-level building instance labels that meet the actual demand size l and the high-resolution remote sensing image data I without pixel-level building instance labels of the corresponding size u ;

[0011] S2: Perform mirror-derived data augmentation on the high-resolution remote sensing image I with pixel-level building instance labels l to obtain the augmented high-resolution remote sensing image data with pixel-level instance labels

[0012] S3: Combine the high-resolution remote sensing image data I with pixel-level building instance labels l and the augmented high-resolution remote sensing image data with pixel-level building instance labels and input them into the first instance segmentation network to be trained for training to obtain the trained first instance segmentation network;

[0013] S4: Input the high-resolution remote sensing image data I without pixel-level building instance labels u into the trained first instance segmentation network for extraction to obtain the instance pseudo-labels of the buildings in the high-resolution remote sensing image data I u ;

[0014] S5: Combine the pseudo - labels of the building instances and the high - resolution remote - sensing image data I with corresponding pixel - free building instance labels u , with the high - resolution remote - sensing image data I with pixel - level building instance labels l , and use them as the input of the second instance - segmentation network for training to obtain the trained second instance - segmentation network.

[0015] S6: Input the high - resolution remote - sensing image to be predicted into the trained second instance - segmentation network to extract building instances.

[0016] The beneficial effects provided by the present invention are as follows: It solves the technical problems in traditional building instance extraction methods, such as difficult acquisition of building labels, high cost of manual annotation, weak generalization ability of extraction methods, and insufficient fineness of the edges of the extracted buildings, and improves the accuracy of building instance extraction while enhancing the generality of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a simple flow diagram of the method of the present invention;

[0018] Figure 2 is a training flow chart of the first instance - segmentation network in the present invention;

[0019] Figure 3 is a structural diagram of the optimized Mask branch of the second instance - segmentation network in the present invention;

[0020] Figure 4 is a comparison diagram of edge optimization results. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described below in conjunction with the accompanying drawings.

[0022] Please refer to Figure 1 , Figure 1 which is a simple flow diagram of the method of the present invention;

[0023] A semi - supervised building instance extraction method based on high - resolution remote - sensing images includes the following steps:

[0024] S1: Obtain the high - resolution remote - sensing image I() and crop it. According to the pixel - level annotation l of the building, obtain the high - resolution remote - sensing image data I with pixel - level building instance labels that meet the actual demand size l and the high - resolution remote - sensing image data I with pixel - free building instance labels of the corresponding size u ;

[0025] In step S1, the high - resolution remote - sensing image I() is specifically represented as follows:

[0026]

[0027] Among them, I l represents the high-resolution remote sensing image data with pixel-level building instance labels, u represents the high-resolution remote sensing image data without pixel-level building instance labels; l represents the annotation boxes of all buildings in the high-resolution remote sensing image, k represents the number of annotation boxes; p refers to the cropped high-resolution remote sensing image, and c represents the number of cropped images.

[0028] S2: Mirror-derived data augmentation is performed on the high-resolution remote sensing image I l with pixel-level building instance labels to obtain the augmented high-resolution remote sensing image data with pixel-level instance labels

[0029] The labeled high-resolution remote sensing image data after enhancement in step S2 Specifically:

[0030]

[0031] Among them, T i represents the enhancement of the i-th data, n represents the number of operations, θ n represents the weight of the n-th operation, η represents the parameter of the operation; g represents the specific enhancement operation, including any one or a combination of flipping, symmetry, and mirror operations.

[0032] S3: The high-resolution remote sensing image data I l with pixel-level building instance labels and the augmented high-resolution remote sensing image data with pixel-level building instance labels are combined and input into the first instance segmentation network to be trained to obtain the trained first instance segmentation network;

[0033] Please refer to Figure 2 , Figure 2 which is the training flowchart of the first instance segmentation network in the present invention; Step S3 is specifically:

[0034] S31: The ResNet network is used as the backbone network of the first instance segmentation network, and the high-resolution remote sensing image data I l with pixel-level building instance labels and the augmented high-resolution remote sensing image data with pixel-level instance labels are input into the backbone network of the first instance segmentation network to obtain multiple feature maps of different scales;

[0035] S32: Input the extracted feature maps of different scales into the Feature Pyramid Network (FPN) module for feature fusion to obtain multiple fused feature maps;

[0036] S33: Input the multiple fused feature maps into the Region Proposal Network (RPN) module to perform foreground and background classification and bounding box regression, obtain candidate regions, and calculate the RPN module loss L ′ , as follows:

[0037] L ′ = L ′ cls + L ′ box (3)

[0038] where L ′ cls represents the classification loss of foreground and background in the RPN module, and L ′ box represents the bounding box regression loss;

[0039] S34: Input the candidate regions into the RoIAlign module for unified sampling to obtain candidate feature maps containing buildings;

[0040] S35: Send the candidate feature maps into the Box branch and the Mask branch respectively. The Box branch obtains the classification and label box of the extracted building, and the Mask branch obtains the corresponding building instance mask, and calculate the losses of the two branches L, specifically:

[0041] L = L cls + L box + L mask (4)

[0042] where L cls represents the classification loss of the extracted building instance, L box represents the regression loss of the label box, and L mask represents the loss of the building instance mask.

[0043] S36: Calculate the total loss L all of the instance segmentation model, perform gradient descent and backpropagation, so as to train the first instance segmentation network to obtain a trained first instance segmentation network.

[0044] The total loss L all of the first instance segmentation network is specifically:

[0045] L all = L + L′ (5)

[0046] Preferably, in step S33, L′ clsThe calculation method is specifically as follows:

[0047]

[0048] Among them, N cls represents the number of categories, represents the true category label, and p i represents the predicted label.

[0049] In step S33, L′ box The calculation method is specifically as follows:

[0050]

[0051] Among them, represents the true bounding box label, and t i represents the predicted bounding box label.

[0052] Preferably, in step S35, L cls and L box have the same calculation method as L′ cls and L′ box in S33. The calculation method of L mask is as follows:

[0053] L mask = -y log y * - (1 - y) log(1 - y * ) (8)

[0054] Among them, y * represents the true building mask label, and y represents the predicted building mask label.

[0055] S4: Input the high-resolution remote sensing image data I u without pixel-level building instance labels into the trained first instance segmentation network for extraction to obtain the instance pseudo-labels of the buildings in the high-resolution remote sensing image data I u without pixel-level building instance labels.

[0056] Step S4 is specifically as follows:

[0057] S41: Input the high-resolution remote sensing image data I u without pixel-level building instance labels into the trained first instance segmentation network for extraction to obtain the predicted bounding box label and confidence score, which are specifically represented as:

[0058]

[0059] Among them, b m is the predicted bounding box position, s m represents the confidence score, and zm denotes the predicted instance mask, and M denotes the unlabeled high-resolution remote sensing image data I u quantity;

[0060] S42: By setting a threshold filter the obtained confidence scores to obtain the labels higher than the threshold as the pseudo-labels of the high-resolution remote sensing image data I without pixel-level building instance labels u pseudo-labels.

[0061] The said threshold The calculation formula is:

[0062]

[0063] S5: Combine the building instance pseudo-labels and the corresponding high-resolution remote sensing image data I without pixel-level building instance labels u , with the high-resolution remote sensing image data I with pixel-level building instance labels l as the input of the second instance segmentation network and train it to obtain the trained second instance segmentation network.

[0064] It should be noted that the training process of the second instance segmentation network in step S5 can refer to step S3.

[0065] Specifically, please refer to Figure 3 , Figure 3 which is the structural diagram of the optimized Mask branch of the second instance segmentation network in the present invention; step S5 is specifically:

[0066] S51: Input the high-resolution remote sensing image data I without pixel-level building instance labels u , and the high-resolution remote sensing image data I with pixel-level building instance labels l into the second instance segmentation network, and the second instance segmentation network has the same structure as the first instance segmentation network; after passing through the feature extractor, FPN module, RPN module and RoIAlign module, obtain the candidate feature map containing the building;

[0067] S52: Send the candidate feature map containing the building into the Box branch to obtain the classification and label box of the extracted building, and calculate the corresponding losses L cls and L box ;

[0068] S53: Send the candidate feature map containing the building into the Mask branch, extract two feature maps f1 and f2 through two sub-branches of different scales, and fuse them to obtain the feature map f m , and the specific dimension is expressed as:

[0069] ;

[0070] S54: Extract the edge features of the building in the horizontal and vertical directions from the fused feature map f m , and fuse them to obtain a feature map f s that refines the building edge, and calculate the corresponding loss L mask ;

[0071] The formula for f s is as follows:

[0072]

[0073] where C1D represents a 1×1 convolution operation, ξ x represents the Sobel operation in the horizontal direction, ξ y represents the Sobel operation in the vertical direction, represents the feature fusion operation.

[0074] S55: Calculate the total loss L all , perform gradient descent and backpropagation, thereby training the instance segmentation model to obtain a trained second instance segmentation network.

[0075] S6: Input the high-resolution remote sensing image to be predicted into the trained second instance segmentation network to extract building instances.

[0076] As an embodiment, this application uses the WHU dataset for experiments. The training set of the WHU dataset is divided into labeled data according to the ratios of 5%, 10%, 20%, 30%, and 40%, and the remaining data is used as unlabeled data. The above method is used for training and tested on the test set of the WHU dataset. The final average mAP accuracy results are shown in Table 1. The supervised row in the table represents the accuracy of training only using labeled data in step S3. The Semi-super row represents the final result accuracy of S5, that is, training using both data with true labels and unlabeled data with pseudo labels.

[0077] By comparing the results in the second row and the third row, it can be seen that the accuracy in the third row has been greatly improved, indicating that the proposed method uses both labeled data and unlabeled data for learning, improving the accuracy of building instance extraction, and completing the task of high-resolution remote sensing image building instance extraction in the scenario of limited data labels.

[0078] Table 1 Experimental result data table

[0079]

[0080] Figure 4 is a comparison chart of edge optimization results; Figure 4 The figure shows the visual comparison of the results before and after the optimization of the Mask branch in S5. Each column uses the same picture for result comparison. The first row shows the result without optimization, with circles marking the areas with poor edge segmentation. The second row shows the result after the Mask branch is optimized by the method described in S5, with circles marking the refined building edges. By comparing the visualization results before and after optimization, it can be seen that the optimized Mask branch method described in S5 can extract building instances more finely.

[0081] In summary, the experimental results show that the method can complete the task of extracting building instances from high-resolution remote sensing images in scenarios with limited data labels, improve the accuracy of building extraction, and refine the extraction of building edges.

[0082] The beneficial effects of the present invention are: solving technical problems in traditional building instance extraction methods, such as difficulty in obtaining building labels, high cost of manual labeling, weak generalization ability of extraction methods, and insufficient fine edges of extracted buildings, thereby improving the accuracy of building instance extraction while improving the versatility of the model.

[0083] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A semi-supervised building instance extraction method based on high-resolution remote sensing images, characterized in that: Including the following steps: S1: Obtain high-resolution remote sensing images and perform cropping according to the pixel-level annotation of buildings , to obtain high-resolution remote sensing image data with pixel-level building instance labels in the size that meets the actual requirements and high-resolution remote sensing image data without pixel-level building instance labels in the corresponding size ; S2: The high-resolution remote sensing image with pixel-level building instance labels is subjected to mirror-derived data augmentation to obtain the high-resolution remote sensing image data with augmented pixel-level instance labels ; S3: Combine the high-resolution remote sensing image data with pixel-level building instance labels and the high-resolution remote sensing image data with pixel-level building instance labels after data augmentation and input them into the first instance segmentation network to be trained for training to obtain a trained first instance segmentation network; S4: Input the high-resolution remote sensing image data without pixel-level building instance labels into the trained first instance segmentation network for extraction to obtain the instance pseudo-labels of the buildings in the high-resolution remote sensing image data without pixel-level building instance labels ; S5: Use the high-resolution remote sensing image data of the building instance pseudo-labels and the corresponding pixel-level building instance labels without , combine it with the high-resolution remote sensing image data with pixel-level building instance labels , use it as the input of the second instance segmentation network and train it to obtain a trained second instance segmentation network; Specifically, step S5 is as follows: S51: The high-resolution remote sensing image data of the building instance label without pixel level , and the high-resolution remote sensing image data of the building instance label with pixel level are input into a second instance segmentation network, which has the same structure as the first instance segmentation network; after passing through a feature extractor, an FPN module, an RPN module, and an RoIAlign module, a candidate feature map containing buildings is obtained; S52: Feed the candidate feature map containing the building into the Box branch to obtain the classification and label boxes for extracting the building, and calculate the corresponding loss and ; S53: Extract two feature maps from the candidate feature map Mask branch containing the building through two sub-branches of different scales and , and fuse them to obtain a feature map ; S54: Take the fused feature map , extract the edge features of the building in the horizontal and vertical directions through the Sobel edge operator, and fuse them to obtain a feature map that refines the building edges , and calculate the corresponding loss ; S55: Calculate the total loss , perform gradient descent and backpropagation to train the instance segmentation model, and obtain a trained second instance segmentation network; S6: Input the high-resolution remote sensing image to be predicted into the trained second instance segmentation network to extract building instances.

2. The semi-supervised building instance extraction method based on high-resolution remote sensing images according to claim 1, wherein: The high-resolution remote sensing image described in step S1 , and the specific representation form is as follows: (1) Among them, represents high-resolution remote sensing image data with pixel-level building instance labels, represents high-resolution remote sensing image data without pixel-level building instance labels; represents the bounding boxes of all buildings in the high-resolution remote sensing image, represents the number of bounding boxes; refers to the cropped high-resolution remote sensing image, represents the number of cropped images.

3. A semi-supervised building instance extraction method based on high-resolution remote sensing images according to claim 1, characterized in that: The enhanced labeled high-resolution remote sensing image data in step S2 , specifically: (2) Among them indicates enhancing the th piece of data, represents the number of operations, represents the weight of the th operation, represents the parameters of the operation; represents the specific enhancement operation, including any one or a combination of flipping, symmetry, and mirroring operations.

4. The semi-supervised building instance extraction method based on high-resolution remote sensing images according to claim 1, wherein: Specifically, step S3 is as follows: S31: Use the ResNet network as the backbone network of the first instance segmentation network, and input the high-resolution remote sensing image data with pixel-level building instance labels and the augmented high-resolution remote sensing image data with pixel-level instance labels into the backbone network of the first instance segmentation network to obtain feature maps of multiple different scales; S32: Input the extracted feature maps of different scales into the Feature Pyramid Network (FPN) module for feature fusion to obtain multiple fused feature maps; S33: Input the multiple fused feature maps into the Region Proposal Network (RPN) module to perform foreground and background classification and bounding box regression, obtain candidate regions, and calculate the RPN module loss , as follows: (3) Among them, represents the classification loss of the foreground and background of the RPN module, represents the bounding box regression loss; S34: Input the candidate regions into the RoIAlign module for unified sampling to obtain candidate feature maps containing buildings; S35: Feed the candidate feature maps into the Box branch and the Mask branch respectively. The Box branch obtains the classification and label boxes for extracting buildings, the Mask branch obtains the corresponding building instance masks, and calculate the losses of the two branches L , specifically: (4) Among them, represents the classification loss of the extracted building instances, represents the regression loss of the label boxes, represents the loss of the building instance mask; S36: Calculate the total loss of the instance segmentation model , perform gradient descent and backpropagation to train the first instance segmentation network, and obtain the trained first instance segmentation network.

5. A semi-supervised building instance extraction method based on high-resolution remote sensing images according to claim 1, characterized in that: Specifically, step S4 is as follows: S41: Input the high-resolution remote sensing image data of the building instance label without pixel level into the trained first instance segmentation network for extraction to obtain the predicted box label and confidence score, specifically expressed as: (5) Among them, Predicted bounding box position, Indicates the confidence score, Indicates the predicted instance mask, Indicates unlabeled high-resolution remote sensing image data The number of; S42: By setting a threshold , filter the obtained confidence scores to obtain the labels higher than the threshold as the pseudo-labels of the high-resolution remote sensing image data of the pixel-level building instance label-free .