Building extraction method from remote sensing images based on adaptive coding and two-stage training
By adopting adaptive coding and two-stage training methods in remote sensing image building extraction, combined with deep learning and interactive segmentation, the problem of building extraction in cross-region images is solved, the extraction accuracy and efficiency are improved, and sample dependence is reduced.
Patent Information
- Application Number
- CN202310620502.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-30
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-05-30
AI Technical Summary
The existing deep learning semantic segmentation network model is difficult to effectively extract buildings in different regions and images, especially when sufficient sample information is lacking in cross-regional images, the model is prone to overfitting and generalization is reduced.
The remote sensing image building extraction method based on adaptive coding and two-stage training is adopted. Interactive features are generated through an adaptive radius encoder, and interactive segmentation is performed in combination with a deep learning model. It is divided into two-stage training: the first stage is to adjust the model parameters through the final extraction result and the loss of the label, and the second stage is to adjust the model parameters through the weighting and adjustment of the IoU loss and iterative loss, reducing the model dependence on the number of clicks.
It improves the accuracy and efficiency of building extraction, reduces the cost of manual interaction, and can achieve building extraction with few samples in cross-regional images, improving the generalization ability of the model.
Smart Images

Figure CN117079121B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the field of remote sensing technology, in particular to a remote sensing image building extraction method based on adaptive coding and two-stage training. Background Art
[0002] The main content of the high-resolution remote sensing image building extraction task is to determine whether the remote sensing image contains building objects, and determine the category of each pixel as a building or non-building, so as to provide effective context information for tasks such as change detection. At present, high-resolution remote sensing image building extraction has been widely used in major fields such as urban expansion research, digital city construction, and post-disaster reconstruction. Unlike ordinary natural image semantic segmentation tasks, high-resolution remote sensing images not only have a significantly reduced proportion of mixed pixels, but also record richer features of ground object shapes, geometric structures, textures, and spatial relationships. A high-resolution remote sensing image scene may contain a variety of ground object types. For example, a dense residential area scene may contain some ground object targets such as houses, roads, and trees. There are problems such as small inter-class differences, large intra-class differences, and buildings being blocked by trees. At the same time, there are also differences in regions (topography), time phases, sensor characteristics, and color styles between buildings in different countries and regions, which leads to the existence of cross-regional image gaps in the existing automatic extraction of buildings. Although this gap can be bridged by annotating buildings in high-resolution images of each time phase and region and building a sample library, manual building annotation often requires a lot of manpower and material resources. High-resolution remote sensing images are updated quickly, and it is impossible to manually annotate buildings on high-resolution images of each region in each period. Mainstream deep learning semantic segmentation network models rely on a large amount of annotation information. Insufficient sample information will lead to overfitting of the model, resulting in a decrease in model generalization. Therefore, mainstream deep learning semantic segmentation network models are difficult to solve the problem of extracting buildings from different images in different regions. Summary of the invention
[0003] In view of this, the purpose of the present invention is to provide a remote sensing image building extraction method based on adaptive coding and two-stage training, which can use human reasoning knowledge to guide the training of network models to achieve better building extraction effects.
[0004] To achieve the above object, the present invention adopts the following technical solution: a remote sensing image building extraction method based on adaptive coding and two-stage training, comprising the following steps:
[0005] Step S1: Design a model including two parts: adaptive radius encoder ARE and building extraction model; first, send the labeled building samples into the adaptive radius encoder, and generate the building / non-building click disk encoding map coords_feature of the maximum adaptive building click radius, and design two training stages;
[0006] Step S2: Set the prediction iteration threshold N and counter n, and generate a blank matrix with the same size as the input original image, fuse it with the original input image and the click disk code map generated in step S1, and input it into the building extraction model for prediction. The counter n is incremented by 1, and the predicted map is sent to ARE to generate new interactive features and update the predicted map.
[0007] Step S3: This step involves two different training stages. In the first stage, step S2 is repeated until n reaches the iteration threshold N, and the prediction map is fused with the interaction feature. In the second stage, step S2 is repeated until the IoU between the prediction map and the label map is greater than 0.9 or reaches the prediction iteration threshold N. Each time step S2 is repeated, the iteration loss function is updated, and finally the prediction map is fused with the interaction feature.
[0008] Step S4: The feature map fused in step S3 and the original image are used as training data for the feature extraction network. The obtained building extraction results are calculated based on the normalized focal loss of the label. The model parameters are adjusted by the loss value in the first stage of training. In the second stage of training, the model parameters are further adjusted by summing the loss with the IoU loss value predicted in each iteration.
[0009] In a preferred embodiment: through an improved click coding scheme, deep learning is combined with interactive segmentation and used for building extraction from high-resolution remote sensing images. Through a "human in the loop" strategy, the cost of labeling is reduced, the quality of labeled samples is improved, and ultimately the extraction of buildings from a small number of samples under cross-regional images is completed.
[0010] In a preferred embodiment: the specific content of the adaptive coding described in step S1 is as follows: the encoder takes the coordinates of the interaction sequence in the two-dimensional space as the center of the circle, and according to the category represented by the pixel point at the coordinate, calculates the minimum distance between the point and other categories, that is, the maximum inscribed circle radius, as the binary disk encoding radius of each interaction feature.
[0011] In a preferred embodiment: The specific contents of the segmented training described in step S3 are as follows: In the first stage, the model parameters are adjusted by using the loss of the final extraction result and the label of the model to promote the model to learn and process the input features. In the second stage, the IoU loss value of each iteration in step S2 is calculated, and then the weighted sum of the iterative loss value is used as the loss value to adjust the model parameters, so as to effectively reduce the number of sample labeling times, improve the sample labeling quality, and enhance the model performance.
[0012] Compared with the prior art, the present invention has the following beneficial effects:
[0013] (1) The method proposed in the present invention can train a model to update the extraction results of buildings in a high-resolution remote sensing image by the user clicking on the image, so as to achieve satisfactory results for the user.
[0014] (2) The present invention aims to solve the problem that fixed-size disk encoding in mainstream interactive semantic segmentation cannot accurately extract buildings in high-resolution remote sensing images due to the various shapes of buildings, mixed distribution, many impurities in target buildings and small target buildings. The adaptive disk encoding is designed to improve the building extraction performance while reducing the cost of manual interaction.
[0015] (3) The present invention improves the training strategy. The model is trained in two stages. The first stage of training is used to promote the model to learn interactive features. In the second stage, the IoU loss function is added during the iteration process to simulate the user's satisfaction with the extraction results to determine whether to continue to increase the click correction results. Secondly, the iterative loss function is added according to the number of iterations to constrain the number of clicks, thereby reducing the model's dependence on the number of clicks and improving the model efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a schematic diagram of the principle of a preferred embodiment of the present invention. DETAILED DESCRIPTION
[0017] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0018] It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.
[0019] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or their combinations.
[0020] like Figure 1 As shown, this embodiment provides a method for interactive building extraction from high-resolution remote sensing images based on adaptive coding and two-stage training, comprising the following steps:
[0021] Step S1: The designed model consists of two parts: Adaptive Radius Encoding (ARE) and building extraction model. First, the labeled building samples are sent to the adaptive radius encoder, and the building / non-building click disk encoding map (coords_feature) with the maximum adaptive building click radius is generated. This method designs two training stages;
[0022] Step S2: Set the prediction iteration threshold N and counter n, and generate a blank matrix with the same size as the input original image, fuse it with the original input image and the click disk code map generated in step S1, and input it into the building extraction model for prediction. The counter n is incremented by 1, and the predicted map is sent to ARE to generate new interactive features and update the predicted map.
[0023] Step S3: This step involves two different training stages. In the first stage, step S2 is repeated until n reaches the iteration threshold N, and the prediction map is fused with the interaction feature. In the second stage, step S2 is repeated until the IoU between the prediction map and the label map is greater than 0.9 or reaches the prediction iteration threshold N. Each time step S2 is repeated, the iteration loss function is updated, and finally the prediction map is fused with the interaction feature.
[0024] Step S4: The feature map fused in step S3 and the original image are used as training data for the feature extraction network. The obtained building extraction results are calculated based on the normalized focal loss of the label. The model parameters are adjusted by the loss value in the first stage of training. In the second stage of training, the model parameters are further adjusted by summing the loss with the IoU loss value predicted in each iteration.
[0025] In this embodiment, an interactive building extraction method for high-resolution remote sensing images based on adaptive coding and two-stage training is adopted. The method combines deep learning with interactive segmentation and is used for building extraction from high-resolution remote sensing images. The building extraction work on cross-regional images is completed through a "human-in-the-loop" computing method.
[0026] In this embodiment, the specific content of the adaptive disk encoding described in step S1 is as follows: This method combines deep learning with interactive segmentation through an improved click encoding scheme, and is used for building extraction from high-resolution remote sensing images. By simulating the "human in the loop" calculation method, it completes the extraction of buildings with a small number of samples under cross-regional images.
[0027] In this embodiment, the specific contents of the segmented training described in step S3 are as follows: In the first stage, the IoU loss extracted at the end of the model is used to adjust the model parameters to promote the model to learn to process the input features. In the second stage, the weighted sum of the iterative IoU loss and the iterative loss according to step 2 is used to adjust the model parameters, which can effectively reduce the model's dependence on increasing clicks and improve the efficiency of the model.
[0028] Preferably, this embodiment combines deep learning with interactive segmentation and uses it for building extraction from high-resolution remote sensing images, completing building extraction on cross-regional images through a "human-in-the-loop" computing method.
[0029] Preferably, the present embodiment includes an adaptive disk code generator, which adaptively generates disk codes of appropriate sizes according to the selected pixels, so as to handle the phenomenon of irregular shapes of buildings in high-resolution remote sensing images.
[0030] Preferably, this embodiment includes a phased training strategy. In the first phase, the model parameters are adjusted by using the loss of the model's final extraction result and the label to promote the model to learn to process the input features. In the second phase, the model parameters are adjusted by using the weighted sum of the iterative IoU loss and the iterative loss according to step 2, which can effectively reduce the model's dependence on increasing clicks and improve the efficiency of the model.
[0031] An improved interactive building extraction strategy for high-resolution remote sensing images with adaptive disk coding. This strategy is divided into two stages of training. In the first stage, a binary disk coding map based on the radius of the maximum inscribed circle is generated by a randomly sampled adaptive disk coding generator to provide more correct semantic information to the network. Then, an iterative threshold N and a counter n are set, and a blank image of the same size as the input original image is generated. After connecting it with the original image and the disk coding map generated by S1, it is input into the building extraction model for prediction. Assign n the value of n + 1, send the prediction map into ARE to generate new interactive features, and update the blank matrix to the prediction map. Repeat the above steps until n reaches the iterative threshold N, and then fuse the prediction map with the interactive features. Then, the fused features are fed into the model for training, and the loss is calculated for the output result of the model to adjust the model parameters. In the second stage, the repeated steps in the first stage are adjusted. The main content is to judge the prediction map. If the IoU between the prediction map and the label is less than 0.9 or n < N, then the IoU is added as a loss to the iterative loss and combined with the loss of the model training result. The specific expression is loss = α(n * iteration_loss) + βNFL to adjust the model parameters.
[0032] Preferably, in this embodiment, by adopting the strategy of adaptive disk coding for the network model, a disk coding of a suitable size can be adaptively generated according to the selected pixel points to handle the phenomenon of abnormal shapes in buildings in high-resolution remote sensing images. The IoU loss function is added during the iteration process to simulate the user's satisfaction with the extraction result to judge whether to continue to increase the click to correct the result. The iterative loss is increased for the number of iterations to constrain the number of clicks, reducing the model's dependence on the number of clicks and thus improving the model efficiency.
[0033] Specifically, it includes the following steps:
[0034] Step S1: The designed model includes two parts: an adaptive radius encoder (ARE) and a building extraction model. First, the labeled building samples are fed into the adaptive radius encoder, and a building / non-building click disk coding map (coords_feature) with the maximum adaptive building click radius is generated. This method designs two training stages;
[0035] Step S2: Set the prediction iteration threshold N and the counter n, and generate a blank matrix of the same size as the input original image. After fusing it with the original input image and the click disk coding map generated in Step S1, it is input into the building extraction model for prediction. The counter n is incremented by 1, the prediction map is fed into ARE to generate new interactive features, and the prediction map is updated;
[0036] Step S3: This step involves two different training stages. In the first stage, step S2 is repeated until n reaches the iteration threshold N, and the prediction map is fused with the interaction feature. In the second stage, step S2 is repeated until the IoU between the prediction map and the label map is greater than 0.9 or reaches the prediction iteration threshold N. Each time step S2 is repeated, the iteration loss function is updated, and finally the prediction map is fused with the interaction feature.
[0037] Step S4: The feature map fused in step S3 and the original image are used as training data for the feature extraction network. The obtained building extraction results are calculated based on the normalized focal loss of the label. The model parameters are adjusted by the loss value in the first stage of training. In the second stage of training, the model parameters are further adjusted by summing the loss with the IoU loss value predicted in each iteration.
[0038] The present invention has the following beneficial effects: the method proposed by the present invention can adapt to the characteristics of buildings of different shapes and the presence of debris in buildings in high-resolution remote sensing images to generate disk codes of different sizes to complete interactive building extraction. In the iterative process, IoU is added to simulate the user's satisfaction with the extraction result to determine whether to continue to increase the click correction result, and iteration_loss is added according to the number of iterations to constrain the number of clicks, reducing the model's dependence on the number of clicks and thus improving the model efficiency.
[0039] In summary, for the problem that directly using existing network models in high-resolution remote sensing image building extraction cannot handle the differences in buildings in different regions, and the problem that the fixed-size disk encoding of existing interactive segmentation strategies is not applicable to high-resolution image buildings, the improved high-resolution remote sensing image interactive building extraction method with adaptive disk encoding proposed in this embodiment is used for high-resolution remote sensing image interactive building extraction. This method is divided into two-stage training. In the first stage, a binary disk encoding map based on the radius of the maximum inscribed circle is generated by a randomly sampled adaptive disk encoding generator to provide more correct semantic information to the network. Then, an iterative threshold N and a counter n are set, and a blank matrix with the same size as the input original image is generated. After connecting it with the original image and the disk encoding map generated by S1, it is input into the building extraction model for prediction. Assign n the value of n+1, send the prediction map into ARE to generate new interactive features, and update the blank matrix to the prediction map. Repeat the above steps until n reaches the iterative threshold N, and then fuse the prediction map with the interactive features. Then, the fused features are fed into the model for training, and the loss is calculated for the output result of the model to adjust the model parameters. In the second stage, the repeated steps in the first stage are adjusted. The main content is to judge the prediction map. If the Iou between the prediction map and the label is less than 0.9 or n<N, then the Iou is added as a loss to Iteration_loss and combined with the loss of the model training result. The specific expression is loss=α(n * iteration_loss) + βNFL to adjust the model parameters. Compared with existing methods, the method of this invention can effectively improve the building extraction accuracy across regions, and at the same time make the positive and negative click samples adaptively encode the size of the building / non-building regions to obtain sufficient and correct semantic information to be fed into the network for training. The experimental data uses the commonly used high-resolution remote sensing image building extraction dataset WHU as the training set, and the click encoding is simulated by machine processing the samples of the dataset to simulate human reasoning knowledge as an additional input source for the training set. The strategy of this embodiment can guide the network to learn how to extract buildings under the guidance of reasoning knowledge. This embodiment can effectively improve the generalization ability of the model, improve the building extraction accuracy across regions, and at the same time reduce the number of clicks.
[0040] The above are only the preferred embodiments of the present invention, and all equivalent changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope of the present invention.
Claims
1. A remote sensing image building extraction method based on adaptive coding and two-stage training, characterized by: The following steps are involved: Step S1: Design a model including two parts: adaptive radius encoder ARE and building extraction model; first, send the labeled building samples into the adaptive radius encoder, and generate the building / non-building click disk encoding map coords_feature of the maximum adaptive building click radius, and design two training stages; Step S2: Set the prediction iteration threshold N and counter n, and generate a blank matrix with the same size as the input original image, fuse it with the original input image and the click disk code map generated in step S1, and input it into the building extraction model for prediction. The counter n is incremented by 1, and the predicted map is sent to ARE to generate new interactive features and update the predicted map. Step S3: This step involves two different training stages. In the first stage, step S2 is repeated until n reaches the iteration threshold N, and the prediction map is fused with the interaction feature. In the second stage, step S2 is repeated until the IoU between the prediction map and the label map is greater than 0.9 or reaches the prediction iteration threshold N. Each time step S2 is repeated, the iteration loss function is updated, and finally the prediction map is fused with the interaction feature. Step S4: The feature map fused in step S3 and the original image are used as training data for the feature extraction network. The obtained building extraction results are calculated based on the normalized focal loss of the label. The model parameters are adjusted by the loss value in the first stage of training. In the second stage of training, the model parameters are further adjusted by summing the loss with the IoU loss value predicted in each iteration. Combining deep learning with interactive segmentation through an improved click coding scheme and using it for building extraction from high-resolution remote sensing images through a "human-in-the-loop" strategy; The specific contents of the adaptive radius encoding in step S1 are as follows: the encoder takes the coordinates of the interaction sequence in the two-dimensional space as the center of the circle, and according to the category represented by the pixel point where the coordinates are located, calculates the minimum distance between the point and other categories, that is, the maximum inscribed circle radius, as the binary disk encoding radius of each interaction feature; The specific contents of the segmented training in step S3 are as follows: in the first stage, the loss of the model's final extraction result and the label is used to adjust the model parameters. In the second stage, the IoU loss value of each iteration in step S2 is calculated, and then the weighted sum of the iterative loss value is used as the loss value to adjust the model parameters.
Citation Information
Patent Citations
Three-stage liver tumor image segmentation method based on adaptive preprocessing
CN115018864A
Boundary-optimized remote sensing image semantic segmentation method and apparatus, and device and medium
WO2023077816A1