Semi-supervised building instance extraction method based on dynamic alignment and saliency constraint
By constructing a semi-supervised extraction network with cross-stage consistency constraints, and utilizing dynamic alignment and saliency constraints, the problems of background interference and pseudo-label noise in semi-supervised building instance extraction methods are solved, thereby improving the performance and accuracy of building instance extraction.
Patent Information
- Application Number
- CN202510140511.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-02-08
AI Technical Summary
Existing semi-supervised building instance extraction methods face challenges in distinguishing background interference, pseudo-label noise, and the complexity of multi-stage training, leading to a decline in extraction performance.
A semi-supervised method based on dynamic alignment and saliency constraints is adopted. By constructing a semi-supervised extraction network with cross-stage consistency constraints (CLC-SIE), the quality of pseudo-labels and the performance of building instance extraction are improved by utilizing the dynamic alignment and pixel saliency constraints of the teacher model and student model.
It improves the accuracy and efficiency of building instance extraction, reduces the complexity of multi-stage training and the difficulty of parameter tuning, enhances the network's ability to distinguish background interference, and mitigates the impact of pseudo-label noise.
Smart Images

Figure CN120088645B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of remote sensing image processing technology, and in particular to a semi-supervised method for extracting building instances based on dynamic alignment and saliency constraints. Background Technology
[0002] Building instance extraction, a fundamental task in remote sensing image processing, serves a wide range of fields, including urban construction and planning, environmental monitoring and management, disaster prevention and mitigation, and map updating. With the rapid development of satellite remote sensing technology, a large amount of high-resolution remote sensing data is constantly being generated, providing valuable resources for building extraction. However, complex designs and buildings of varying sizes, as well as natural elements such as trees and shadows, introduce additional complexity. Therefore, achieving accurate and automatic extraction of building instances from high-resolution remote sensing imagery remains a significant challenge.
[0003] Traditional building extraction methods manually select more discriminative semantic features, such as edges, spectrum, color, or texture. These features rely on extensive expert experience and handcrafted rules, exhibiting limited robustness in complex and varied environments. In recent years, with the rapid development of deep learning technology, deep learning-based methods have achieved remarkable results in building instance extraction tasks. These methods automatically learn and interpret abstract high-dimensional features from low-dimensional features, extracting building instances in a detection-then-segmentation manner. However, these methods are significantly limited by their reliance on pixel-level annotations, as obtaining such complex annotations is typically time-consuming and labor-intensive, impractical for rapid applications.
[0004] In recent years, semi-supervised learning has attracted researchers' attention as a method to alleviate dependence on data annotation, allowing learning from a limited amount of labeled data and a large amount of unlabeled data. In existing semi-supervised building instance extraction work, confidence-based pseudo-labeling methods have shown significant performance advantages. In this paradigm, pseudo-labels are assigned to unlabeled data after being filtered by a preset threshold. The model is then trained on both labeled and pseudo-labeled images. Although these methods can extract buildings, they still face significant challenges. First, remote sensing images often consist of multiple ground features, leading to complex background interference. Directly distinguishing between buildings and interfering targets with similar features based on confidence is challenging. Second, it is worth noting that these confidence-based methods heavily rely on the quality of the pseudo-labels. When a large number of noisy pseudo-labels are present, the model exhibits fitting bias, leading to incorrect building extraction. Finally, multi-stage methods that predict pseudo-labels before co-training require more parameter tuning and manual intervention, a complex process prone to error accumulation, resulting in degraded extraction performance. Summary of the Invention
[0005] This invention provides a semi-supervised building instance extraction method based on dynamic alignment and saliency constraints, which solves the shortcomings of existing semi-supervised building instance extraction methods in distinguishing background interference, the influence of pseudo-label noise, and the complexity of multi-stage training. It improves the performance of building instance extraction while improving the quality of pseudo-labels.
[0006] In a first aspect, the present invention provides a semi-supervised method for extracting building instances based on dynamic alignment and saliency constraints, comprising:
[0007] High-resolution remote sensing images are acquired, cropped and divided, and instance-level annotations are performed on some of the high-resolution remote sensing images to obtain a small amount of data containing building instance labels and a large amount of data without building instance labels.
[0008] The unlabeled instance data is augmented to different degrees to obtain heavily augmented unlabeled data and normally augmented unlabeled data.
[0009] Construct CLC-SIE by inputting the data containing building instance labels, the heavily enhanced unlabeled data, and the ordinary enhanced unlabeled data into CLC-SIE for training, and obtain a building instance extraction model;
[0010] The remote sensing image of the building to be extracted is input into the building instance extraction model, and the building instance prediction result is output.
[0011] According to the present invention, a semi-supervised building instance extraction method based on dynamic alignment and saliency constraints is provided. This method acquires high-resolution remote sensing images, crops and divides the high-resolution remote sensing images, and performs instance-level annotation on buildings in a portion of the high-resolution remote sensing images, resulting in a small amount of data containing building instance labels and a large amount of data without building instance labels, including:
[0012] The high-resolution remote sensing images are obtained using existing satellite remote sensing image data or aerial photography equipment.
[0013] The high-resolution remote sensing images were manually annotated at the instance level to obtain a small amount of building instance label data, which included category labels, bounding box labels, and mask labels for all buildings. The remaining images were used as the building instance label data.
[0014] According to a semi-supervised building instance extraction method based on dynamic alignment and saliency constraints provided by the present invention, the unlabeled building instance data is augmented to different degrees to obtain heavily augmented unlabeled data and normally augmented unlabeled data, including:
[0015] The unlabeled data of the building instances is subjected to size transformation, rotation, sharpness transformation, Gaussian noise and image inversion on the image according to random probability to obtain the heavily enhanced unlabeled data.
[0016] The unlabeled data of the building instances is subjected to size transformation, random rotation, sharpness transformation and brightness transformation of the image according to random probability to obtain the ordinary enhanced unlabeled data.
[0017] According to the present invention, a semi-supervised building instance extraction method based on dynamic alignment and saliency constraints is provided, which constructs CLC-SIE, including:
[0018] The CLC-SIE is defined as including a building instance prediction module, a target dynamic alignment module, and a pixel saliency constraint module;
[0019] The building instance prediction module includes a teacher model and a student model, used to obtain target-level predictions of building categories and bounding boxes in remote sensing images, as well as pixel-level predictions of building masks.
[0020] The target dynamic alignment module is used to calculate the consistency loss of the teacher model and the student model, and output the building category and bounding box.
[0021] The pixel saliency constraint module is used to calculate the consistency loss of the building mask for the teacher model and the student model.
[0022] According to the present invention, a semi-supervised building instance extraction method based on dynamic alignment and saliency constraints is provided, which inputs the labeled data containing building instances, the heavily enhanced unlabeled data, and the normally enhanced unlabeled data into the CLC-SIE for training to obtain a building instance extraction model, including:
[0023] The data containing building instance labels is input into the student model to predict the building categories, bounding boxes, and masks in the remote sensing image, and the supervised loss is calculated by comparing these with the ground truth labels. :
[0024]
[0025] in, This indicates the number of labeled data. , and This represents the category, bounding box, and mask predicted by the student model. , and This represents the category label, bounding box label, and mask label of the real buildings in the labeled data. Represents any tag data;
[0026] The ordinary augmented unlabeled data is input into the teacher model to predict data including building categories. Bounding box and mask pseudo-tags ;
[0027] The heavily augmented unlabeled data is input into the student model to predict scores including building categories. Bounding box and mask Predicted value ;
[0028] Determine the threshold The predicted values of the student model are divided into foreground targets according to their category scores. and background objectives :
[0029]
[0030] Based on background objectives Calculate dynamic weights :
[0031]
[0032] in, This represents the classification score of the j-th background target. Indicates the number of background targets;
[0033] pseudo-tags Building category score and the foreground target in the forecast Background and Objectives and dynamic weights Input the target dynamic alignment module and calculate the unsupervised class loss. :
[0034]
[0035] in, Indicates the number of foreseeable targets. Represents the cross-entropy classification loss;
[0036] Based on the building boundary box in the pseudo-label and the foreground target in the forecast Filter the regression box with the largest IOU. :
[0037]
[0038] Based on the building boundary box in the pseudo-label and regression box Calculate unsupervised regression loss :
[0039]
[0040] in, This represents the bounding box regression loss;
[0041] Based on unlabeled data Extracting significance boundaries :
[0042]
[0043] in, Indicates the Laplace boundary extraction operator;
[0044] Building mask in pseudo-tags Mask in student model predictions and significance boundary Input pixel saliency constraint module to calculate unsupervised mask loss :
[0045]
[0046] in, This indicates the number of building masks in the pseudo-tags. This represents the cross-entropy mask loss;
[0047] Based on the above unsupervised category losses Unsupervised regression loss and unsupervised mask loss Calculate the total unsupervised loss :
[0048]
[0049] in, Indicates the number of unlabeled data;
[0050] According to the monitoring loss and unsupervised total loss Calculate the total loss of the student model, and train the weights of the student model using gradient descent and backpropagation. The total loss was obtained. :
[0051]
[0052] Based on the weights of the student model Update the weights of the teacher model :
[0053]
[0054] in, Indicates the smoothing hyperparameter;
[0055] The building instance extraction model is obtained by iteratively training the student model and the teacher model in an end-to-end manner.
[0056] Secondly, the present invention also provides a semi-supervised building instance extraction system based on dynamic alignment and saliency constraints, comprising:
[0057] The preprocessing module is used to acquire high-resolution remote sensing images, crop and divide the high-resolution remote sensing images, and perform instance-level annotation on some of the high-resolution remote sensing images to obtain a small amount of data containing building instance labels and a large amount of data without building instance labels.
[0058] The enhancement module is used to perform different degrees of data enhancement on the unlabeled instance data to obtain heavily enhanced unlabeled data and normally enhanced unlabeled data;
[0059] The training module is used to construct a semi-supervised extraction network CLC-SIE with cross-stage consistency constraints. The data containing building instance labels, the heavily enhanced unlabeled data, and the ordinary enhanced unlabeled data are input into the CLC-SIE for training to obtain a building instance extraction model.
[0060] The processing module is used to input the remote sensing image of the building to be extracted into the building instance extraction model and output the building instance prediction result.
[0061] According to the present invention, a semi-supervised building instance extraction system based on dynamic alignment and saliency constraints is provided, wherein the preprocessing module is specifically used for:
[0062] The high-resolution remote sensing images are obtained using existing satellite remote sensing image data or aerial photography equipment.
[0063] The high-resolution remote sensing images were manually annotated at the instance level to obtain a small amount of building instance label data, which included category labels, bounding box labels, and mask labels for all buildings. The remaining images were used as the building instance label data.
[0064] Thirdly, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the semi-supervised building instance extraction method based on dynamic alignment and saliency constraints as described above.
[0065] Fourthly, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the semi-supervised building instance extraction method based on dynamic alignment and saliency constraints as described above.
[0066] Fifthly, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the semi-supervised building instance extraction method based on dynamic alignment and saliency constraints as described above.
[0067] The semi-supervised building instance extraction method based on dynamic alignment and saliency constraints provided by this invention constructs an end-to-end semi-supervised building instance extraction method based on the idea of cross-stage consistency. It solves the problems of complex training process, difficult parameter tuning and easy error accumulation of existing multi-stage methods. It also uses the idea of dynamic alignment to increase the network's ability to distinguish between buildings and background interference, uses saliency boundaries to mitigate the influence of noise in pseudo-labels, and improves the quality of pseudo-labels through end-to-end joint training, thereby improving the performance of building instance extraction. Attached Figure Description
[0068] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0069] Figure 1 This is one of the flowcharts of the semi-supervised building instance extraction method based on dynamic alignment and saliency constraints provided by the present invention;
[0070] Figure 2 This is the second flowchart of the semi-supervised building instance extraction method based on dynamic alignment and saliency constraints provided by the present invention;
[0071] Figure 3 This is a diagram of the CLC-SIE semi-supervised extraction network architecture provided by the present invention;
[0072] Figure 4 This is a schematic diagram of target dynamic alignment provided by the present invention;
[0073] Figure 5 This is a schematic diagram of pixel saliency constraints provided by the present invention;
[0074] Figure 6 This is a schematic diagram of the semi-supervised building instance extraction system based on dynamic alignment and saliency constraints provided by the present invention;
[0075] Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0076] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0077] Figure 1 This is one of the flowcharts illustrating the semi-supervised building instance extraction method based on dynamic alignment and saliency constraints provided in this embodiment of the invention, such as... Figure 1 As shown, it includes:
[0078] Step 100: Acquire high-resolution remote sensing images, crop and divide the high-resolution remote sensing images, and perform instance-level annotation on some of the buildings in the high-resolution remote sensing images to obtain a small amount of data containing building instance labels and a large amount of data without building instance labels.
[0079] Step 200: Perform different degrees of data augmentation on the unlabeled instance data to obtain heavily augmented unlabeled data and normally augmented unlabeled data;
[0080] Step 300: Construct CLC-SIE by inputting the data containing building instance labels, the heavily enhanced unlabeled data, and the ordinary enhanced unlabeled data into CLC-SIE for training to obtain a building instance extraction model;
[0081] Step 400: Input the remote sensing image of the building to be extracted into the building instance extraction model, and output the building instance prediction result.
[0082] Specifically, refer to Figure 2 As shown, the embodiments of the present invention mainly include the following steps:
[0083] S1: Acquire a large number of high-resolution remote sensing images, crop and segment them, and perform instance-level annotation on buildings in a small portion of the images to obtain a small amount of data with building instance labels. and a large amount of data without building instance labels ;
[0084] Specifically, S1 includes:
[0085] A large number of high-resolution remote sensing images were obtained using existing satellite remote sensing image data or aerial photography equipment. Some images were labeled at the instance level by manual annotation, resulting in a small amount of labeled data including category labels, bounding box labels, and mask labels for all buildings. The remaining images were treated as unlabeled data.
[0086] S2, the unlabeled data Different data augmentations were performed to obtain heavily augmented unlabeled data and ordinary augmented unlabeled data;
[0087] Specifically, S2 includes:
[0088] Heavy enhancement applies random probability to an image by resizing, rotating, sharpening, adding Gaussian noise, and inverting the image; ordinary enhancement applies random probability to an image by resizing, rotating, sharpening, and brightness.
[0089] S3. Construct a semi-supervised extraction network (CLC-SIE) with cross-stage consistency constraints. Input the labeled data and the enhanced unlabeled data into the CLC-SIE network for training. After training, the building instance extraction model is obtained.
[0090] Among them, such as Figure 3 As shown, S3 specifically includes:
[0091] S31. Construct a semi-supervised extraction network with cross-stage consistency constraints (CLC-SIE), including a building instance prediction module, a target dynamic alignment module, and a pixel saliency constraint module. The building instance prediction module comprises two instance extraction models, a teacher model and a student model, used to obtain target-level predictions of building categories and bounding boxes in remote sensing imagery, as well as pixel-level predictions of building masks. The target dynamic alignment module calculates the consistency loss of the building categories and bounding boxes output by the teacher and student models. The pixel saliency constraint module calculates the consistency loss of the building masks output by the teacher and student models.
[0092] S32, Input labeled data into the student model to predict building categories, bounding boxes, and masks in the image, and calculate the supervised loss with the ground truth labels. The formula is as follows:
[0093]
[0094] in, This indicates the number of labeled data. , and This represents the category, bounding box, and mask predicted by the student model. , and This represents the category label, bounding box label, and mask label of the real buildings in the labeled data.
[0095] S33, Input unlabeled data with standard augmentation into the teacher model to predict data including building categories. Bounding box and mask pseudo-tags Simultaneously, heavily augmented unlabeled data is input into the student model, and the output includes building category scores. Bounding box and mask Predicted value .
[0096] S34, Set threshold The student model predictions are then categorized into foreground targets based on their class scores. and background objectives The formula is as follows:
[0097]
[0098] S35, based on the aforementioned background objective Calculate dynamic weights The formula is as follows:
[0099]
[0100] in, This represents the classification score of the j-th background target. Indicates the number of background targets.
[0101] S36, such as Figure 4 As shown, the building category scores in the above pseudo-labels are... Foreground target in the predicted value and background objectives and dynamic weights Input the target dynamic alignment module and calculate the unsupervised class loss. The formula is as follows:
[0102]
[0103] in, Indicates the number of foreseeable targets. This represents the cross-entropy classification loss.
[0104] S37, based on the building boundary box in the aforementioned pseudo-label and the foreground target in the forecast Filter the regression box with the largest IOU. The formula is as follows:
[0105]
[0106] S38, based on the building boundary box in the above pseudo-label and regression box Calculate unsupervised regression loss The formula is as follows:
[0107]
[0108] in, This represents the bounding box regression loss.
[0109] S39, based on the above unlabeled data Extracting saliency boundaries The formula is as follows:
[0110]
[0111] in, This represents the Laplace boundary extraction operator.
[0112] S310, such as Figure 5 As shown, the building mask in the above pseudo-labels Mask in student model predictions and significance boundary Input pixel saliency constraint module to calculate unsupervised mask loss The formula is as follows:
[0113]
[0114] in, This indicates the number of building masks in the pseudo-tags. This represents the cross-entropy mask loss.
[0115] S311, based on the above unsupervised category loss Unsupervised regression loss and unsupervised mask loss Calculate the total unsupervised loss :
[0116]
[0117] in, This indicates the number of unlabeled data items.
[0118] S312, based on the above-mentioned monitoring losses and unsupervised total loss The total loss of the student model is calculated, and the weights of the student model are trained using gradient descent and backpropagation. The total loss formula is as follows:
[0119]
[0120] S313, based on the weights of the student model described above. Update the weights of the teacher model The formula is as follows:
[0121]
[0122] in, This represents the smoothing hyperparameter.
[0123] S314. The student model and teacher model are iteratively trained in an end-to-end manner to obtain the building instance extraction model.
[0124] S4. Input the remote sensing image of the building to be extracted into the trained building instance extraction model to extract the prediction results of the corresponding building instances in the remote sensing image.
[0125] The semi-supervised building instance extraction system based on dynamic alignment and saliency constraints provided by the present invention will be described below. The semi-supervised building instance extraction system based on dynamic alignment and saliency constraints described below can be referred to in correspondence with the semi-supervised building instance extraction method based on dynamic alignment and saliency constraints described above.
[0126] Figure 6 This is a schematic diagram of the semi-supervised building instance extraction system based on dynamic alignment and saliency constraints provided in an embodiment of the present invention, as shown below. Figure 6 As shown, it includes: a preprocessing module 61, an enhancement module 62, a training module 63, and a processing module 64, wherein:
[0127] The preprocessing module 61 is used to acquire high-resolution remote sensing images, crop and segment the high-resolution remote sensing images, and perform instance-level annotation on some buildings in the high-resolution remote sensing images to obtain a small amount of data containing building instance labels and a large amount of data without building instance labels. The enhancement module 62 is used to perform data enhancement on the data without building instance labels to different degrees to obtain heavily enhanced unlabeled data and normally enhanced unlabeled data. The training module 63 is used to construct a semi-supervised extraction network CLC-SIE with cross-stage consistency constraints, and input the data containing building instance labels, the heavily enhanced unlabeled data, and the normally enhanced unlabeled data into the CLC-SIE for training to obtain a building instance extraction model. The processing module 64 is used to input the remote sensing images of the buildings to be extracted into the building instance extraction model and output the building instance prediction results.
[0128] Figure 7 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 7 As shown, the electronic device may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740, wherein the processor 710, the communication interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 can call logic instructions in the memory 730 to execute a semi-supervised building instance extraction method based on dynamic alignment and saliency constraints. This method includes: acquiring high-resolution remote sensing images; cropping and dividing the high-resolution remote sensing images; and performing instance-level annotation on buildings in a portion of the high-resolution remote sensing images to obtain a small amount of data containing building instance labels and a large amount of data containing no building instance labels; performing different degrees of data augmentation on the data containing no building instance labels to obtain heavily augmented unlabeled data and normally augmented unlabeled data; constructing a CLC-SIE model; inputting the data containing building instance labels, the heavily augmented unlabeled data, and the normally augmented unlabeled data into the CLC-SIE model for training to obtain a building instance extraction model; inputting the remote sensing images of the buildings to be extracted into the building instance extraction model, and outputting building instance prediction results.
[0129] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0130] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the semi-supervised building instance extraction method based on dynamic alignment and saliency constraints provided by the above methods. The method includes: acquiring high-resolution remote sensing images; cropping and dividing the high-resolution remote sensing images; and performing instance-level annotation on buildings in a portion of the high-resolution remote sensing images to obtain a small amount of data containing building instance labels and a large amount of data without building instance labels; performing data augmentation on the data without building instance labels to different degrees to obtain heavily augmented unlabeled data and normally augmented unlabeled data; constructing a CLC-SIE; inputting the data containing building instance labels, the heavily augmented unlabeled data, and the normally augmented unlabeled data into the CLC-SIE for training to obtain a building instance extraction model; inputting the remote sensing image of the building to be extracted into the building instance extraction model and outputting the building instance prediction result.
[0131] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements a semi-supervised building instance extraction method based on dynamic alignment and saliency constraints provided by the above methods. The method includes: acquiring high-resolution remote sensing images; cropping and dividing the high-resolution remote sensing images; and performing instance-level annotation on buildings in a portion of the high-resolution remote sensing images to obtain a small amount of data containing building instance labels and a large amount of data without building instance labels; performing data augmentation on the data without building instance labels to different degrees to obtain heavily augmented unlabeled data and normally augmented unlabeled data; constructing a CLC-SIE; inputting the data containing building instance labels, the heavily augmented unlabeled data, and the normally augmented unlabeled data into the CLC-SIE for training to obtain a building instance extraction model; inputting the remote sensing image of the building to be extracted into the building instance extraction model, and outputting the building instance prediction result.
[0132] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0133] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A semi-supervised method for extracting building instances based on dynamic alignment and saliency constraints, characterized in that, include: High-resolution remote sensing images are acquired, cropped and divided, and instance-level annotations are performed on some of the high-resolution remote sensing images to obtain a small amount of data containing building instance labels and a large amount of data without building instance labels. The unlabeled instance data is augmented to different degrees to obtain heavily augmented unlabeled data and normally augmented unlabeled data. A semi-supervised extraction network CLC-SIE with cross-stage consistency constraints is constructed. The data containing building instance labels, the heavily enhanced unlabeled data, and the ordinary enhanced unlabeled data are input into the CLC-SIE for training to obtain a building instance extraction model. Input the remote sensing image of the building to be extracted into the building instance extraction model, and output the building instance prediction result; Building CLC-SIE includes: The CLC-SIE is defined as including a building instance prediction module, a target dynamic alignment module, and a pixel saliency constraint module; The building instance prediction module includes a teacher model and a student model, used to obtain target-level predictions of building categories and bounding boxes in remote sensing images, as well as pixel-level predictions of building masks. The target dynamic alignment module is used to calculate the consistency loss of the teacher model and the student model, and output the building category and bounding box. The pixel saliency constraint module is used to calculate the consistency loss of the building mask for the teacher model and the student model.
2. The semi-supervised building instance extraction method based on dynamic alignment and saliency constraints according to claim 1, characterized in that, High-resolution remote sensing imagery is acquired, cropped, and segmented. Instance-level annotations are then performed on buildings within a portion of the high-resolution remote sensing imagery, resulting in a small amount of data containing building instance labels and a large amount of data without building instance labels, including: The high-resolution remote sensing images are obtained using existing satellite remote sensing image data or aerial photography equipment. The high-resolution remote sensing images were manually annotated at the instance level to obtain a small amount of building instance label data, which included category labels, bounding box labels, and mask labels for all buildings. The remaining images were used as the building instance label data.
3. The semi-supervised building instance extraction method based on dynamic alignment and saliency constraints according to claim 1, characterized in that, The unlabeled instance data is augmented to different degrees to obtain heavily augmented unlabeled data and moderately augmented unlabeled data, including: The unlabeled data of the building instances is subjected to size transformation, rotation, sharpness transformation, Gaussian noise and image inversion on the image according to random probability to obtain the heavily enhanced unlabeled data. The unlabeled data of the building instances is subjected to size transformation, random rotation, sharpness transformation and brightness transformation of the image according to random probability to obtain the ordinary enhanced unlabeled data.
4. The semi-supervised building instance extraction method based on dynamic alignment and saliency constraints according to claim 1, characterized in that, The labeled data containing building instances, the heavily augmented unlabeled data, and the normally augmented unlabeled data are input into the CLC-SIE for training to obtain a building instance extraction model, including: The data containing building instance labels is input into the student model to predict the building categories, bounding boxes, and masks in the remote sensing image, and the supervised loss is calculated by comparing these with the ground truth labels. : in, This indicates the number of labeled data. , and This represents the category, bounding box, and mask predicted by the student model. , and This represents the category label, bounding box label, and mask label of the real buildings in the labeled data. Represents any tag data; The ordinary augmented unlabeled data is input into the teacher model to predict data including building categories. Bounding box and mask pseudo-tags ; The heavily augmented unlabeled data is input into the student model to predict scores including building categories. Bounding box and mask Predicted value ; Determine the threshold The predicted values of the student model are divided into foreground targets according to their category scores. and background objectives : Based on background objectives Calculate dynamic weights : in, This represents the classification score of the j-th background target. Indicates the number of background targets; pseudo-tags Building category score and the foreground target in the forecast Background and Objectives and dynamic weights Input the target dynamic alignment module and calculate the unsupervised class loss. : in, Indicates the number of foreseeable targets. Represents the cross-entropy classification loss; Based on the building boundary box in the pseudo-label and the foreground target in the forecast Filter the regression box with the largest IOU. : Based on the building boundary box in the pseudo-label and regression box Calculate unsupervised regression loss : in, This represents the bounding box regression loss; Based on unlabeled data Extracting significance boundaries : in, Indicates the Laplace boundary extraction operator; Building mask in pseudo-tags Mask in student model predictions and significance boundary Input pixel saliency constraint module to calculate unsupervised mask loss : in, This indicates the number of building masks in the pseudo-tags. This represents the cross-entropy mask loss; Based on the above unsupervised category losses Unsupervised regression loss and unsupervised mask loss Calculate the total unsupervised loss : in, Indicates the number of unlabeled data; According to the monitoring loss and unsupervised total loss Calculate the total loss of the student model, and train the weights of the student model using gradient descent and backpropagation. The total loss was obtained. : Based on the weights of the student model Update the weights of the teacher model : in, Indicates the smoothing hyperparameter; The building instance extraction model is obtained by iteratively training the student model and the teacher model in an end-to-end manner.
5. A semi-supervised building instance extraction system based on dynamic alignment and saliency constraints, based on the semi-supervised building instance extraction method based on dynamic alignment and saliency constraints as described in any one of claims 1 to 4, characterized in that, include: The preprocessing module is used to acquire high-resolution remote sensing images, crop and divide the high-resolution remote sensing images, and perform instance-level annotation on some of the high-resolution remote sensing images to obtain a small amount of data containing building instance labels and a large amount of data without building instance labels. The enhancement module is used to perform different degrees of data enhancement on the unlabeled instance data to obtain heavily enhanced unlabeled data and normally enhanced unlabeled data; The training module is used to construct a semi-supervised extraction network CLC-SIE with cross-stage consistency constraints. The data containing building instance labels, the heavily enhanced unlabeled data, and the ordinary enhanced unlabeled data are input into the CLC-SIE for training to obtain a building instance extraction model. The processing module is used to input the remote sensing image of the building to be extracted into the building instance extraction model and output the building instance prediction result.
6. The semi-supervised building instance extraction system based on dynamic alignment and saliency constraints according to claim 5, characterized in that, The preprocessing module is specifically used for: The high-resolution remote sensing images are obtained using existing satellite remote sensing image data or aerial photography equipment. The high-resolution remote sensing images were manually annotated at the instance level to obtain a small amount of building instance label data, which included category labels, bounding box labels, and mask labels for all buildings. The remaining images were used as the building instance label data.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the semi-supervised building instance extraction method based on dynamic alignment and saliency constraints as described in any one of claims 1 to 4.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the semi-supervised building instance extraction method based on dynamic alignment and saliency constraints as described in any one of claims 1 to 4.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the semi-supervised building instance extraction method based on dynamic alignment and saliency constraints as described in any one of claims 1 to 4.