Semi-supervised building instance extraction method based on dynamic alignment and significance constraint
By introducing technical means of dynamic alignment and significance constraints in the semi-supervised building instance extraction method, the CLC-SIE network is built, which solves the problems of background interference, pseudo-label noise and multi-stage training complexity, and achieves more efficient and accurate building instance extraction.
Patent Information
- Application Number
- CN202510140511.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-02-08
AI Technical Summary
Existing semi-supervised building instance extraction methods have challenges in distinguishing background interference, handling pseudo-label noise, and multi-stage training complexity, resulting in degradation in extraction performance.
Using a semi-supervised approach based on dynamic alignment and significance constraints, a semi-supervised extraction network (CLC-SIE) with cross-stage consistency constraints is constructed, and end-to-end training is combined with teacher model and student model to dynamically align buildings and backgrounds, and using significance boundaries to improve pseudo-label quality.
It improves the pseudo-label quality and building instance extraction performance, reduces background interference and noise impact, simplifies the training process, and improves the stability and accuracy of the extraction model.
Smart Images

Figure CN120088645A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image processing, and in particular, to a semi-supervised building instance extraction method based on dynamic alignment and saliency constraints. Background Art
[0002] Building instance extraction, as a fundamental task in remote sensing image processing, widely serves fields such as urban construction and planning, environmental monitoring and management, disaster prevention and mitigation, and map updating. With the rapid development of satellite remote sensing technology, a large amount of high-resolution remote sensing data has been continuously generated, providing valuable resources for building extraction. However, the complex design and buildings of different scales, as well as natural elements such as trees and shadows, bring additional complexities. Therefore, achieving accurate and automatic extraction of building instances from high-resolution remote sensing images still faces significant challenges.
[0003] Traditional building extraction methods are to manually select more discriminative semantic features, such as edges, spectra, colors, or textures. These features rely on rich expert experience and handcrafted rules, and show limited robustness in complex and changing environments. In recent years, with the rapid development of deep learning technology, deep learning-based methods have achieved remarkable results in building instance extraction tasks. These methods automatically learn and interpret abstract high-dimensional features from low-dimensional features, and extract building instances in a way of first detecting and then segmenting. However, the dependence of these methods on pixel-level annotations constitutes a significant limitation, because obtaining such complex annotations is usually time-consuming and laborious, and is impractical for rapid applications.
[0004] In recent years, semi-supervised learning, as a method to reduce the dependence on data annotation, has attracted the attention of researchers. It allows learning from a limited number of labeled data and a large amount of unlabeled data. In existing semi-supervised building instance extraction work, the confidence-based pseudo-labeling method has significant performance advantages. In this paradigm, pseudo-labels are assigned to unlabeled data after being filtered by a preset threshold. Then, the labeled images and pseudo-labeled images are used to train the model. Although these methods can extract buildings, they still face huge challenges. First, remote sensing images usually consist of multiple ground objects, resulting in complex background interference. Directly distinguishing buildings with similar features and interference targets according to confidence is somewhat challenging. Second, it is worth noting that these confidence-based methods rely heavily on the quality of pseudo-labels. When there are a large number of noisy pseudo-labels, the model will have a fitting bias, resulting in incorrect building extraction. Finally, the multi-stage method of first predicting pseudo-labels and then co-training requires more parameter tuning and manual intervention. The process is complex and error accumulation is easy, leading to a decline in extraction performance. Summary of the Invention
[0005] The present invention provides a semi-supervised building instance extraction method based on dynamic alignment and saliency constraints, which is used to solve the defects of difficult background interference discrimination, pseudo-label noise influence, and complex multi-stage training in the existing semi-supervised building instance extraction methods, and realizes improving the performance of building instance extraction while enhancing the quality of pseudo-labels.
[0006] In a first aspect, the present invention provides a semi-supervised building instance extraction method based on dynamic alignment and saliency constraints, including: Obtain a high-resolution remote sensing image, crop and divide the high-resolution remote sensing image, and perform instance-level annotation on the buildings in part of the high-resolution remote sensing image to obtain a small amount of building instance label data and a large amount of non-building instance label data; Perform data augmentation on the non-building instance label data to different degrees to obtain severely augmented unlabeled data and normally augmented unlabeled data; Construct a CLC-SIE, input the building instance label data, the severely augmented unlabeled data, and the normally augmented unlabeled data into the CLC-SIE for training to obtain a building instance extraction model; Input the building remote sensing image to be extracted into the building instance extraction model, and output the building instance prediction result.
[0007] According to the semi-supervised building instance extraction method based on dynamic alignment and saliency constraints provided by the present invention, obtaining a high-resolution remote sensing image, cropping and dividing the high-resolution remote sensing image, and performing instance-level annotation on the buildings in part of the high-resolution remote sensing image to obtain a small amount of building instance label data and a large amount of non-building instance label data includes: Obtain the high-resolution remote sensing image through existing satellite remote sensing image data or aerial photography devices; Adopt an artificial annotation method to perform instance-level annotation on the high-resolution remote sensing image to obtain a small amount of building instance label data including class labels, bounding box labels, and mask labels of all buildings, and the remaining images are used as the non-building instance label data.
[0008] According to the semi-supervised building instance extraction method based on dynamic alignment and saliency constraints provided by the present invention, performing data augmentation on the non-building instance label data to different degrees to obtain severely augmented unlabeled data and normally augmented unlabeled data includes: Perform size transformation, rotation, sharpness transformation, Gaussian noise, and image inversion on the non-building instance label data according to a random probability to obtain the severely augmented unlabeled data; Transform the building - free instance label data to perform size transformation, random rotation, sharpness transformation, and brightness transformation on the image according to a random probability to obtain the general enhanced unlabeled data.
[0009] According to a semi - supervised building instance extraction method based on dynamic alignment and saliency constraints provided by the present invention, construct a CLC - SIE, including: Determine that the CLC - SIE includes a building instance prediction module, a target dynamic alignment module, and a pixel saliency constraint module; The building instance prediction module includes a teacher model and a student model, and is used to obtain object - level predictions of building categories and bounding boxes in the remote sensing image, as well as pixel - level predictions of building masks; The target dynamic alignment module is used to calculate the teacher model and the student model, and output the consistency loss of the building category and the bounding box; The pixel saliency constraint module is used to calculate the teacher model and the student model, and output the consistency loss of the building mask.
[0010] According to a semi - supervised building instance extraction method based on dynamic alignment and saliency constraints provided by the present invention, input the building instance label data, the heavily enhanced unlabeled data, and the general enhanced unlabeled data into the CLC - SIE for training to obtain a building instance extraction model, including: Input the building instance label data into the student model, predict the building category, bounding box, and mask in the remote sensing image, and calculate the supervision loss with the ground - truth label :
[0011] Among them, represents the number of labeled data, , and represent the category, bounding box, and mask predicted by the student model, , and represent the category label, bounding box label, and mask label of the real building in the labeled data, represents any label data; Input the general enhanced unlabeled data into the teacher model, and predict the pseudo - label including the building category , bounding box and mask ; ; Input the heavily enhanced unlabeled data into the student model, and predict to obtain the building category score , bounding box And mask The predicted value ; Determine the threshold , and divide the predicted values of the student model into foreground targets according to the class scores And background targets :
[0012] Calculate the dynamic weight according to the background target : :
[0013] Among them, Represents the classification score of the jth background target, Represents the number of background targets; Put the pseudo-label The building class score in And the foreground targets in the predicted value , background targets And the dynamic weight Input into the target dynamic alignment module to calculate the unsupervised class loss :
[0014] Among them, Represents the number of foreground targets, Represents the cross-entropy classification loss; According to the building bounding box in the pseudo-label And the foreground targets in the predicted value , screen the regression box with the largest IOU :
[0015] According to the building bounding box in the pseudo-label And the regression box Calculate the unsupervised regression loss :
[0016] Among them, Represents the bounding box regression loss; Extract the saliency boundary according to the unlabeled data : :
[0017] Among them, Represents the Laplacian boundary extraction operator; The building mask in the pseudo-label , the mask in the predicted value of the student model and the saliency boundary are input into the pixel saliency constraint module to calculate the unsupervised mask loss :
[0018] where represents the number of building masks in the pseudo-label, represents the cross-entropy mask loss; According to the above unsupervised class loss , the unsupervised regression loss and the unsupervised mask loss calculate the unsupervised total loss :
[0019] where represents the number of unlabeled data; According to the supervised loss and the unsupervised total loss calculate the total loss of the student model, and train the weights of the student model through gradient descent and backpropagation to obtain the total loss :
[0020] According to the weights of the student model , update the weights of the teacher model :
[0021] where represents the smoothing hyperparameter; Iteratively train the student model and the teacher model in an end-to-end manner to obtain the building instance extraction model.
[0022] In a second aspect, the present invention also provides a semi-supervised building instance extraction system based on dynamic alignment and saliency constraint, including: A preprocessing module for obtaining a high-resolution remote sensing image, cropping and dividing the high-resolution remote sensing image, and performing instance-level annotation on buildings in part of the high-resolution remote sensing image to obtain a small amount of building instance label data and a large amount of non-building instance label data; An enhancement module for performing data enhancement on the non-building instance label data to different degrees to obtain severely enhanced unlabeled data and normally enhanced unlabeled data; A training module, which is used to construct a semi-supervised extraction network CLC-SIE with cross-stage consistency constraints, input the building instance label data, the heavily augmented unlabeled data, and the normally augmented unlabeled data into the CLC-SIE for training, and obtain a building instance extraction model; A processing module, which is used to input the remotely sensed image of the building to be extracted into the building instance extraction model and output a building instance prediction result.
[0023] According to a semi-supervised building instance extraction system based on dynamic alignment and saliency constraints provided by the present invention, the preprocessing module is specifically used for: Obtaining the high-resolution remotely sensed image through existing satellite remotely sensed image data or aerial photography devices; Performing instance-level annotation on the high-resolution remotely sensed image in an artificial annotation manner to obtain a small amount of building instance label data including class labels, bounding box labels, and mask labels of all buildings, and the remaining images are used as the building instance label-free data.
[0024] In a third aspect, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the semi-supervised building instance extraction method based on dynamic alignment and saliency constraints as described in any one of the above.
[0025] In a fourth aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the semi-supervised building instance extraction method based on dynamic alignment and saliency constraints as described in any one of the above.
[0026] In a fifth aspect, the present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the semi-supervised building instance extraction method based on dynamic alignment and saliency constraints as described in any one of the above.
[0027] The semi-supervised building instance extraction method based on dynamic alignment and saliency constraints provided by the present invention constructs an end-to-end semi-supervised building instance extraction method based on the idea of cross-stage consistency, solves the problems of complex training process, difficult parameter tuning, and easy error accumulation in existing multi-stage methods, and also uses the idea of dynamic alignment to increase the network's ability to distinguish buildings from background interference, uses saliency boundaries to alleviate the influence of noise in pseudo-labels, and improves the quality of pseudo-labels through end-to-end joint training, thereby improving the performance of building instance extraction. Description of the Drawings
[0028] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0029] Figure 1 is one of the schematic flowcharts of the semi-supervised building instance extraction method based on dynamic alignment and saliency constraint provided by the present invention; Figure 2 is the second schematic flowchart of the semi-supervised building instance extraction method based on dynamic alignment and saliency constraint provided by the present invention; Figure 3 is the architecture diagram of the semi-supervised extraction network CLC-SIE provided by the present invention; Figure 4 is the schematic diagram of target dynamic alignment provided by the present invention; Figure 5 is the schematic diagram of pixel saliency constraint provided by the present invention; Figure 6 is the structural schematic diagram of the semi-supervised building instance extraction system based on dynamic alignment and saliency constraint provided by the present invention; Figure 7 is the structural schematic diagram of the electronic device provided by the present invention. Detailed implementation manners
[0030] To make the purpose, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are some embodiments of the present invention, rather than all embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0031] Figure 1 is one of the schematic flowcharts of the semi-supervised building instance extraction method based on dynamic alignment and saliency constraint provided by the embodiments of the present invention. As Figure 1 shown, it includes: Step 100: Obtain a high-resolution remote sensing image, crop and divide the high-resolution remote sensing image, and perform instance-level annotation on the buildings in part of the high-resolution remote sensing image to obtain a small amount of building instance label data and a large amount of unlabeled building instance data; Step 200: Perform data augmentation on the unlabeled building instance data to different degrees to obtain severely augmented unlabeled data and normally augmented unlabeled data; Step 300: Construct a CLC-SIE, input the building instance label data, the heavily augmented unlabeled data, and the normally augmented unlabeled data into the CLC-SIE for training to obtain a building instance extraction model; Step 400: Input the remote sensing image of the building to be extracted into the building instance extraction model, and output the building instance prediction result.
[0032] Specifically, referring to Figure 2 As shown, the embodiments of the present invention mainly include the following steps: S1. Obtain a large number of high-resolution remote sensing images, perform cropping and partitioning, and perform instance-level annotation on the buildings in a small part of the images to obtain a small amount of data with building instance labels and a large amount of data without building instance labels ; Among them, S1 specifically includes: Obtain a large number of high-resolution remote sensing images through existing satellite remote sensing image data or aerial photography devices, and perform instance-level annotation on some images through manual annotation to obtain a small amount of labeled data including class labels, bounding box labels, and mask labels of all buildings, and the remaining images are used as unlabeled data.
[0033] S2. Perform different data augmentations on the unlabeled data to obtain heavily augmented unlabeled data and normally augmented unlabeled data; Among them, S2 specifically includes: Heavy augmentation performs size transformation, rotation, sharpness transformation, Gaussian noise, and image inversion on the image according to a random probability; normal augmentation performs size transformation, random rotation, sharpness transformation, and brightness transformation on the image according to a random probability.
[0034] S3. Construct a cross-stage consistency constraint semi-supervised extraction network (CLC-SIE), input the labeled data and the augmented unlabeled data into the CLC-SIE network for training, and obtain a building instance extraction model after training is completed.
[0035] Among them, as Figure 3 shown, S3 specifically includes: S31. Construct a semi-supervised extraction network with cross-stage consistency constraints (CLC-SIE), including a building instance prediction module, an object dynamic alignment module, and a pixel saliency constraint module. Among them, the building instance prediction module includes two instance extraction models, a teacher model and a student model, which are used to obtain object-level predictions of building categories and bounding boxes in remote sensing images, as well as pixel-level predictions of building masks; the object dynamic alignment module is used to calculate the consistency loss of the building categories and bounding boxes output by the teacher model and the student model; the pixel saliency constraint module is used to calculate the consistency loss of the building masks output by the teacher model and the student model.
[0036] S32. Input the labeled data into the student model to predict the building categories, bounding boxes, and masks in the image, and calculate the supervision loss with the ground truth labels , the formula is as follows:
[0037] Among them, represents the number of labeled data, , and represent the categories, bounding boxes, and masks predicted by the student model, , and represent the category labels, bounding box labels, and mask labels of the real buildings in the labeled data.
[0038] S33. Input the normally augmented unlabeled data into the teacher model to predict the pseudo-labels including building categories , bounding boxes and masks ; at the same time, input the heavily augmented unlabeled data into the student model to output the predicted values including building category scores , bounding boxes and masks .
[0039] S34. Set a threshold , and divide the above predicted values of the student model into foreground objects and background objects according to the category scores, the formula is as follows:
[0040] S35. Calculate the dynamic weight according to the above background objects , the formula is as follows:
[0041] Among them, represents the classification score of the j-th background target, represents the number of background targets.
[0042] S36, as Figure 4 shown, input the building class scores in the above pseudo-labels , the foreground targets in the prediction values and background targets , as well as the dynamic weights into the target dynamic alignment module to calculate the unsupervised class loss , the formula is as follows:
[0043] where, represents the number of foreground targets, represents the cross-entropy classification loss.
[0044] S37, according to the building bounding boxes in the above pseudo-labels and the foreground targets in the prediction values , filter the regression box with the maximum IOU , the formula is as follows:
[0045] S38, according to the building bounding boxes in the above pseudo-labels and the regression box calculate the unsupervised regression loss , the formula is as follows:
[0046] where, represents the bounding box regression loss.
[0047] S39, extract the saliency boundary according to the above unlabeled data , the formula is as follows:
[0048] where, represents the Laplacian boundary extraction operator.
[0049] Figure 5 S310, as Figure 5 shown, input the building mask in the above pseudo-labels , the mask in the student model prediction values and the saliency boundary into the pixel saliency constraint module to calculate the unsupervised mask loss , the formula is as follows:
[0050] Among them, represents the number of building masks in the pseudo-label, represents the cross-entropy mask loss.
[0051] S311. According to the above-mentioned unsupervised class loss , unsupervised regression loss and unsupervised mask loss calculate the unsupervised total loss :
[0052] Among them, represents the number of unlabeled data.
[0053] S312. According to the above-mentioned supervised loss and unsupervised total loss , calculate the total loss of the student model, and through gradient descent and backpropagation, thus train the weights of the student model . The total loss formula is as follows:
[0054] S313. According to the weights of the above-mentioned student model , update the weights of the teacher model , the formula is as follows:
[0055] Among them, represents the smoothing hyperparameter.
[0056] S314. Iteratively train the student model and the teacher model in an end-to-end manner to obtain a building instance extraction model.
[0057] S4. Input the remote sensing image to be subjected to building extraction into the trained building instance extraction model, and extract the corresponding building instance prediction result in the remote sensing image.
[0058] Next, the semi-supervised building instance extraction system provided by the present invention based on dynamic alignment and saliency constraint will be described. The semi-supervised building instance extraction system based on dynamic alignment and saliency constraint described below can be mutually corresponded and referred to with the semi-supervised building instance extraction method described above.
[0059] Figure 6 is a schematic structural diagram of the semi-supervised building instance extraction system provided by the embodiment of the present invention. As Figure 6 shown, it includes: a preprocessing module 61, an enhancement module 62, a training module 63, and a processing module 64, where: The preprocessing module 61 is used to obtain high-resolution remote sensing images, crop and divide the high-resolution remote sensing images, and perform instance-level annotation on buildings in some of the high-resolution remote sensing images, so as to obtain a small amount of data with building instance labels and a large amount of data without building instance labels; the enhancement module 62 is used to perform data enhancement on the data without building instance labels to different degrees to obtain severely enhanced unlabeled data and normally enhanced unlabeled data; the training module 63 is used to construct a semi-supervised extraction network CLC-SIE with cross-stage consistency constraints, and input the data with building instance labels, the severely enhanced unlabeled data, and the normally enhanced unlabeled data into the CLC-SIE for training to obtain a building instance extraction model; the processing module 64 is used to input the building remote sensing image to be extracted into the building instance extraction model and output a building instance prediction result.
[0060] Figure 7 An example of the physical structure diagram of an electronic device is shown as Figure 7 shown. The electronic device may include: a processor 710, a communications interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communications interface 720, and the memory 730 complete mutual communication through the communication bus 740. The processor 710 can call the logical instructions in the memory 730 to execute a semi-supervised building instance extraction method based on dynamic alignment and saliency constraints. The method includes: obtaining high-resolution remote sensing images, cropping and dividing the high-resolution remote sensing images, and performing instance-level annotation on buildings in some of the high-resolution remote sensing images to obtain a small amount of data with building instance labels and a large amount of data without building instance labels; performing data enhancement on the data without building instance labels to different degrees to obtain severely enhanced unlabeled data and normally enhanced unlabeled data; constructing CLC-SIE, inputting the data with building instance labels, the severely enhanced unlabeled data, and the normally enhanced unlabeled data into the CLC-SIE for training to obtain a building instance extraction model; inputting the building remote sensing image to be extracted into the building instance extraction model and outputting a building instance prediction result.
[0061] In addition, when the logical instructions in the above-mentioned memory 730 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0062] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the semi-supervised building instance extraction method based on dynamic alignment and saliency constraints provided by the above-mentioned various methods. The method includes: acquiring a high-resolution remote sensing image, cropping and dividing the high-resolution remote sensing image, and performing instance-level annotation on buildings in part of the high-resolution remote sensing image to obtain a small amount of building instance label data and a large amount of non-building instance label data; performing data augmentation on the non-building instance label data to different degrees to obtain severely augmented unlabeled data and normally augmented unlabeled data; constructing a CLC-SIE, inputting the building instance label data, the severely augmented unlabeled data, and the normally augmented unlabeled data into the CLC-SIE for training to obtain a building instance extraction model; inputting the building remote sensing image to be extracted into the building instance extraction model and outputting a building instance prediction result.
[0063] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the semi-supervised building instance extraction method based on dynamic alignment and saliency constraints provided by the above-mentioned various methods. The method includes: obtaining a high-resolution remote sensing image, cropping and dividing the high-resolution remote sensing image, and performing instance-level annotation on buildings in part of the high-resolution remote sensing image to obtain a small amount of building instance label data and a large amount of unlabeled building instance data; performing data augmentation on the unlabeled building instance data to different degrees to obtain severely augmented unlabeled data and normally augmented unlabeled data; constructing a CLC-SIE, inputting the building instance label data, the severely augmented unlabeled data, and the normally augmented unlabeled data into the CLC-SIE for training to obtain a building instance extraction model; inputting the remote sensing image of the building to be extracted into the building instance extraction model, and outputting a building instance prediction result.
[0064] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0065] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0066] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A semi-supervised building instance extraction method based on dynamic alignment and saliency constraints, characterized in that: include: Acquire high-resolution remote sensing images, crop and divide the high-resolution remote sensing images, and perform instance-level annotation on buildings in some high-resolution remote sensing images to obtain a small amount of data containing building instance labels and a large amount of data without building instance labels; Performing different degrees of data enhancement on the unlabeled building instance data to obtain heavily enhanced unlabeled data and generally enhanced unlabeled data; Constructing a semi-supervised extraction network CLC-SIE with cross-stage consistency constraints, inputting the building instance label data, the heavily enhanced unlabeled data and the ordinary enhanced unlabeled data into the CLC-SIE for training, and obtaining a building instance extraction model; The remote sensing image of the building to be extracted is input into the building instance extraction model, and the building instance prediction result is output.
2. The semi-supervised building instance extraction method based on dynamic alignment and saliency constraints according to claim 1, characterized in that: Obtain high-resolution remote sensing images, crop and divide the high-resolution remote sensing images, and perform instance-level annotation on buildings in some high-resolution remote sensing images to obtain a small amount of data containing building instance labels and a large amount of data without building instance labels, including: Obtaining the high-resolution remote sensing image through existing satellite remote sensing image data or aerial photography equipment; The high-resolution remote sensing images are annotated at the instance level by manual annotation to obtain a small amount of building instance label data including category labels, bounding box labels and mask labels of all buildings, and the remaining images are used as the data without building instance labels.
3. The semi-supervised building instance extraction method based on dynamic alignment and saliency constraints according to claim 1, characterized in that: The unlabeled building instance data is subjected to different degrees of data enhancement to obtain heavily enhanced unlabeled data and ordinary enhanced unlabeled data, including: The image without building instance label data is subjected to size transformation, rotation, definition transformation, Gaussian noise and image inversion according to random probability to obtain the heavily enhanced unlabeled data; The image without building instance label data is resized, randomly rotated, sharpened and brightened according to random probability to obtain the common enhanced unlabeled data.
4. The semi-supervised building instance extraction method based on dynamic alignment and saliency constraints according to claim 1, characterized in that: Construct CLC-SIE, including: Determining that the CLC-SIE includes a building instance prediction module, a target dynamic alignment module and a pixel saliency constraint module; The building instance prediction module includes a teacher model and a student model for obtaining object-level predictions of building categories and bounding boxes in remote sensing images, as well as pixel-level predictions of building masks; The target dynamic alignment module is used to calculate the consistency loss of the teacher model and the student model, and output the building category and the bounding box; The pixel saliency constraint module is used to calculate the consistency loss of the teacher model and the student model and output the building mask.
5. The semi-supervised building instance extraction method based on dynamic alignment and saliency constraints according to claim 4, characterized in that: Inputting the building instance label data, the heavily enhanced unlabeled data and the ordinary enhanced unlabeled data into the CLC-SIE for training to obtain a building instance extraction model, including: The data containing the building instance labels is input into the student model to predict the building category, bounding box and mask in the remote sensing image, and the supervision loss is calculated with the true value label. : in, Indicates the number of labeled data. , and represents the categories, bounding boxes, and masks predicted by the student model, , and Represents the category labels, bounding box labels, and mask labels of real buildings in the labeled data. Represents any tag data; The ordinary enhanced unlabeled data is input into the teacher model, and the prediction is obtained including the building category , Bounding Box and mask Pseudo-labels ; The heavily enhanced unlabeled data is input into the student model to predict the score of the building category. , Bounding Box and mask The predicted value of ; Determine the threshold , the predicted values of the student model are divided into foreground targets according to the category scores and background targets : According to the background target Calculating dynamic weights : in, represents the classification score of the jth background target, Indicates the number of background targets; The pseudo label Building category score in and the prospect target in the predicted value , Background target And dynamic weight Input target dynamic alignment module to calculate unsupervised category loss : in, represents the number of foreground objects, represents the cross entropy classification loss; Based on the building bounding boxes in the pseudo-labels and the prospect target in the predicted value , filter the regression box with the largest IOU : Based on the building bounding boxes in the pseudo-labels and regression box Calculating unsupervised regression loss : in, represents the bounding box regression loss; Based on unlabeled data Extracting salient boundaries : in, represents the Laplace boundary extraction operator; Mask the buildings in the pseudo-labels , Masks in student model predictions and saliency boundaries Input pixel saliency constraint module to calculate unsupervised mask loss : in, represents the number of building masks in the pseudo-label, represents the cross entropy mask loss; According to the above unsupervised category loss , unsupervised regression loss and unsupervised mask loss Calculate the unsupervised total loss : in, Represents the number of unlabeled data; According to the supervision loss and the unsupervised total loss Calculate the total loss of the student model and train the weights of the student model through gradient descent and backpropagation , and the total loss is : According to the weight of the student model , update the weights of the teacher model : in, represents the smoothing hyperparameter; The student model and the teacher model are iteratively trained in an end-to-end manner to obtain a building instance extraction model.
6. A semi-supervised building instance extraction system based on dynamic alignment and saliency constraints, characterized in that: include: A preprocessing module is used to obtain high-resolution remote sensing images, crop and divide the high-resolution remote sensing images, and perform instance-level annotation on buildings in some high-resolution remote sensing images to obtain a small amount of data containing building instance labels and a large amount of data without building instance labels; An enhancement module, used for performing different degrees of data enhancement on the unlabeled building instance data to obtain heavily enhanced unlabeled data and ordinary enhanced unlabeled data; A training module is used to construct a semi-supervised extraction network CLC-SIE with cross-stage consistency constraints, and the building instance label data, the heavily enhanced unlabeled data and the ordinary enhanced unlabeled data are input into the CLC-SIE for training to obtain a building instance extraction model; The processing module is used to input the remote sensing image of the building to be extracted into the building instance extraction model and output the building instance prediction result.
7. The semi-supervised building instance extraction system based on dynamic alignment and saliency constraints according to claim 6, characterized in that: The preprocessing module is specifically used for: Obtaining the high-resolution remote sensing image through existing satellite remote sensing image data or aerial photography equipment; The high-resolution remote sensing images are annotated at the instance level by manual annotation to obtain a small amount of building instance label data including category labels, bounding box labels and mask labels of all buildings, and the remaining images are used as the data without building instance labels.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the semi-supervised building instance extraction method based on dynamic alignment and saliency constraints is implemented as described in any one of claims 1 to 5.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the semi-supervised building instance extraction method based on dynamic alignment and saliency constraints as described in any one of claims 1 to 5 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the semi-supervised building instance extraction method based on dynamic alignment and saliency constraints as described in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Weak supervision remote sensing image rotating target detection method based on view consistency network
CN118097447A
Computer System and Method for Batch Data Alignment with Active Learning In Batch Process Modeling, Monitoring, And Control
US20220035348A1
Mutual learning-based semi-supervised medical image segmentation method and system
WO2023116635A1