Building semantic segmentation method for correcting unsupervised domain adaptive pseudo tag by using segmentation large model

By segmenting the large model to correct the unsupervised domain adaptation pseudo-label, the problem of feature changes caused by different imaging conditions and sensor characteristics during building extraction in remote sensing images is solved, and high-precision building extraction under no labeling conditions is achieved, improving automation level and economic benefits.

CN120014643APending Publication Date: 2025-05-16ZHEJIANG COLLEGE OF SECURITY TECH +2
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411912461.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the prior art, when extracting buildings in remote sensing images, there are feature changes due to different imaging conditions and sensor characteristics. The fully supervised semantic segmentation method relies on high-quality labels. When the training data is inconsistent with the test data, the model performance is severely degraded, making it difficult to meet the actual application needs.

Method used

The semantic segmentation method of building adapted to pseudo-labels for unsupervised domains is adopted to correct the unsupervised domain. By collecting the building data sets of the source domain and the target domain, the Advent model is trained to generate the initial pseudo-labels, and the segmentation results of the SAM model are used for label correction. Finally, the Deeplabv3+ semantic segmentation network is input for training and building semantic segmentation.

Benefits of technology

High-precision building extraction under no labeling conditions is achieved, reducing the problem of inaccurate building boundaries, improving the level of automation, reducing labor costs, and improving economic benefits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014643A_ABST
    Figure CN120014643A_ABST
Patent Text Reader

Abstract

The invention discloses a building semantic segmentation method for correcting an unsupervised domain adaptive pseudo tag by using a large segmentation model. The method comprises the following steps: collecting a source domain building data set with similar building characteristics with a target domain building data set; generating an initial target domain pseudo tag and an SAM segmentation result based on the source domain building data set and the target domain building data set; correcting the initial target domain pseudo-tag based on a tag correction algorithm by using an SAM segmentation result to obtain a final target domain pseudo-tag; a Deeplabv3 + semantic segmentation network is constructed; training a Deeplabv3 + semantic segmentation network based on the final target domain pseudo tag and the target domain image; and inputting a to-be-detected image to the trained Deeplabv3 + semantic segmentation network, and outputting the position and boundary information of the building in the to-be-detected image to realize building semantic segmentation. According to the invention, full-automatic high-precision building segmentation on the remote sensing image is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of remote sensing image processing and information extraction, and relates to a building detection method for high-resolution remote sensing images, and specifically to a building semantic segmentation method that uses a segmentation macro model (SAM) to correct unsupervised domain adaptation pseudo labels. Background Art

[0002] Building extraction plays a key role in urban planning, urban dynamic monitoring, urban growth analysis, illegal building identification, and geographic information system updates. Traditional manual extraction methods based on ground surveys and censuses are time-consuming and costly. With the development of deep learning technology and the wide availability of satellite image data, it has become possible to automatically extract building information from satellite images. However, due to different imaging conditions and sensor characteristics, buildings may show significant changes in different remote sensing images. In addition, due to the influence of weather conditions, terrain, and differences between sensors, even if the remote sensing images are taken in the same area, the characteristics of the same building vary greatly between remote sensing images.

[0003] At present, the mainstream method is the fully supervised semantic segmentation method represented by the fully convolutional neural network, which realizes the automatic extraction of remote sensing ground objects in an end-to-end mode. This type of semantic segmentation method relies on sufficient high-quality labels to ensure the accuracy of the method. When the source of training data and test data is inconsistent, the model performance is seriously degraded. Weakly supervised learning (WSL) technology, which can use incomplete, inaccurate or imprecise labels, is a technical approach to solve the problem of promoting deep learning models. Among them, domain adaptation (UDA) refers to the transfer learning technology for source domains with labels and target domains without labels. The UDA method reduces domain differences by designing models from the aspects of data, features, and output space, and can effectively transfer source domain knowledge to the target domain, thereby improving the classification performance and generalization ability of the model in the target domain. However, the accuracy of this type of weakly supervised method is still significantly lower than that of the fully supervised method, and it is difficult to meet the needs of actual applications.

[0004] Recently, a method using the Segment anything model (SAM) has emerged. However, as an unsupervised image segmentation method, the segmentation patches obtained by SAM lack semantic information. In order to obtain semantic information and achieve instance or semantic level segmentation, it is necessary to combine large language models such as CLIP to generate specific category masks using prompt points, prompt boxes, or prompt words. This processing is highly dependent on prompt conditions and is often not accurate in complex scene conditions.

[0005] Subsequently, some methods combining weakly supervised learning with SAM emerged, aiming to give full play to the respective technical advantages of the two and achieve high-quality semantic information extraction at a lower annotation cost. The basic idea is to generate "pseudo-labels" with a certain degree of accuracy as prompt information through weakly supervised methods, guiding SAM to more accurately locate and segment the target objects in the image to complete the process of information extraction from coarse to fine. For example, (He et al,.2024) et al. in "Anefficient urban flood mapping framework towards disaster response driven by weakly supervised semantic segmentation with decoupled training samples" (see ISPRS Journal of Photogrammetry and Remote Sensing, 2024, 207, 338-358) input the prompt point information into the SAM model by manual clicking, and the SAM mask decoder generates the object segmentation result, so as to annotate the flood training samples, speed up the production of training data, and thus serve the urban flood disaster response faster. (Yang et al,.2024) et al., in the weakly supervised segmentation framework proposed in "A novel weakly-supervised method based on the segmentanything model for seamless transition from classification to segmentation: A case study in segmenting latent photovoltaic locations" (see International Journal of Applied Earth Observation and Geoinformation, 2024, 130, 103929), the label after image classification is used as a rough pseudo-label, and a sampling prompt point density algorithm is designed for the label and the original image, and the rough pseudo-label is converted into a point prompt input into SAM, so that SAM is used to accurately refine the rough pseudo-label so that it is aligned with the actual contour of the target object. Finally, the refined pseudo-label is input into UNET training to complete the extraction of the target area.(Sun et al., 2024) et al. in "G2LDIE: Global-to-Local Dynamic Information Enhancement Framework for Weakly Supervised Building Extraction From Remote Sensing Images" (IEEE Transactions on Geoscience and Remote Sensing, 2024, 62, 1-14) designed a global-to-local weakly supervised segmentation framework based on image annotation. The randomly cropped global image (256×256) and local image (64×64) were input into the ViT classification network to obtain contextual information and local details. At the same time, the local contrast loss was introduced to strengthen the perceptual consistency of the model when processing images of different scales, and to promote the model to generate more accurate maximum activation responses of buildings, and generate pseudo labels for buildings accordingly. Finally, the pseudo labels were post-processed by combining the edge refinement capability of the SAM model to significantly enhance the model's building extraction performance. (Kim et al., 2024) et al. in "Integrated Framework for Unsupervised Building Segmentation with SegmentAnything Model-Based Pseudo-Labeling and Weakly Supervised Learning" (RemoteSensing, 2024, 16) combined SAM, spectral index and CANNY edge detection to design screening rules to generate pseudo labels, and simultaneously input pseudo labels and corresponding edge map training models in the weakly supervised learning stage to complete the accurate segmentation of building areas. Existing methods combine SAM with CAM activation maps, image classification, or knowledge rule methods to obtain pseudo labels, but no research has focused on the combination of SAM and unsupervised UDA to generate pseudo labels. Summary of the invention

[0006] Purpose of the invention: The purpose of the present invention is to provide a building semantic segmentation method that uses a large segmentation model to correct unsupervised domain adaptation pseudo-labels to achieve fully automatic and high-precision building segmentation on remote sensing images.

[0007] Technical solution: The method for semantic segmentation of buildings using a large segmentation model to correct unsupervised domain adaptation pseudo-labels of the present invention comprises the following steps:

[0008] S1. Collect a source domain building dataset that has similar building features to the target domain building dataset. The target domain building dataset is unlabeled, while the source domain building dataset is labeled.

[0009] S2. Train the Advent model based on the source domain building dataset and the target domain building dataset, use the trained Advent model to predict the target domain building data, and binarize the prediction results to obtain the initial target domain pseudo-label; input the target domain building dataset into the SAM model to obtain the remote sensing image segmentation result; use the remote sensing image segmentation result to correct the initial target domain pseudo-label based on the label correction algorithm to obtain the final target domain pseudo-label;

[0010] S3, build a Deeplabv3+ semantic segmentation network for building segmentation in the target domain building dataset;

[0011] S4, training the Deeplabv3+ semantic segmentation network, inputting the final target domain pseudo-label and the target domain building dataset generated in step S2 into the Deeplabv3+ semantic segmentation network constructed in step S3; training the network to obtain the trained Deeplabv3+ semantic segmentation network;

[0012] S5. Input the target domain building dataset to be detected into the Deeplabv3+ semantic segmentation network trained in step S4, directly output the location and boundary information of the buildings in the target domain building dataset to be detected, and realize the semantic segmentation of the buildings.

[0013] Furthermore, in step S1, unified data preprocessing is performed on the target domain building dataset and the source domain building dataset, including cropping images of the same pixel size using non-overlapping sliders, and dividing the target domain building dataset and the source domain building dataset into a training dataset and a test dataset.

[0014] Furthermore, the objective function used by the generator network F of the Advent model in step S2 is:

[0015]

[0016] Among them, θ # are the parameters of the generator network F, is the source domain building data sample set, is the number of source domain building data samples, x s is a single sample of source domain building data, y s The label corresponding to a single sample of the source domain building data, λ adv is the adversarial weight in the DCGAN network training process, is the target domain building data sample set, is the number of building data samples in the target domain, x t is a single sample of the building data in the target domain, is the loss of a single sample of the source domain building dataset, is the cross entropy loss of the discriminator D for judging the target domain building data sample as the source domain building data sample, is a single sample x of the target domain building data t Weighted self-information graph of;

[0017] The objective function used by the discriminator network D is:

[0018]

[0019] in, is the cross entropy loss of the discriminator D on the source domain building data samples, is the cross entropy loss of the discriminator D on the target domain building data samples, is a single sample x of the source domain building data s The weighted self-information graph of D are the parameters of the discriminator network D.

[0020] Furthermore, in step S2, the SAM model includes an image encoder, a hint encoder and a mask decoder, and the image encoder uses a pre-trained visual ViT structure to extract feature representations of the image to obtain the final image embedding result;

[0021] The hint encoder encodes the hint provided by the user into an embedding vector; its role is to map the hint into a feature space with the same number of channels as the image embedding;

[0022] The mask decoder generates the final SAM segmentation result based on the image embedding and hint embedding.

[0023] Furthermore, the label correction algorithm in step S2 includes:

[0024] (1) Create an empty set to represent the target domain pseudo label Y t_ref+,ed ;

[0025] (2) Traverse all SAM segmentation results S , , for the i-th SAM segmentation patch S + Do the following calculations: First, calculate the initial target domain pseudo label Y t Segment patch S with the i-th SAM + The overlapping area U + :

[0026] U + =Y t ∩S +

[0027] If the overlapping area U +If it is an empty set, the initial target domain pseudo-label is not processed; otherwise, the overlapping area U is calculated. + and SAM segmentation patch S + The ratio of the number of elements is called the overlap ratio and is calculated as:

[0028]

[0029] When SAM divides the patch S + When the initial target domain pseudo-label is completely covered, if the overlap ratio is greater than the threshold, t_ref+,ed Add SAM segmentation patch S + , that is, Y t_ref+,ed =Y t ∪S + , otherwise add the overlapping area U + , that is, Y t_ref+,ed =Y t ∪U + When SAM divides the patch S + When the initial target domain pseudo-label is partially covered, if the overlap ratio is greater than the threshold, t_ref+,ed Add SAM segmentation patch S + , otherwise add the overlapping area U + ;

[0030] (3) Output the final target domain pseudo label Y t_f+,a= .

[0031] Furthermore, the Deeplabv3+ semantic segmentation network in step S3 adopts an encoder-decoder structure, and also uses dilated convolution, depthwise separable convolution and ASPP modules, where the encoder part is responsible for extracting multi-scale features of remote sensing images in the target domain building dataset, while the decoder part is used to restore the spatial resolution of remote sensing images to obtain more accurate segmentation boundaries; dilated convolution is used to expand the receptive field of the Deeplabv3+ semantic segmentation network, so that the Deeplabv3+ semantic segmentation network can capture more contextual information without reducing the resolution; depthwise separable convolution decomposes the standard convolution operation into depthwise convolution and pointwise convolution, which is used to reduce the number of parameters and computational complexity of the model while maintaining the performance of the Deeplabv3+ semantic segmentation network; the ASPP module captures multi-scale contextual information by applying dilated convolutions with different expansion coefficients in parallel, which is used to enhance the recognition ability of the Deeplabv3+ semantic segmentation network for objects of different scales.

[0032] Furthermore, in step S4, the optimizer selects Adam when training the Deeplabv3+ semantic segmentation network.

[0033] The system corresponding to the method includes:

[0034] The data unit is used to collect a source domain building dataset that has similar building features to the target domain building dataset. The target domain building dataset is unlabeled, and the source domain building dataset is labeled.

[0035] The target domain pseudo-label correction unit is used to train the Advent model based on the source domain building data set and the target domain building data set, use the trained Advent model to predict the target domain building data, and binarize the prediction results to obtain the initial target domain pseudo-label; input the target domain building data set into the SAM model to obtain the remote sensing image segmentation result; use the remote sensing image segmentation result to correct the initial target domain pseudo-label based on the label correction algorithm to obtain the final target domain pseudo-label;

[0036] Deeplabv3+ semantic segmentation network construction unit, used to build Deeplabv3+ semantic segmentation network for building segmentation in target domain building dataset;

[0037] Deeplabv3+ semantic segmentation network training unit, used to train the Deeplabv3+ semantic segmentation network, input the final target domain pseudo-label and the target domain building dataset into the constructed Deeplabv3+ semantic segmentation network; train the network to obtain the trained Deeplabv3+ semantic segmentation network;

[0038] The semantic segmentation unit is used to input the target domain building dataset to be detected into the trained Deeplabv3+ semantic segmentation network, directly output the location and boundary information of the buildings in the target domain building dataset to be detected, and realize the semantic segmentation of the buildings.

[0039] An electronic device for storing and executing the method, the device comprising:

[0040] A memory storing executable program code;

[0041] a processor coupled to the memory;

[0042] The processor calls the executable program code stored in the memory to execute the steps of the building semantic segmentation method using a large segmentation model to correct unsupervised domain adaptation pseudo-labels.

[0043] A computer-readable storage medium for storing and executing the method, wherein the computer-readable storage medium stores computer instructions, and when the computer instructions are called, they are used to execute the steps of the building semantic segmentation method using a large segmentation model to correct unsupervised domain adaptation pseudo-labels.

[0044] Beneficial effects: Compared with the prior art, the present invention has the following significant technical effects: (1) Taking buildings on remote sensing images as extraction objects, the advantages of weak supervision technology and SAM image segmentation are fully utilized to achieve high-precision building extraction under unlabeled conditions in the target domain, reducing the problem of inaccurate building boundaries; (2) Efficient remote sensing image building segmentation is achieved, effectively improving the automation level of relevant implementation units, reducing labor costs, and improving their economic benefits. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is a flow chart of the method of the present invention;

[0046] Figure 2 The original image of the remote sensing image in the target domain of the method embodiment of the present invention;

[0047] Figure 3 It is the initial target domain pseudo label map of the embodiment of the method of the present invention;

[0048] Figure 4 The figure is a segmentation result diagram of the SAM model of the method embodiment of the present invention;

[0049] Figure 5 The pseudo label map of the target domain corrected by the method embodiment of the present invention;

[0050] Figure 6 It is the final semantic segmentation result diagram of the method of the present invention;

[0051] Figure 7 Schematic diagram of the label correction process of the method of the present invention. DETAILED DESCRIPTION

[0052] The present invention is described in detail below in conjunction with the accompanying drawings and embodiments.

[0053] like Figure 1 As shown, the method of the present invention comprises the following steps:

[0054] S1, data preparation and preprocessing;

[0055] In order to achieve high-precision building extraction of target domain building data (unlabeled), source domain building data (labeled) with similar building features to target domain building data are collected. The source domain building data and target domain building data will be used to train the unsupervised domain adaptation semantic segmentation network Advent to achieve the purpose of transferring the domain knowledge of source domain building data to target domain building data. The source domain building data can select any labeled remote sensing image building data set. The target domain building data is an unlabeled remote sensing image building data set.

[0056] In the embodiment of the present invention, the source domain building data selects the INRIA building data set (with annotations), and the target domain building data selects the WHU building data set. In order to realize the extraction of buildings in the unlabeled target domain building data, the present invention first performs unified data preprocessing on the target domain building data and the source domain building data, and crops the target domain building data and the source domain building data to a size of 512×512 pixels without overlapping sliders, and divides the training data set and the test data set into a ratio of 7:3.

[0057] S2, generating target domain pseudo labels, including training an unsupervised domain adaptation (UDA) semantic segmentation network Advent to generate initial target domain pseudo labels; then inputting the initial target domain pseudo labels and the segmentation results of the SAM model into a label correction algorithm to obtain the final target domain pseudo labels;

[0058] The present invention uses the Advent method to minimize entropy through adversarial learning and generate initial target domain pseudo labels with high confidence. Specifically, source domain building data (labeled) and target domain building data (unlabeled) with similar building features to target domain building data (unlabeled) are input to generate initial target domain pseudo labels. The Advent method is an entropy minimization semi-supervised semantic segmentation method based on domain adaptation (UDA) and adversarial training.

[0059] Advent is specifically composed of a semantic segmentation network based on DeepLabV2 and an adversarial generation network based on DCGAN. The optimization problem can be expressed as:

[0060]

[0061] Among them, θ # are the parameters of the generator network F, is the source domain building data sample set, is the number of source domain building data samples, x s is a single sample of source domain building data, t s The label corresponding to a single sample of the source domain building data, λ adv is the adversarial weight in the DCGAN network training process, is the target domain building data sample set, is the number of building data samples in the target domain, x t is a single sample of the building data in the target domain, is the loss of a single sample of the source domain building dataset, is the cross entropy loss of the discriminator D for judging the target domain building data sample as the source domain building data sample, is a single sample x of the target domain building data t Weighted self-infographic. The expression is:

[0062]

[0063] Among them, H is the height of a single sample of source domain building data, W is the width of a single sample of source domain building data, and C is the number of channels of a single sample of source domain building data. Represents the one-hot label value at height h, width w and channel c. The semantic segmentation network based on DeepLabV2 is input to the single sample x of the source domain building data s The predicted probability of class c at position (h,w).

[0064] Weighted self-information graph of target domain building data I P The expression is:

[0065] I P =-P P ·logP P

[0066] Among them, P T The discriminator network D predicts the score for the pixel-level category, i.e., self-information. T Determine whether it belongs to the source domain or the target domain. During the training process, the discriminator network D and the generator network F are optimized alternately. The generator network F uses the above optimization problem as the objective function. The objective function used by the discriminator network D is:

[0067]

[0068] in, is the cross entropy loss of the discriminator D on the source domain building data samples, is the cross entropy loss of the discriminator D on the target domain building data samples, is a single sample x of the source domain building data s The weighted self-information graph of D are the parameters of the discriminator network D.

[0069] Training Advent requires source domain building data and target domain building data. This method inputs the source domain training dataset and the target domain training dataset into Advent for network training. The trained Advent model is used to predict the target domain test dataset, and the prediction results are binarized to obtain the initial target domain pseudo-label.

[0070] Then the present invention uses the SAM method to obtain the image segmentation result of the target domain image according to the default point prompt. The core of the SAM segmentation model is that it can perform image segmentation by prompting, and can be transferred to new image distributions and tasks through zero-sample learning without additional training. The SAM model consists of three main parts, namely the image encoder, the prompt encoder and the mask decoder. The image encoder usually uses a pre-trained visual Transformer (ViT) as a basis to extract the feature representation of the image. The ViT structure used by the SAM model is relatively simple, and the process of input image processing by ViT can be summarized into four main steps:

[0071] (1) Input the remote sensing image in the target domain building dataset into the SAM model and perform patch embedding through a convolutional base layer. This process divides the 16×16 pixel area into a patch, and the step size is also 16, so that the size of the feature map generated after convolution of the original remote sensing image in the target domain building dataset is reduced by 16 times, and the number of channels is mapped from 3 to 768.

[0072] (2) After completing the patch embedding, add positional embedding. Positional embedding is a learnable parameter matrix whose initial value is zero.

[0073] (3) Subsequently, the feature map with position embedding is processed by 16 Transformer blocks, of which 12 are window-based attention modules (i.e., the feature map is divided into 14×14 windows to implement the local attention mechanism), and the remaining 4 are global attention modules, which are evenly distributed between the window-based attention modules.

[0074] (4) Finally, the number of channels is reduced to 256 through two layers of convolution operations, and the final remote sensing image embedding result of the target domain building dataset is obtained.

[0075] The prompt encoder encodes the prompts provided by the user (such as points, boxes, masks) into an embedding vector. Its function is to map the prompt to a feature space with the same number of channels as the remote sensing image embedding (image embedding) in the target domain building dataset generated by the image encoder, so that it can be fused through the attention mechanism later.

[0076] The mask decoder generates the final segmentation result of the SAM model based on the image embedding and the hint embedding. The mask decoder usually adopts the Transformer architecture and can predict the masks of each object in the image.

[0077] Then, the present invention inputs the initial target domain pseudo-label generated by the unsupervised domain adaptation (UDA) semantic segmentation model Advent and the segmentation result of the SAM model obtained from the mask decoder into the label correction algorithm to obtain the final target domain pseudo-label. The label correction algorithm can be further divided into three steps:

[0078] Step 1: Create an empty set to represent the target domain pseudo label Y t_refined ;

[0079] Step 2: Traverse the segmentation results S of all SAM models , (including n SAM segmentation patches), for the i-th SAM segmentation patch S + Do the following calculations. First, calculate the initial target domain pseudo label Y t Segment patch S with the i-th SAM + The overlapping area U + :

[0080] U + =Y t ∩S +

[0081] If the overlapping area U + If is an empty set, the initial target domain pseudo-label is not processed. Otherwise, the overlap area U is calculated. + and SAM segmentation patch S + The ratio of the number of elements is called the overlap ratio and is calculated as:

[0082]

[0083] When SAM divides the patch S + When the initial target domain pseudo-label is completely covered, if the overlap ratio is greater than the threshold, t_ref+,ed Add SAM segmentation patch S + , that is, Y t_ref+,ed =Y t ∪S + , otherwise add the overlapping area U + , that is, Y t_ref+,ed =Y t ∪U + When SAM divides the patch S + When the initial target domain pseudo-label is partially covered, if the overlap ratio is greater than the threshold, t_refined Add SAM segmentation patch S i , otherwise add the overlapping area U i .

[0084] Step 3: Traverse the segmentation results S of all SAM models , , output the final target domain pseudo label Y t_final .

[0085] In the embodiment of the present invention, AUSTIN is used as the source domain building dataset and WHU is used as the target domain building dataset. A 512×512 pixel remote sensing image sample is randomly selected from the target domain building dataset WHU as the target domain building remote sensing image sample X, such as Figure 2 The target domain building remote sensing image sample X is input into the Advent model and the SAM model respectively (the Advent model is a model trained by the source domain building data and the target domain building data, and the SAM model loads the public model sam_vit_h_4b8939.pth) to output the target domain SAM segmentation result and obtain the initial target domain pseudo label Y of the building. t And SAM segmentation result S , , where Y t is the binary map of the prediction results of the Advent model, that is, the corrected target domain pseudo-label map, such as Figure 3 As shown. The segmentation result S of the SAM model , like Figure 4 As shown, it contains n SAM segmentation patches. Then, according to the label correction algorithm, the corresponding final target domain pseudo label Y of the building remote sensing image sample X previously selected from the WHU training data set is obtained. t_f+,a= , the results are as follows Figure 5 As shown in the figure, the schematic diagram of the label correction process is shown in Figure 7 .

[0086] S3. Construct a Deeplabv3+ semantic segmentation network for building segmentation in the target domain building dataset. The DeepLabV3+ semantic segmentation network adopts an encoder-decoder architecture, and also uses dilated convolution, depthwise separable convolution, and Atrous Spatial Pyramid Pooling (ASPP) modules. The encoder part is responsible for extracting the multi-scale features of the remote sensing images in the target domain building dataset, while the decoder part is used to restore the spatial resolution of the remote sensing images to obtain more accurate segmentation boundaries. Dilated convolution is used to expand the receptive field of the Deeplabv3+ semantic segmentation network, which allows the Deeplabv3+ semantic segmentation network to capture more contextual information without reducing the resolution. The depthwise separable convolution decomposes the standard convolution operation into two parts: depthwise convolution and pointwise convolution, reducing the number of parameters and computational complexity of the Deeplabv3+ semantic segmentation network while maintaining the performance of the Deeplabv3+ semantic segmentation network. The ASPP module captures multi-scale contextual information by applying dilated convolutions with different dilation coefficients in parallel, enhancing the Deeplabv3+ semantic segmentation network’s ability to recognize objects of different scales.

[0087] The encoder part usually consists of a pre-trained deep convolutional neural network as the backbone network, such as ResNet or Xception. In the DeepLabV3+ semantic segmentation network, the feature map output by the backbone network passes through the Atrous Spatial Pyramid Pooling (ASPP) module to capture multi-scale contextual information. The decoder part is mainly responsible for upsampling and feature fusion of the feature map output by the encoder to restore the spatial information of the image. In the DeepLabV3+ semantic segmentation network, the decoder first receives the low-resolution feature map from the encoder and the high-resolution feature map from the middle layer of the backbone network. Then, the number of channels of the high-resolution feature map is adjusted by performing a convolution operation on the high-resolution feature map, and it is spliced ​​with the upsampled low-resolution feature map. Finally, the final segmentation result is obtained through a series of convolution operations.

[0088] S4, train the Deeplabv3+ semantic segmentation network and convert the final target domain pseudo label Y generated in step 2 t_f+,a= The remote sensing images in the target domain building dataset are input together into the Deeplabv3+ semantic segmentation network constructed in step S33. The network is trained to obtain a Deeplabv3+ semantic segmentation network suitable for the target domain building dataset (WHU dataset in this embodiment). The optimizer is Adam, the learning rate is set to 0.001, and the training batch size is set to 4.

[0089] S5. Input the target domain building dataset to be detected into the Deeplabv3+ semantic segmentation network trained in step 4, which can directly output the location and boundary information of the buildings in the target domain building dataset to be detected, thereby realizing semantic segmentation of the buildings. Figure 6 This is the final semantic segmentation result diagram in an embodiment of the present invention. It can be observed that the edges of the buildings are relatively clear and complete.

[0090] The system corresponding to the method includes:

[0091] The data unit is used to collect a source domain building dataset that has similar building features to the target domain building dataset. The target domain building dataset is unlabeled, and the source domain building dataset is labeled.

[0092] The target domain pseudo-label correction unit is used to train the Advent model based on the source domain building data set and the target domain building data set, use the trained Advent model to predict the target domain building data, and binarize the prediction results to obtain the initial target domain pseudo-label; input the target domain building data set into the SAM model to obtain the remote sensing image segmentation result; use the remote sensing image segmentation result to correct the initial target domain pseudo-label based on the label correction algorithm to obtain the final target domain pseudo-label;

[0093] Deeplabv3+ semantic segmentation network construction unit, used to build Deeplabv3+ semantic segmentation network for building segmentation in target domain building dataset;

[0094] Deeplabv3+ semantic segmentation network training unit, used to train the Deeplabv3+ semantic segmentation network, input the final target domain pseudo-label and the target domain building dataset into the constructed Deeplabv3+ semantic segmentation network; train the network to obtain the trained Deeplabv3+ semantic segmentation network;

[0095] The semantic segmentation unit is used to input the target domain building dataset to be detected into the trained Deeplabv3+ semantic segmentation network, directly output the location and boundary information of the buildings in the target domain building dataset to be detected, and realize the semantic segmentation of the buildings.

[0096] An electronic device for storing and executing the method, the device comprising:

[0097] A memory storing executable program code;

[0098] a processor coupled to the memory;

[0099] The processor calls the executable program code stored in the memory to execute the steps of the building semantic segmentation method using a large segmentation model to correct unsupervised domain adaptation pseudo-labels.

[0100] A computer-readable storage medium for storing and executing the method, wherein the computer-readable storage medium stores computer instructions, and when the computer instructions are called, they are used to execute the steps of the building semantic segmentation method using a large segmentation model to correct unsupervised domain adaptation pseudo-labels.

[0101] Figure 2 Target domain remote sensing image for implementation example. Figure 3 This is the initial target domain pseudo-label image of the embodiment of the method of the present invention. The black area is the background and the white area is the building. It can be observed that the edge of the building is not clear. Figure 4 The figure shows the SAM segmentation result in an embodiment of the present invention, where each color block is a SAM spot. Figure 5 It is the corrected pseudo label map of the target domain in the embodiment of the present invention. Figure 6 This is the final semantic segmentation result diagram in an embodiment of the present invention. It can be observed that the edges of the buildings are relatively clear and complete. Figure 7 Schematic diagram of the label correction process.

Claims

1. A method for semantic segmentation of buildings using a large segmentation model to modify unsupervised domain adaptation pseudo-labels, characterized in that: The following steps are involved: S1. Collect a source domain building dataset that has similar building features to the target domain building dataset. The target domain building dataset is unlabeled, while the source domain building dataset is labeled. S2. Train the Advent model based on the source domain building dataset and the target domain building dataset, use the trained Advent model to predict the target domain building data, and binarize the prediction results to obtain the initial target domain pseudo-label; The target domain building dataset is input into the SAM model to obtain the remote sensing image segmentation result; the initial target domain pseudo label is corrected based on the label correction algorithm using the remote sensing image segmentation result to obtain the final target domain pseudo label; S3, build a Deeplabv3+ semantic segmentation network for building segmentation in the target domain building dataset; S4, training the Deeplabv3+ semantic segmentation network, inputting the final target domain pseudo-label and the target domain building dataset generated in step S2 into the Deeplabv3+ semantic segmentation network constructed in step S3; Train the network to obtain the trained Deeplabv3+ semantic segmentation network; S5. Input the target domain building dataset to be detected into the Deeplabv3+ semantic segmentation network trained in step S4, directly output the location and boundary information of the buildings in the target domain building dataset to be detected, and realize the semantic segmentation of the buildings.

2. According to claim 1, a method for semantic segmentation of buildings using a large segmentation model to correct unsupervised domain adaptation pseudo-labels is characterized in that: In step S1, unified data preprocessing is performed on the target domain building dataset and the source domain building dataset, including cropping images of the same pixel size without overlapping sliders, and dividing the target domain building dataset and the source domain building dataset into a training dataset and a test dataset.

3. The method for semantic segmentation of buildings using a large segmentation model to correct unsupervised domain adaptation pseudo-labels according to claim 1, characterized in that: The objective function used by the generator network F of the Advent model in step S2 is: Among them, θ " is the parameter of the generator network F, x s is the source domain building data sample set, |x s | is the number of source domain building data samples, x s is a single sample of source domain building data, y s The label corresponding to a single sample of the source domain building data, λ adv is the adversarial weight in the DCGAN network training process, X t is the target domain building data sample set, |x t | is the number of building data samples in the target domain, x t is a single sample of the building data in the target domain, is the loss of a single sample of the source domain building dataset, is the cross entropy loss of the discriminator D for judging the target domain building data sample as the source domain building data sample, is a single sample x of the target domain building data t Weighted self-information graph of; The objective function used by the discriminator network D is: in, is the cross entropy loss of the discriminator D on the source domain building data samples, is the cross entropy loss of the discriminator D on the target domain building data samples, is a single sample x of the source domain building data s The weighted self-information graph of D are the parameters of the discriminator network D.

4. The method for semantic segmentation of buildings using a large segmentation model to correct unsupervised domain adaptation pseudo-labels according to claim 1, characterized in that: In step S2, the SAM model includes an image encoder, a hint encoder, and a mask decoder. The image encoder uses the pre-trained visual ViT structure to extract the feature representation of the image and obtain the final image embedding result. The prompt encoder encodes the user-provided prompts into embedding vectors; Its role is to map the cue into a feature space with the same number of channels as the image embedding; The mask decoder generates the final SAM segmentation result based on the image embedding and hint embedding.

5. The method for semantic segmentation of buildings using a large segmentation model to correct unsupervised domain adaptation pseudo-labels according to claim 1, characterized in that: The label correction algorithm in step S2 includes: (1) Create an empty set to represent the target domain pseudo label Y t_(efi+ed ; (2) Traverse the segmentation results S of all SAM models + , for the i-th SAM segmentation patch S i Do the following calculations: First, calculate the initial target domain pseudo label Y t Segment patch S with the i-th SAM i The overlapping area U i : U i =Y t ∩S i If the overlapping area U i If it is an empty set, the initial target domain pseudo-label is not processed; otherwise, the overlapping area U is calculated. i and SAM segmentation patch S i The ratio of the number of elements is called the overlap ratio and is calculated as: When SAM divides the patch S i When the initial target domain pseudo-label is completely covered, if the overlap ratio is greater than the threshold, t_(efi+ed Add SAM segmentation patch S i , that is, Y t_(efi+ed =Y t ∪S i , otherwise add the overlapping area U i , that is, Y t_(efi+ed =Y t ∪U i ; When SAM divides the patch S i When the initial target domain pseudo-label is partially covered, if the overlap ratio is greater than the threshold, t_(efi+ed Add SAM segmentation patch S i , otherwise add the overlapping area U i ; (3) Traverse the segmentation results S of all SAM models + , output the final target domain pseudo label Y t_fi+a< .

6. The method for semantic segmentation of buildings using a large segmentation model to correct unsupervised domain adaptation pseudo-labels according to claim 1, characterized in that: In step S3, the Deeplabv3+ semantic segmentation network adopts an encoder-decoder structure, and also uses dilated convolution, depthwise separable convolution and ASPP modules. The encoder part is responsible for extracting multi-scale features of remote sensing images in the target domain building dataset, while the decoder part is used to restore the spatial resolution of remote sensing images to obtain more accurate segmentation boundaries; dilated convolution is used to expand the receptive field of the Deeplabv3+ semantic segmentation network, so that the Deeplabv3+ semantic segmentation network can capture more contextual information without reducing the resolution; depthwise separable convolution decomposes the standard convolution operation into depthwise convolution and pointwise convolution, which is used to reduce the number of parameters and computational complexity of the model while maintaining the performance of the Deeplabv3+ semantic segmentation network; the ASPP module captures multi-scale contextual information by applying dilated convolutions with different expansion coefficients in parallel, which is used to enhance the recognition ability of the Deeplabv3+ semantic segmentation network for objects of different scales.

7. The method for semantic segmentation of buildings using a large segmentation model to correct unsupervised domain adaptation pseudo-labels according to claim 1, characterized in that: In step S4, the optimizer is selected as Adam when training the Deeplabv3+ semantic segmentation network.

8. A building semantic segmentation system using a large segmentation model to correct unsupervised domain adaptation pseudo-labels, characterized in that: include: The data unit is used to collect a source domain building dataset that has similar building features to the target domain building dataset. The target domain building dataset is unlabeled, and the source domain building dataset is labeled. A target domain pseudo-label correction unit is used to train an Advent model based on a source domain building dataset and a target domain building dataset, use the trained Advent model to predict target domain building data, and binarize the prediction results to obtain an initial target domain pseudo-label; The target domain building dataset is input into the SAM model to obtain the remote sensing image segmentation result; the initial target domain pseudo label is corrected based on the label correction algorithm using the remote sensing image segmentation result to obtain the final target domain pseudo label; Deeplabv3+ semantic segmentation network construction unit, used to build Deeplabv3+ semantic segmentation network for building segmentation in target domain building dataset; Deeplabv3+ semantic segmentation network training unit, used to train the Deeplabv3+ semantic segmentation network, and input the final target domain pseudo-label and target domain building dataset into the constructed Deeplabv3+ semantic segmentation network; Train the network to obtain the trained Deeplabv3+ semantic segmentation network; The semantic segmentation unit is used to input the target domain building dataset to be detected into the trained Deeplabv3+ semantic segmentation network, directly output the location and boundary information of the buildings in the target domain building dataset to be detected, and realize the semantic segmentation of the buildings.

9. An electronic device, characterized in that: The device comprises: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the steps of the building semantic segmentation method using a large segmentation model to correct unsupervised domain adaptation pseudo-labels as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, which, when called, are used to execute the steps of the building semantic segmentation method for correcting unsupervised domain adaptation pseudo-labels using a large segmentation model as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Remote sensing image building semantic segmentation system based on visual language model

    CN120388178A

  • Power transmission and transformation line inspection method, system and equipment and storage medium

    CN120808065A

  • Passive domain adaptive semantic segmentation method and device for diffusion-guided pseudo-label enhancement

    CN121190764A