A cross-domain building extraction method and related device based on multi-level feature fusion

By constructing a cross-domain building extraction framework model, combining contrastive learning and SAM modules to generate high-quality pseudo labels, and deeply integrating scene-level and object-level features, the problem of poor cross-domain migration performance of building extraction in urban and rural remote sensing images is solved, and higher extraction accuracy and robustness are achieved.

CN120599482BActive Publication Date: 2025-10-03GUANGZHOU UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511097439.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-10-03
Estimated Expiration
2045-08-06

AI Technical Summary

Technical Problem

Existing deep learning methods have poor cross-domain migration performance and insufficient model generalization ability in building extraction from urban and rural remote sensing images, especially in rural areas. The extraction effect is poor, and the quality of pseudo-label generation and feature extraction is insufficient. It is difficult to effectively combine global scene information with fine-grained target information, resulting in insufficient extraction accuracy and robustness.

Method used

A cross-domain building extraction method with multi-level feature fusion is adopted. By constructing a cross-domain building extraction framework model, combining contrastive learning, SAM module and bidirectional cross-attention module, high-quality object-level masks and pseudo-labels are generated, and scene-level and object-level features are integrated to optimize the model's cross-domain adaptability in the absence of annotations.

Benefits of technology

The model significantly improves the accuracy and robustness of building extraction in rural areas, and can achieve higher extraction accuracy and stronger cross-domain adaptability on multiple urban and rural remote sensing image datasets. It solves the problem of unstable pseudo-label quality in traditional methods and ensures the accurate positioning and extraction of building boundaries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599482B_ABST
    Figure CN120599482B_ABST
Patent Text Reader

Abstract

The present invention discloses a cross-domain building extraction method and related device with multi-level feature fusion, which includes: constructing a cross-domain building extraction framework model and performing scene-level feature learning; performing object-level feature training and learning to form a trained cross-domain building extraction framework model; inputting the scene-level feature data and object-level feature data extracted during training and learning into a bidirectional cross-attention module for feature fusion processing to form fused feature data; adjusting the trained cross-domain building extraction framework model based on the fused feature data to form an adjusted cross-domain building extraction framework model; and inputting the remote sensing image to be identified into the adjusted cross-domain building extraction framework model to perform cross-domain building feature extraction processing. In an embodiment of the present invention, higher building extraction accuracy can be achieved on multiple urban and rural remote sensing image datasets, demonstrating stronger cross-domain adaptability and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a multi-level feature fusion cross-domain building extraction method and related devices. Background Art

[0002] Remote sensing images, as an important type of earth observation data, can efficiently provide extensive spatial information about the earth's surface and are widely used in various fields such as urban planning, resource management, and environmental monitoring. In high-resolution remote sensing images, buildings are an important type of ground feature, and their extraction plays a key role in various applications. Traditional building extraction methods can be roughly divided into feature-based methods, image segmentation-based methods, and deep learning-based methods. Among them, feature-based methods rely on manually designed features, especially features such as color, texture, and shape, to distinguish buildings from the background. The advantages of this type of method are intuitiveness and high computational efficiency, but its performance is relatively limited in complex scenes, especially when there are various building types and complex backgrounds, making it difficult to provide accurate extraction results.

[0003] With the rapid development of deep learning technology, convolutional neural network (CNN)-based methods have gradually become the mainstream approach for building extraction from remote sensing imagery. By training deep neural networks, models can automatically learn effective features from large amounts of data, eliminating the need for manual feature design, significantly improving the accuracy and efficiency of building extraction. In particular, network structures such as U-Net and FCN have been widely used in remote sensing image analysis, enabling effective pixel-level building segmentation. However, while deep learning-based methods have achieved promising results on large-scale datasets, existing deep learning methods exhibit poor transferability between domains due to significant differences between urban and rural areas in terms of building morphology, background, and building density. For example, many models excel in building extraction in urban areas but often experience performance degradation in rural areas. This is primarily due to insufficient training data and the lack of generalization capabilities caused by differences in building types between urban and rural areas. Furthermore, with the continuous advancement of remote sensing technology, access to high-resolution remote sensing imagery has become increasingly common. However, efficiently and accurately extracting buildings in rural areas or against complex backgrounds remains a major challenge in remote sensing image processing.

[0004] The challenges of building extraction primarily include the following: First, the morphology of urban and rural buildings differs significantly. Urban buildings are often high-rise or densely packed, while rural buildings are typically lower and more scattered. Second, the complex background environment in rural areas can easily create texture features similar to those of buildings. Finally, existing methods typically rely on large amounts of labeled data when addressing cross-domain urban and rural issues. However, in practice, the target domain (e.g., rural areas) often lacks sufficient labeled data, making it difficult for trained models to adapt to new environments. To address these issues, domain adaptation (DA) methods have been gradually introduced to the task of extracting buildings from remote sensing images in recent years. DA methods attempt to transfer knowledge from a source domain (e.g., an urban dataset) to a target domain (e.g., a rural dataset), using unlabeled or minimally labeled data for training in the target domain, thereby reducing reliance on labeled data from the target domain. However, existing DA methods still face numerous challenges in addressing the differences between urban and rural buildings, extracting fine-grained features, and migrating across domains. In particular, effectively extracting building boundaries and details in rural areas remains an urgent problem.

[0005] Existing technologies rely heavily on target domain data in transfer learning and fail to fully optimize the building extraction effect in the target domain. Secondly, the quality control of pseudo-label generation and feature extraction is insufficient, resulting in poor cross-domain adaptability and robustness of the model when there are large urban-rural differences. Finally, existing methods fail to effectively combine global scene information with fine-grained target information, limiting the accuracy of the model in extracting rural buildings. Summary of the Invention

[0006] The purpose of the present invention is to overcome the shortcomings of the existing technology. The present invention provides a cross-domain building extraction method and related devices with multi-level feature fusion, which can achieve higher building extraction accuracy on multiple urban and rural remote sensing image data sets, and show stronger cross-domain adaptability and robustness.

[0007] In order to solve the above technical problems, an embodiment of the present invention provides a cross-domain building extraction method using multi-level feature fusion, the method comprising:

[0008] A cross-domain building extraction framework model is constructed, which processes scene-level features in unlabeled urban and rural remote sensing image sets based on contrastive learning.

[0009] The SAM module is used to process unlabeled urban and rural remote sensing images to generate object-level masks; the cross-domain building extraction framework model uses the object-level masks as supervision signals to perform object-level feature training and learning to form a trained cross-domain building extraction framework model;

[0010] The scene-level feature data and object-level feature data extracted during training are input into the bidirectional cross-attention module of the trained cross-domain building extraction framework model for feature fusion processing to form fused feature data;

[0011] Adjusting the trained cross-domain building extraction framework model based on the fused feature data to form an adjusted cross-domain building extraction framework model;

[0012] A remote sensing image to be identified is obtained, and the remote sensing image to be identified is input into the adjusted cross-domain building extraction framework model to perform feature extraction processing on the cross-domain buildings.

[0013] Optionally, the cross-domain building extraction framework model performs scene-level feature learning processing on an unlabeled urban and rural remote sensing image set based on contrastive learning, including:

[0014] The cross-domain building extraction framework model performs random data enhancement processing on any urban and rural remote sensing image in the input unlabeled urban and rural remote sensing image set to form a first image and a second image with similar content but different forms of expression;

[0015] Inputting the first image and the second image into a shared encoder for high-dimensional feature extraction processing to obtain a first high-dimensional feature vector corresponding to the first image and a second high-dimensional feature vector corresponding to the second image;

[0016] Mapping the first high-dimensional feature vector and the second high-dimensional feature vector to a low-dimensional feature space using a multi-layer perception mechanism to obtain a first low-dimensional feature vector corresponding to the first high-dimensional feature vector and a second low-dimensional feature vector corresponding to the second high-dimensional feature vector;

[0017] The first low-dimensional feature vector and the second low-dimensional feature vector are used to learn scene-level features based on the InfoNCE contrast loss function.

[0018] Optionally, the formula of the InfoNCE contrast loss function is as follows:

[0019] ;

[0020] in, is the first low-dimensional eigenvector; is the second lowest dimensional eigenvector; It is a hyperparameter used to control the impact of negative samples; is the number of samples in a batch, .

[0021] Optionally, the SAM module is used to process unlabeled urban and rural remote sensing images to generate object-level masks; the cross-domain building extraction framework model uses the object-level masks as supervision signals to perform object-level feature training and learning, including:

[0022] In the cross-domain building extraction framework model, the SAM module is used to perform target object extraction processing on each urban and rural remote sensing image in the unlabeled urban and rural remote sensing image set under unsupervised conditions to generate a target object mask;

[0023] Performing target and background segmentation processing on the target object mask to form a segmented target and a segmented background;

[0024] The cross-domain building extraction framework model uses a binary cross entropy loss function to perform object-level feature learning on the segmentation target and the segmentation background.

[0025] Optionally, the binary cross entropy loss function is as follows:

[0026] ;

[0027] in, The cross-domain building extraction framework model predicts the The probability that a pixel belongs to the target; For the The supervised labels for pixels; is the number of pixels;

[0028] Taking into account the problem of class imbalance, Focal Loss is added to strengthen the focus on difficult-to-classify samples, as follows:

[0029] ;

[0030] in, The cross-domain building extraction framework model predicts the probability of the target category; is the category balance factor; To adjust the focus parameters of difficult and easy samples; therefore, the total loss function is as follows:

[0031] .

[0032] Optionally, the scene-level feature data and object-level feature data extracted during training and learning are input into the bidirectional cross-attention module of the trained cross-domain architecture extraction framework model for feature fusion processing to form fused feature data, including:

[0033] Obtaining scene-level feature data and object-level feature data extracted during training and learning of the cross-domain building extraction framework model;

[0034] After inputting the scene-level feature data and the object-level feature data into the bidirectional cross-attention module, convolution processing is performed through a convolution layer with convolution kernels of different sizes in the bidirectional cross-attention module to obtain first multi-scale information corresponding to the scene-level feature data and second multi-scale information corresponding to the object-level feature data;

[0035] performing enhancement processing on the first multi-scale information using object-level feature data based on the query vector, the key vector, and the value vector to form enhanced scene-level feature data;

[0036] performing enhancement processing on the second multi-scale information using scene-level feature data based on the query vector, the key vector, and the value vector to form enhanced object-level feature data;

[0037] The enhanced scene-level feature data and the enhanced object-level feature data are fused by adding them together to form fused feature data.

[0038] Optionally, adjusting the trained cross-domain building extraction framework model based on the fused feature data to form an adjusted cross-domain building extraction framework model includes:

[0039] The trained cross-domain building extraction framework model is used to predict and process unlabeled urban and rural remote sensing images to generate pseudo labels in the target domain.

[0040] The target domain pseudo-labels are combined with the real label data of the source domain as supervisory signals, and pixel-level supervised learning is performed on the fused feature data to adjust the trained cross-domain building extraction framework model.

[0041] In addition, an embodiment of the present invention further provides a cross-domain building extraction device using multi-level feature fusion, the device comprising:

[0042] The first learning module is used to build a cross-domain building extraction framework model, which is based on contrastive learning to learn scene-level features in unlabeled urban and rural remote sensing image sets; at the same time,

[0043] The second learning module is used to process the unlabeled urban and rural remote sensing images using the SAM module to generate object-level masks; the cross-domain building extraction framework model uses the object-level masks as supervision signals to perform object-level feature training and learning processing to form a trained cross-domain building extraction framework model;

[0044] Feature fusion module: used to input the scene-level feature data and object-level feature data extracted during training into the bidirectional cross-attention module of the trained cross-domain building extraction framework model for feature fusion processing to form fused feature data;

[0045] Adjustment module: used to adjust the trained cross-domain building extraction framework model based on the fused feature data to form an adjusted cross-domain building extraction framework model;

[0046] Feature extraction module: used to obtain the remote sensing image to be identified, and input the remote sensing image to be identified into the adjusted cross-domain building extraction framework model to perform feature extraction processing on cross-domain buildings.

[0047] In addition, an embodiment of the present invention further provides an electronic device, including a processor and a memory, wherein the processor runs a computer program or code stored in the memory to implement the cross-domain building extraction method as described in any one of the above.

[0048] In addition, an embodiment of the present invention further provides a computer-readable storage medium for storing a computer program or code. When the computer program or code is executed by a processor, the cross-domain building extraction method as described in any one of the above is implemented.

[0049] In an embodiment of the present invention, the global semantic information and fine-grained target features of urban and rural buildings are captured, thereby significantly improving the extraction accuracy and robustness of the model in rural areas; by introducing a contrastive learning framework and SegmentAnything Model (SAM), high-quality target-level pseudo labels are generated without target domain annotation, further optimizing the cross-domain adaptability and solving the problem of unstable pseudo-label quality in traditional methods; a bidirectional cross-attention mechanism deeply integrates global scene features with local target features, ensuring the subsequent precise positioning and extraction of building boundaries; thereby achieving higher building extraction accuracy on multiple urban and rural remote sensing image datasets, demonstrating stronger cross-domain adaptability and robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0051] Figure 1 1 is a flow chart of a cross-domain building extraction method using multi-level feature fusion according to an embodiment of the present invention;

[0052] Figure 2 1 is a flow chart of a cross-domain building extraction method using multi-level feature fusion in another embodiment of the present invention;

[0053] Figure 3Schematic diagram of the structure of a cross-domain building extraction device with multi-level feature fusion in an embodiment of the present invention;

[0054] Figure 4 is a schematic diagram of the structure of an electronic device in an embodiment of the present invention;

[0055] Figure 5 This is an architectural diagram of the cross-domain building extraction framework model in an embodiment of the present invention. DETAILED DESCRIPTION

[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0057] For example 1, please refer to Figure 1 , Figure 1 4 is a flow chart of a cross-domain building extraction method using multi-level feature fusion in an embodiment of the present invention.

[0058] like Figure 1 As shown, a cross-domain building extraction method based on multi-level feature fusion, the method includes:

[0059] S101: Constructing a cross-domain building extraction framework model, wherein the cross-domain building extraction framework model performs scene-level feature learning processing in a set of unlabeled urban and rural remote sensing images based on contrastive learning; at the same time,

[0060] In the specific implementation process of the present invention, the cross-domain building extraction framework model performs scene-level feature learning processing in an unlabeled urban and rural remote sensing image set based on contrastive learning, including: the cross-domain building extraction framework model performs random data enhancement processing on any urban and rural remote sensing image in the input unlabeled urban and rural remote sensing image set to form two first images and second images with similar content but different forms of expression; the first image and the second image are input into a shared encoder for high-dimensional feature extraction processing to obtain a first high-dimensional feature vector corresponding to the first image and a second high-dimensional feature vector corresponding to the second image; the first high-dimensional feature vector and the second high-dimensional feature vector are mapped to a low-dimensional feature space using a multi-layer perception mechanism to obtain a first low-dimensional feature vector corresponding to the first high-dimensional feature vector and a second low-dimensional feature vector corresponding to the second high-dimensional feature vector; based on the InfoNCE contrast loss function, the first low-dimensional feature vector and the second low-dimensional feature vector are used to perform scene-level feature learning processing.

[0061] Furthermore, the formula of the InfoNCE contrast loss function is as follows:

[0062] ;

[0063] in, is the first low-dimensional eigenvector; is the second lowest dimensional eigenvector; It is a hyperparameter used to control the impact of negative samples; is the number of samples in a batch, .

[0064] Specifically, we first need to build a cross-domain building extraction framework model, which can be referenced Figure 5 , a cross-domain building extraction framework model is constructed, aiming to optimize the cross-domain adaptability of buildings in remote sensing images by gradually extracting and fusing features at different levels, and effectively solve the problem of model performance degradation caused by the difference in urban and rural building styles. It includes two core stages: pre-training and fine-tuning; first, in the pre-training stage, the model learns common scene-level features from large-scale unlabeled urban and rural remote sensing images through comparative learning; at the same time, SAM is introduced to generate category-independent object-level pseudo labels for all images, guiding the model to grasp the boundary features of the target and background; then, in the fine-tuning stage, a bidirectional cross attention module (BCAM) is designed to deeply fuse scene and target features, and the target domain pseudo labels generated under the guidance of the source domain are used for pixel-level joint training, thereby significantly improving the accuracy of building extraction in the absence of target domain annotations.

[0065] During the pre-training phase, i.e., learning scene-level features, this embodiment uses the SimCLR contrastive learning framework to extract high-quality, generalizable scene-level features from a large number of unlabeled urban and rural remote sensing images to enable the model to understand the global structure and semantic relationships of remote sensing scenes. Specifically, the cross-domain building extraction framework model performs random data augmentation on any urban and rural remote sensing image in the input unlabeled urban and rural remote sensing image set, generating a first image and a second image with similar content but different representations. These first and second images are then fed into a shared encoder (e.g., ResNet-50) for high-dimensional feature extraction, obtaining a first high-dimensional feature vector corresponding to the first image and a second high-dimensional feature vector corresponding to the second image. Subsequently, a lightweight multi-layer perceptron (MLP) maps these high-dimensional vectors into a low-dimensional feature space, obtaining a first low-dimensional feature vector corresponding to the first high-dimensional feature vector and a second low-dimensional feature vector corresponding to the second high-dimensional feature vector. The training goal is to use the InfoNCE contrastive loss function to bring two views of the same scene (positive pairs) closer together in the feature space while simultaneously increasing the distance between two views of different scenes (negative pairs).

[0066] The formula of InfoNCE contrast loss function is as follows:

[0067] ;

[0068] in, is the first low-dimensional eigenvector; is the second lowest dimensional eigenvector; is a hyperparameter used to control the impact of negative samples, in this embodiment, Set to 0.5; is the number of samples in a batch, .

[0069] Through this unsupervised learning, the model is able to capture the intrinsic invariant properties of the scene rather than specific visual details, thereby extracting rich scene-level information covering local texture, overall layout, and environmental characteristics.

[0070] S102: Processing the unlabeled urban and rural remote sensing images using the SAM module to generate object-level masks; the cross-domain building extraction framework model uses the object-level masks as supervision signals to perform object-level feature training and learning to form a trained cross-domain building extraction framework model;

[0071] In the specific implementation process of the present invention, the SAM module is used to process the unlabeled urban and rural remote sensing images to generate object-level masks; the cross-domain building extraction framework model uses the object-level masks as supervision signals to perform training and learning processing of object-level features, including: using the SAM module in the cross-domain building extraction framework model to perform target object extraction processing on each urban and rural remote sensing image in the unlabeled urban and rural remote sensing image set under unsupervised conditions to generate a target object mask; performing target and background segmentation processing on the target object mask to form a segmented target and a segmented background; the cross-domain building extraction framework model uses a binary cross entropy loss function to perform object-level feature learning processing on the segmented target and the segmented background.

[0072] Furthermore, the binary cross entropy loss function is as follows:

[0073] ;

[0074] in, The cross-domain building extraction framework model predicts the The probability that a pixel belongs to the target; For the The supervised labels for pixels; is the number of pixels;

[0075] Taking into account the problem of class imbalance, the Focal Loss loss function is added to strengthen the focus on difficult-to-classify samples, as follows:

[0076] ;

[0077] in, The cross-domain building extraction framework model predicts the probability of the target category; is the category balance factor; To adjust the focus parameters of difficult and easy samples; therefore, the total loss function is as follows:

[0078] .

[0079] Specifically, target-level feature extraction is performed based on SAM to generate semantic-free pseudo-labels. Target-level features focus on detailed attributes such as the shape, edge, and texture of specific objects in the image, which are crucial for accurate segmentation. In this embodiment, the Segment Anything Model (SAM) is introduced to generate high-quality, semantic-free object masks (i.e., pseudo-labels) for all remote sensing images in an unsupervised manner. These masks accurately depict the boundaries of all potential targets in the image, serving as a guide for target-level feature learning. SAM processes the original image to generate an initial object mask. In order to improve mask accuracy and optimize the model's target detection capabilities, this problem is converted into a binary segmentation task of target and background. In the target and background segmentation task, a binary cross-entropy loss function is used to classify the target and background, thereby improving the accuracy of the mask and further optimizing the model's target detection capabilities. The binary cross-entropy loss function is as follows:

[0080] The binary cross entropy loss function is as follows:

[0081] ;

[0082] in, The cross-domain building extraction framework model predicts the The probability that a pixel belongs to the target; For the The supervised labels for pixels; is the number of pixels;

[0083] Taking into account the problem of class imbalance, the Focal Loss loss function is added to strengthen the focus on difficult-to-classify samples, as follows:

[0084] ;

[0085] in, The cross-domain building extraction framework model predicts the probability of the target category; is the category balance factor; To adjust the focus parameters of difficult and easy samples; therefore, the total loss function is as follows:

[0086] .

[0087] S103: Inputting the scene-level feature data and object-level feature data extracted during training into the bidirectional cross-attention module of the trained cross-domain building extraction framework model for feature fusion processing to form fused feature data;

[0088] In the specific implementation process of the present invention, the scene-level feature data and object-level feature data extracted during training and learning are input into the bidirectional cross-attention module of the trained cross-domain building extraction framework model for feature fusion processing to form fused feature data, including: obtaining the scene-level feature data and object-level feature data extracted by the cross-domain building extraction framework model during training and learning; after inputting the scene-level feature data and object-level feature data into the bidirectional cross-attention module, convolution processing is performed through a convolution layer of convolution kernels of different sizes in the bidirectional cross-attention module to obtain first multi-scale information corresponding to the scene-level feature data and second multi-scale information corresponding to the object-level feature data; based on the query vector, the key vector and the value vector, the first multi-scale information is enhanced using the object-level feature data to form enhanced scene-level feature data; based on the query vector, the key vector and the value vector, the second multi-scale information is enhanced using the scene-level feature data to form enhanced object-level feature data; the enhanced scene-level feature data and the enhanced object-level feature data are fused by addition to form fused feature data.

[0089] Specifically, in the fine-tuning stage, in order to give full play to the complementary advantages of scene-level and target-level features, the present invention designs a bidirectional cross-attention module for deep feature fusion, and introduces target domain pseudo-labels for pixel-level guidance to improve the final segmentation accuracy.

[0090] In order to efficiently combine the scene-level feature data (scene global features) and object-level feature data (target global features) extracted in the pre-training stage, a bidirectional cross-attention module (BCAM) is designed in the fine-tuning stage. This module uses the cross-attention mechanism to complement and enhance the two features. Specifically, the input scene feature map and target feature map are first processed by convolution layers with different sizes of convolution kernels (3x3 and 5x5) to capture multi-scale information. Then, the respective query (Q), key (K) and value (V) matrices are generated through 1x1 convolution. The calculation process of cross attention is as follows: First, the query (Q) of the scene feature is used to obtain the key (K) and value (V) matrices. s ) and the key of the target feature (K o ) and value (Vo ) Calculate attention and obtain the scene feature A after the target feature is enhanced s :

[0091] ;

[0092] Then, symmetrically, the query (Q o ) and the key of scene features (K s ) and value (V s ) Calculate attention and obtain the target feature A after scene feature enhancement o :

[0093] ;

[0094] Finally, the two enhanced features are fused by adding them together to form the final fusion feature F fused :

[0095] ;

[0096] This bidirectional mechanism ensures the effective interaction and integration of global context and local detail information.

[0097] S104: Adjusting the trained cross-domain building extraction framework model based on the fused feature data to form an adjusted cross-domain building extraction framework model;

[0098] In the specific implementation process of the present invention, the trained cross-domain building extraction framework model is adjusted based on the fused feature data to form an adjusted cross-domain building extraction framework model, including: using the trained cross-domain building extraction framework model to predict and process unlabeled urban and rural remote sensing images to generate target domain pseudo labels; combining the target domain pseudo labels with the real label data of the source domain as supervision signals, and performing pixel-level supervised learning on the fused feature data to adjust the trained cross-domain building extraction framework model.

[0099] Specifically, pixel-level pseudo-labels are used to perform guided fine-tuning; that is, in order to further improve the migration ability and segmentation accuracy of the model in the target domain (such as rural areas), the present invention introduces pixel-level pseudo-labels for guidance after feature fusion; that is, the model trained on the source domain (such as city) data is used to predict the unlabeled images of the target domain to generate pixel-level pseudo-labels; these pseudo-labels provide supervision information of the target domain, and then the supervision information and the fused feature data are used to perform pixel-level adjustment processing on the trained cross-domain building extraction framework model to form an adjusted cross-domain building extraction framework model; it can effectively help the model optimize the fine-grained feature learning in the target domain; finally, these generated target domain pseudo-labels are combined with the real label data of the source domain and used together for the final training of the model, thereby significantly improving the building extraction accuracy of the model in the target domain.

[0100] S105: Obtain a remote sensing image to be identified, and input the remote sensing image to be identified into the adjusted cross-domain building extraction framework model to perform feature extraction processing on cross-domain buildings.

[0101] In the specific implementation process of the present invention, after obtaining the remote sensing image to be identified, the remote sensing image to be identified can be input into the adjusted cross-domain building extraction framework model to perform feature extraction processing of cross-domain buildings; in this way, the cross-domain building features in the remote sensing image to be identified can be extracted more accurately.

[0102] To verify the effectiveness and versatility of this example, we conducted a large number of experiments on multiple public remote sensing datasets and conducted a comprehensive comparison with various existing semantic segmentation models and domain adaptation methods:

[0103] (1) Dataset, LoveDA dataset: This is a high-resolution remote sensing image dataset for domain adaptation semantic segmentation with a resolution of 0.3 meters. The characteristics of this dataset are diverse target scales, complex scenes, and inconsistent feature distributions between urban and rural scenes, which brings challenges to urban and rural knowledge transfer. In the experiment, all categories other than buildings were set as backgrounds, and images containing buildings were screened out. Finally, the urban and rural datasets each contained 650 1024×1024 pixel images, of which 500 were used for training and 150 were used for testing. Inria building dataset: This dataset contains 180 high-resolution images from five cities, with a size of 5000×5000 pixels and a resolution of 0.3 meters. It was also cropped into 650 1024×1024 sub-images for experiments, of which 500 were used for training and 150 were used for testing.

[0104] (2) Evaluation indicators: Intersection over Union (IoU) and F1 score are used to evaluate the performance of the model. IoU is sensitive to the overall segmentation quality, while F1 score is more sensitive to the class imbalance problem. The calculation formula of the relevant indicators is as follows:

[0105]

[0106] Among them, TP, FP, and FN represent true positive, false positive, and false negative, respectively; is the recall rate; For accuracy.

[0107] (3) Comparison method: The method of this embodiment is compared with the following methods, including: classic semantic segmentation models: Pspnet, Psanet, Deeplabv3+, Danet, Ocrnet; unsupervised domain adaptation (UDA) SOTA methods: CBST, PyCDA, Iast, DCA, Weakly, SiamSeg, ST-DASegNet.

[0108] (4) Experimental setup: All experiments were implemented on the mmsegmentation and mmselfsup frameworks and trained on a single NVIDIA RTX 6000 GPU; scene feature extraction: using the LARS optimizer with an initial learning rate of 0.1, a linear warm-up and cosine annealing learning rate scheduling strategy, and training for a total of 200 epochs; target-level feature extraction: when SAM generates masks, the predicted IoU threshold is set to 0.7 and the stability score threshold is set to 0.85 to ensure that only high-confidence masks are retained; segmentation training: using the Adam optimizer with an initial learning rate of 0.00003 and a weight decay of 0.003.

[0109] (5) Experimental results. Compared with existing domain adaptation technologies, we selected Ocrnet with the best performance as the benchmark model and compared the method given in this embodiment (Ocrnet+SOP) with various advanced domain adaptation methods. As shown in Table 1, in the same-source migration task of LoveDA_urban (LoveDA dataset urban area) → LoveDA_rural (LoveDA dataset rural area), the IoU of the proposed method reached 68.26%, an increase of 38.28% compared with the baseline (29.98%) trained only using the source domain, and outperformed all compared domain adaptation methods; in the more challenging different-source migration task of Inria_urban (Inria building dataset urban area) → LoveDA_rural (LoveDA dataset rural area), the source domain difference caused the performance of all methods to decline, but the proposed method still achieved an IoU of 65.01%, an increase of 40.34% compared with the baseline (24.67%), and significantly better than the second-best method (58.75%), fully demonstrating the effectiveness of the proposed method in dealing with huge domain differences; it can more accurately segment buildings and effectively reduce missed detections and false detections in complex backgrounds.

[0110] Table 1

[0111]

[0112] In an embodiment of the present invention, the global semantic information and fine-grained target features of urban and rural buildings are captured, thereby significantly improving the extraction accuracy and robustness of the model in rural areas; by introducing a contrastive learning framework and SegmentAnything Model (SAM), high-quality target-level pseudo labels are generated without target domain annotation, further optimizing the cross-domain adaptability and solving the problem of unstable pseudo-label quality in traditional methods; a bidirectional cross-attention mechanism deeply integrates global scene features with local target features, ensuring the subsequent precise positioning and extraction of building boundaries; thereby achieving higher building extraction accuracy on multiple urban and rural remote sensing image datasets, demonstrating stronger cross-domain adaptability and robustness.

[0113] For example 2, please refer to Figure 2 , Figure 2 4 is a flow chart of a cross-domain building extraction method using multi-level feature fusion in another embodiment of the present invention.

[0114] like Figure 2 As shown, a cross-domain building extraction method based on multi-level feature fusion, the method includes:

[0115] S201: Constructing a cross-domain building extraction framework model, wherein the cross-domain building extraction framework model performs scene-level feature learning processing in a set of unlabeled urban and rural remote sensing images based on contrastive learning;

[0116] S202: Processing the unlabeled urban and rural remote sensing images using the SAM module to generate object-level masks; the cross-domain building extraction framework model uses the object-level masks as supervision signals to perform object-level feature training and learning to form a trained cross-domain building extraction framework model;

[0117] S203: Obtaining scene-level feature data and object-level feature data extracted during training of the cross-domain building extraction framework model;

[0118] S204: After inputting the scene-level feature data and the object-level feature data into the bidirectional cross-attention module, performing convolution processing through a convolution layer with convolution kernels of different sizes in the bidirectional cross-attention module to obtain first multi-scale information corresponding to the scene-level feature data and second multi-scale information corresponding to the object-level feature data;

[0119] S205: performing enhancement processing on the first multi-scale information using the object-level feature data based on the query vector, the key vector, and the value vector to form enhanced scene-level feature data;

[0120] S206: Based on the query vector, the key vector, and the value vector, the second multi-scale information is enhanced using the scene-level feature data to form enhanced object-level feature data;

[0121] S207: fusing the enhanced scene-level feature data and the enhanced object-level feature data by adding them together to form fused feature data;

[0122] S208: Adjusting the trained cross-domain building extraction framework model based on the fused feature data to form an adjusted cross-domain building extraction framework model;

[0123] S209: Obtain a remote sensing image to be identified, and input the remote sensing image to be identified into the adjusted cross-domain building extraction framework model to perform feature extraction processing on cross-domain buildings.

[0124] The specific implementation of the second embodiment can be found in the first embodiment, which will not be described in detail here.

[0125] For example three, please refer to Figure 3 , Figure 3 Schematic diagram of the structure of a cross-domain building extraction device with multi-level feature fusion in an embodiment of the present invention.

[0126] like Figure 3As shown, a cross-domain building extraction device with multi-level feature fusion, the device includes:

[0127] The first learning module 301 is used to construct a cross-domain building extraction framework model, wherein the cross-domain building extraction framework model learns scene-level features in a set of unlabeled urban and rural remote sensing images based on contrastive learning;

[0128] In the specific implementation process of the present invention, the cross-domain building extraction framework model performs scene-level feature learning processing in an unlabeled urban and rural remote sensing image set based on contrastive learning, including: the cross-domain building extraction framework model performs random data enhancement processing on any urban and rural remote sensing image in the input unlabeled urban and rural remote sensing image set to form two first images and second images with similar content but different forms of expression; the first image and the second image are input into a shared encoder for high-dimensional feature extraction processing to obtain a first high-dimensional feature vector corresponding to the first image and a second high-dimensional feature vector corresponding to the second image; the first high-dimensional feature vector and the second high-dimensional feature vector are mapped to a low-dimensional feature space using a multi-layer perception mechanism to obtain a first low-dimensional feature vector corresponding to the first high-dimensional feature vector and a second low-dimensional feature vector corresponding to the second high-dimensional feature vector; based on the InfoNCE contrast loss function, the first low-dimensional feature vector and the second low-dimensional feature vector are used to perform scene-level feature learning processing.

[0129] Furthermore, the formula of the InfoNCE contrast loss function is as follows:

[0130] ;

[0131] in, is the first low-dimensional eigenvector; is the second lowest dimensional eigenvector; It is a hyperparameter used to control the impact of negative samples; is the number of samples in a batch, .

[0132] Specifically, we first need to build a cross-domain building extraction framework model, which can be referenced Figure 5, a cross-domain building extraction framework model is constructed, aiming to optimize the cross-domain adaptability of buildings in remote sensing images by gradually extracting and fusing features at different levels, and effectively solve the problem of model performance degradation caused by the difference in urban and rural building styles. It includes two core stages: pre-training and fine-tuning; first, in the pre-training stage, the model learns common scene-level features from large-scale unlabeled urban and rural remote sensing images through comparative learning; at the same time, SAM is introduced to generate category-independent object-level pseudo labels for all images, guiding the model to grasp the boundary features of the target and background; then, in the fine-tuning stage, a bidirectional cross attention module (BCAM) is designed to deeply fuse scene and target features, and the target domain pseudo labels generated under the guidance of the source domain are used for pixel-level joint training, thereby significantly improving the accuracy of building extraction in the absence of target domain annotations.

[0133] During the pre-training phase, i.e., learning scene-level features, this embodiment uses the SimCLR contrastive learning framework to extract high-quality, generalizable scene-level features from a large number of unlabeled urban and rural remote sensing images to enable the model to understand the global structure and semantic relationships of remote sensing scenes. Specifically, the cross-domain building extraction framework model performs random data augmentation on any urban and rural remote sensing image in the input unlabeled urban and rural remote sensing image set, generating a first image and a second image with similar content but different representations. These first and second images are then fed into a shared encoder (e.g., ResNet-50) for high-dimensional feature extraction, obtaining a first high-dimensional feature vector corresponding to the first image and a second high-dimensional feature vector corresponding to the second image. Subsequently, a lightweight multi-layer perceptron (MLP) maps these high-dimensional vectors into a low-dimensional feature space, obtaining a first low-dimensional feature vector corresponding to the first high-dimensional feature vector and a second low-dimensional feature vector corresponding to the second high-dimensional feature vector. The training goal is to use the InfoNCE contrastive loss function to bring two views of the same scene (positive pairs) closer together in the feature space while simultaneously increasing the distance between two views of different scenes (negative pairs).

[0134] The formula of InfoNCE contrast loss function is as follows:

[0135] ;

[0136] in, is the first low-dimensional eigenvector; is the second lowest dimensional eigenvector; is a hyperparameter used to control the impact of negative samples, in this embodiment, Set to 0.5; is the number of samples in a batch, .

[0137] Through this unsupervised learning, the model is able to capture the intrinsic invariant properties of the scene rather than specific visual details, thereby extracting rich scene-level information covering local texture, overall layout, and environmental characteristics.

[0138] The second learning module 302 is used to process the unlabeled urban and rural remote sensing images using the SAM module to generate object-level masks; the cross-domain building extraction framework model uses the object-level masks as supervision signals to perform object-level feature training and learning processing to form a trained cross-domain building extraction framework model;

[0139] In the specific implementation process of the present invention, the SAM module is used to process the unlabeled urban and rural remote sensing images to generate object-level masks; the cross-domain building extraction framework model uses the object-level masks as supervision signals to perform training and learning processing of object-level features, including: using the SAM module in the cross-domain building extraction framework model to perform target object extraction processing on each urban and rural remote sensing image in the unlabeled urban and rural remote sensing image set under unsupervised conditions to generate a target object mask; performing target and background segmentation processing on the target object mask to form a segmented target and a segmented background; the cross-domain building extraction framework model uses a binary cross entropy loss function to perform object-level feature learning processing on the segmented target and the segmented background.

[0140] Furthermore, the binary cross entropy loss function is as follows:

[0141] ;

[0142] in, The cross-domain building extraction framework model predicts the The probability that a pixel belongs to the target; For the The supervised labels for pixels; is the number of pixels;

[0143] Taking into account the problem of class imbalance, the Focal Loss loss function is added to strengthen the focus on difficult-to-classify samples, as follows:

[0144] ;

[0145] in, The cross-domain building extraction framework model predicts the probability of the target category; is the category balance factor; To adjust the focus parameters of difficult and easy samples; therefore, the total loss function is as follows:

[0146] .

[0147] Specifically, target-level feature extraction is performed based on SAM to generate semantic-free pseudo-labels. Target-level features focus on detailed attributes such as the shape, edge, and texture of specific objects in the image, which are crucial for accurate segmentation. In this embodiment, the Segment Anything Model (SAM) is introduced to generate high-quality, semantic-free object masks (i.e., pseudo-labels) for all remote sensing images in an unsupervised manner. These masks accurately depict the boundaries of all potential targets in the image, serving as a guide for target-level feature learning. SAM processes the original image to generate an initial object mask. In order to improve mask accuracy and optimize the model's target detection capabilities, this problem is converted into a binary segmentation task of target and background. In the target and background segmentation task, a binary cross-entropy loss function is used to classify the target and background, thereby improving the accuracy of the mask and further optimizing the model's target detection capabilities. The binary cross-entropy loss function is as follows:

[0148] The binary cross entropy loss function is as follows:

[0149] ;

[0150] in, The cross-domain building extraction framework model predicts the The probability that a pixel belongs to the target; For the The supervised labels for pixels; is the number of pixels;

[0151] Taking into account the problem of class imbalance, the Focal Loss loss function is added to strengthen the focus on difficult-to-classify samples, as follows:

[0152] ;

[0153] in, The cross-domain building extraction framework model predicts the probability of the target category; is the category balance factor; To adjust the focus parameters of difficult and easy samples; therefore, the total loss function is as follows:

[0154] .

[0155] Feature fusion module 303: used to input the scene-level feature data and object-level feature data extracted during training and learning into the bidirectional cross-attention module of the trained cross-domain building extraction framework model for feature fusion processing to form fused feature data;

[0156] In the specific implementation process of the present invention, the scene-level feature data and object-level feature data extracted during training and learning are input into the bidirectional cross-attention module of the trained cross-domain building extraction framework model for feature fusion processing to form fused feature data, including: obtaining the scene-level feature data and object-level feature data extracted by the cross-domain building extraction framework model during training and learning; after inputting the scene-level feature data and object-level feature data into the bidirectional cross-attention module, convolution processing is performed through a convolution layer of convolution kernels of different sizes in the bidirectional cross-attention module to obtain first multi-scale information corresponding to the scene-level feature data and second multi-scale information corresponding to the object-level feature data; based on the query vector, the key vector and the value vector, the first multi-scale information is enhanced using the object-level feature data to form enhanced scene-level feature data; based on the query vector, the key vector and the value vector, the second multi-scale information is enhanced using the scene-level feature data to form enhanced object-level feature data; the enhanced scene-level feature data and the enhanced object-level feature data are fused by addition to form fused feature data.

[0157] Specifically, in the fine-tuning stage, in order to give full play to the complementary advantages of scene-level and target-level features, the present invention designs a bidirectional cross-attention module for deep feature fusion, and introduces target domain pseudo-labels for pixel-level guidance to improve the final segmentation accuracy.

[0158] In order to efficiently combine the scene-level feature data (scene global features) and object-level feature data (target global features) extracted in the pre-training stage, a bidirectional cross-attention module (BCAM) is designed in the fine-tuning stage. This module uses the cross-attention mechanism to complement and enhance the two features. Specifically, the input scene feature map and target feature map are first processed by convolution layers with different sizes of convolution kernels (3x3 and 5x5) to capture multi-scale information. Then, the respective query (Q), key (K) and value (V) matrices are generated through 1x1 convolution. The calculation process of cross attention is as follows: First, the query (Q) of the scene feature is used to obtain the key (K) and value (V) matrices. s ) and the key of the target feature (K o ) and value (V o ) Calculate attention and obtain the scene feature A after the target feature is enhanced s :

[0159] ;

[0160] Then, symmetrically, the query (Q o ) and the key of scene features (K s ) and value (V s ) Calculate attention and obtain the target feature A after scene feature enhancemento :

[0161] ;

[0162] Finally, the two enhanced features are fused by adding them together to form the final fusion feature F fused :

[0163] ;

[0164] This bidirectional mechanism ensures the effective interaction and integration of global context and local detail information.

[0165] Adjustment module 304: used to adjust the trained cross-domain building extraction framework model based on the fused feature data to form an adjusted cross-domain building extraction framework model;

[0166] In the specific implementation process of the present invention, the trained cross-domain building extraction framework model is adjusted based on the fused feature data to form an adjusted cross-domain building extraction framework model, including: using the trained cross-domain building extraction framework model to predict and process unlabeled urban and rural remote sensing images to generate target domain pseudo labels; combining the target domain pseudo labels with the real label data of the source domain as supervision signals, and performing pixel-level supervised learning on the fused feature data to adjust the trained cross-domain building extraction framework model.

[0167] Specifically, pixel-level pseudo-labels are used to perform guided fine-tuning; that is, in order to further improve the migration ability and segmentation accuracy of the model in the target domain (such as rural areas), the present invention introduces pixel-level pseudo-labels for guidance after feature fusion; that is, the model trained on the source domain (such as city) data is used to predict the unlabeled images of the target domain to generate pixel-level pseudo-labels; these pseudo-labels provide supervision information of the target domain, and then the supervision information and the fused feature data are used to perform pixel-level adjustment processing on the trained cross-domain building extraction framework model to form an adjusted cross-domain building extraction framework model; it can effectively help the model optimize the fine-grained feature learning in the target domain; finally, these generated target domain pseudo-labels are combined with the real label data of the source domain and used together for the final training of the model, thereby significantly improving the building extraction accuracy of the model in the target domain.

[0168] Feature extraction module 305: used to obtain the remote sensing image to be identified, and input the remote sensing image to be identified into the adjusted cross-domain building extraction framework model to perform feature extraction processing on the cross-domain buildings.

[0169] In the specific implementation process of the present invention, after obtaining the remote sensing image to be identified, the remote sensing image to be identified can be input into the adjusted cross-domain building extraction framework model to perform feature extraction processing of cross-domain buildings; in this way, the cross-domain building features in the remote sensing image to be identified can be extracted more accurately.

[0170] To verify the effectiveness and versatility of this example, we conducted a large number of experiments on multiple public remote sensing datasets and conducted a comprehensive comparison with various existing semantic segmentation models and domain adaptation methods:

[0171] (1) Dataset, LoveDA dataset: This is a high-resolution remote sensing image dataset for domain adaptation semantic segmentation with a resolution of 0.3 meters. The characteristics of this dataset are diverse target scales, complex scenes, and inconsistent feature distributions between urban and rural scenes, which brings challenges to urban and rural knowledge transfer. In the experiment, all categories other than buildings were set as backgrounds, and images containing buildings were screened out. Finally, the urban and rural datasets each contained 650 1024×1024 pixel images, of which 500 were used for training and 150 were used for testing. Inria building dataset: This dataset contains 180 high-resolution images from five cities, with a size of 5000×5000 pixels and a resolution of 0.3 meters. It was also cropped into 650 1024×1024 sub-images for experiments, of which 500 were used for training and 150 were used for testing.

[0172] (2) Evaluation indicators: Intersection over Union (IoU) and F1 score are used to evaluate the performance of the model. IoU is sensitive to the overall segmentation quality, while F1 score is more sensitive to the class imbalance problem. The calculation formula of the relevant indicators is as follows:

[0173]

[0174] Among them, TP, FP, and FN represent true positive, false positive, and false negative, respectively; is the recall rate; For accuracy.

[0175] (3) Comparison method: The method of this embodiment is compared with the following methods, including: classic semantic segmentation models: Pspnet, Psanet, Deeplabv3+, Danet, Ocrnet; unsupervised domain adaptation (UDA) SOTA methods: CBST, PyCDA, Iast, DCA, Weakly, SiamSeg, ST-DASegNet.

[0176] (4) Experimental setup: All experiments were implemented on the mmsegmentation and mmselfsup frameworks and trained on a single NVIDIA RTX 6000 GPU; scene feature extraction: using the LARS optimizer with an initial learning rate of 0.1, a linear warm-up and cosine annealing learning rate scheduling strategy, and training for a total of 200 epochs; target-level feature extraction: when SAM generates masks, the predicted IoU threshold is set to 0.7 and the stability score threshold is set to 0.85 to ensure that only high-confidence masks are retained; segmentation training: using the Adam optimizer with an initial learning rate of 0.00003 and a weight decay of 0.003.

[0177] (5) Experimental results. Compared with existing domain adaptation technologies, we selected Ocrnet with the best performance as the benchmark model and compared the method given in this embodiment (Ocrnet+SOP) with various advanced domain adaptation methods. As shown in Table 1, in the same-source migration task of LoveDA_urban (LoveDA dataset urban area) → LoveDA_rural (LoveDA dataset rural area), the IoU of the proposed method reached 68.26%, an increase of 38.28% compared with the baseline (29.98%) trained only using the source domain, and outperformed all compared domain adaptation methods; in the more challenging different-source migration task of Inria_urban (Inria building dataset urban area) → LoveDA_rural (LoveDA dataset rural area), the source domain difference caused the performance of all methods to decline, but the proposed method still achieved an IoU of 65.01%, an increase of 40.34% compared with the baseline (24.67%), and significantly better than the second-best method (58.75%), fully demonstrating the effectiveness of the proposed method in dealing with huge domain differences; it can more accurately segment buildings and effectively reduce missed detections and false detections in complex backgrounds.

[0178] Table 1

[0179]

[0180] In an embodiment of the present invention, the global semantic information and fine-grained target features of urban and rural buildings are captured, thereby significantly improving the extraction accuracy and robustness of the model in rural areas; by introducing a contrastive learning framework and SegmentAnything Model (SAM), high-quality target-level pseudo labels are generated without target domain annotation, further optimizing the cross-domain adaptability and solving the problem of unstable pseudo-label quality in traditional methods; a bidirectional cross-attention mechanism deeply integrates global scene features with local target features, ensuring the subsequent precise positioning and extraction of building boundaries; thereby achieving higher building extraction accuracy on multiple urban and rural remote sensing image datasets, demonstrating stronger cross-domain adaptability and robustness.

[0181] An embodiment of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the cross-domain building extraction method described in any of the above embodiments. The computer-readable storage medium includes, but is not limited to, any type of disk (including floppy disks, hard disks, optical disks, CD-ROMs, and magneto-optical disks), ROM (Read-Only Memory), RAM (Random Access Memory), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory, magnetic cards, or optical cards. In other words, a storage device includes any medium that stores or transmits information in a readable form by a device (e.g., a computer or mobile phone), and may include a read-only memory, a disk, or an optical disk.

[0182] An embodiment of the present invention further provides a computer application program that runs on a computer and is used to execute the cross-domain building extraction method of any one of the above embodiments.

[0183] also, Figure 4 It is a schematic diagram of the structure of an electronic device in an embodiment of the present invention.

[0184] The embodiment of the present invention further provides an electronic device, such as Figure 4 The electronic device includes a processor 402, a memory 403, an input unit 404, a display unit 405 and other components. Those skilled in the art will understand that Figure 4The structural components of the electronic device shown do not constitute a limitation on all devices, and may include more or fewer components than shown, or combine certain components. The memory 403 can be used to store the application 401 and various functional modules, and the processor 402 runs the application 401 stored in the memory 403, thereby executing various functional applications and data processing of the device. The memory can be an internal memory or an external memory, or include both internal and external memories. The internal memory may include a read-only memory (ROM), a programmable ROM (PROM), an electrically programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a flash memory, or a random access memory. The external memory may include a hard disk, a floppy disk, a ZIP disk, a USB flash drive, a magnetic tape, etc. The memory disclosed in the present invention includes but is not limited to these types of memories. The memory disclosed in the present invention is only an example and not a limitation.

[0185] The input unit 404 is used to receive input signals and keywords entered by the user. The input unit 404 may include a touch panel and other input devices. The touch panel can detect user touch operations on or near it (e.g., operations performed on or near the touch panel using a finger, stylus, or any other suitable object or accessory) and activate corresponding connected devices according to pre-set programs. Other input devices may include, but are not limited to, one or more of a physical keyboard, function keys (e.g., playback control keys, on / off buttons, etc.), a trackball, a mouse, and a joystick. The display unit 405 is used to display user input or information provided to the user, as well as various menus of the terminal device. The display unit 405 may be in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or other devices. The processor 402 is the control center of the terminal device. It connects various components of the device using various interfaces and circuits. It performs various functions and processes data by running or executing software programs and / or modules stored in the memory 403 and accessing data stored in the memory.

[0186] As an embodiment, the electronic device includes: one or more processors 402, a memory 403, and one or more applications 401, wherein the one or more applications 401 are stored in the memory 403 and are configured to be executed by the one or more processors 402, and the one or more applications 401 are configured to execute the corresponding cross-domain building extraction method in any one of the above embodiments.

[0187] In an embodiment of the present invention, the global semantic information and fine-grained target features of urban and rural buildings are captured, thereby significantly improving the extraction accuracy and robustness of the model in rural areas; by introducing a contrastive learning framework and SegmentAnything Model (SAM), high-quality target-level pseudo labels are generated without target domain annotation, further optimizing the cross-domain adaptability and solving the problem of unstable pseudo-label quality in traditional methods; a bidirectional cross-attention mechanism deeply integrates global scene features with local target features, ensuring the subsequent precise positioning and extraction of building boundaries; thereby achieving higher building extraction accuracy on multiple urban and rural remote sensing image datasets, demonstrating stronger cross-domain adaptability and robustness.

[0188] In addition, the above is a detailed introduction to a cross-domain building extraction method and related devices with multi-level feature fusion provided by an embodiment of the present invention. Specific examples have been used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those skilled in the art, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A cross-domain building extraction method based on multi-level feature fusion, characterized in that: The method comprises: A cross-domain building extraction framework model is constructed, which processes scene-level features in unlabeled urban and rural remote sensing image sets based on contrastive learning. The SAM module is used to process unlabeled urban and rural remote sensing images to generate object-level masks; the cross-domain building extraction framework model uses the object-level masks as supervision signals to perform object-level feature training and learning to form a trained cross-domain building extraction framework model; The scene-level feature data and object-level feature data extracted during training are input into the bidirectional cross-attention module of the trained cross-domain building extraction framework model for feature fusion processing to form fused feature data; Adjusting the trained cross-domain building extraction framework model based on the fused feature data to form an adjusted cross-domain building extraction framework model; Obtain a remote sensing image to be identified, and input the remote sensing image to be identified into the adjusted cross-domain building extraction framework model to perform feature extraction processing on the cross-domain buildings; The cross-domain building extraction framework model is based on contrastive learning to learn scene-level features in unlabeled urban and rural remote sensing image sets, including: The cross-domain building extraction framework model performs random data enhancement processing on any urban and rural remote sensing image in the input unlabeled urban and rural remote sensing image set to form a first image and a second image with similar content but different forms of expression; Inputting the first image and the second image into a shared encoder for high-dimensional feature extraction processing to obtain a first high-dimensional feature vector corresponding to the first image and a second high-dimensional feature vector corresponding to the second image; Mapping the first high-dimensional feature vector and the second high-dimensional feature vector to a low-dimensional feature space using a multi-layer perception mechanism to obtain a first low-dimensional feature vector corresponding to the first high-dimensional feature vector and a second low-dimensional feature vector corresponding to the second high-dimensional feature vector; Performing scene-level feature learning processing using the first low-dimensional feature vector and the second low-dimensional feature vector based on the InfoNCE contrast loss function; The SAM module is used to process unlabeled urban and rural remote sensing images to generate object-level masks; the cross-domain building extraction framework model uses the object-level masks as supervision signals to perform object-level feature training and learning, including: In the cross-domain building extraction framework model, the SAM module is used to perform target object extraction processing on each urban and rural remote sensing image in the unlabeled urban and rural remote sensing image set under unsupervised conditions to generate a target object mask; Performing target and background segmentation processing on the target object mask to form a segmented target and a segmented background; The cross-domain building extraction framework model uses a binary cross entropy loss function to perform object-level feature learning on the segmentation target and the segmentation background; The step of adjusting the trained cross-domain building extraction framework model based on the fused feature data to form an adjusted cross-domain building extraction framework model includes: The trained cross-domain building extraction framework model is used to predict and process unlabeled urban and rural remote sensing images to generate pseudo labels in the target domain. The target domain pseudo-labels are combined with the real label data of the source domain as supervisory signals, and pixel-level supervised learning is performed on the fused feature data to adjust the trained cross-domain building extraction framework model.

2. The cross-domain building extraction method according to claim 1, characterized in that: The formula of the InfoNCE contrast loss function is as follows: ; in, is the first low-dimensional eigenvector; is the second lowest dimensional eigenvector; It is a hyperparameter used to control the impact of negative samples; is the number of samples in a batch, .

3. The cross-domain building extraction method according to claim 1, characterized in that: The binary cross entropy loss function is as follows: ; in, The cross-domain building extraction framework model predicts the The probability that a pixel belongs to the target; For the The supervised labels for pixels; is the number of pixels; Taking into account the problem of class imbalance, the Focal Loss loss function is added to strengthen the focus on difficult-to-classify samples, as follows: ; in, The cross-domain building extraction framework model predicts the probability of the target category; is the category balance factor; To adjust the focus parameters of difficult and easy samples; therefore, the total loss function is as follows: 。 4. The cross-domain building extraction method according to claim 1, characterized in that: The scene-level feature data and object-level feature data extracted during training are input into the bidirectional cross-attention module of the trained cross-domain building extraction framework model for feature fusion processing to form fused feature data, including: Obtaining scene-level feature data and object-level feature data extracted during training and learning of the cross-domain building extraction framework model; After inputting the scene-level feature data and the object-level feature data into the bidirectional cross-attention module, convolution processing is performed through a convolution layer with convolution kernels of different sizes in the bidirectional cross-attention module to obtain first multi-scale information corresponding to the scene-level feature data and second multi-scale information corresponding to the object-level feature data; performing enhancement processing on the first multi-scale information using object-level feature data based on the query vector, the key vector, and the value vector to form enhanced scene-level feature data; performing enhancement processing on the second multi-scale information using scene-level feature data based on the query vector, the key vector, and the value vector to form enhanced object-level feature data; The enhanced scene-level feature data and the enhanced object-level feature data are fused by adding them together to form fused feature data.

5. A multi-level feature fusion cross-domain building extraction device, the device is applied to the cross-domain building extraction method according to any one of claims 1 to 4, characterized in that: The device comprises: The first learning module is used to build a cross-domain building extraction framework model, which is based on contrastive learning to learn scene-level features in unlabeled urban and rural remote sensing image sets; at the same time, The second learning module is used to process the unlabeled urban and rural remote sensing images using the SAM module to generate object-level masks; the cross-domain building extraction framework model uses the object-level masks as supervision signals to perform object-level feature training and learning processing to form a trained cross-domain building extraction framework model; Feature fusion module: used to input the scene-level feature data and object-level feature data extracted during training into the bidirectional cross-attention module of the trained cross-domain building extraction framework model for feature fusion processing to form fused feature data; Adjustment module: used to adjust the trained cross-domain building extraction framework model based on the fused feature data to form an adjusted cross-domain building extraction framework model; Feature extraction module: used to obtain the remote sensing image to be identified, and input the remote sensing image to be identified into the adjusted cross-domain building extraction framework model to perform feature extraction processing on cross-domain buildings.

6. An electronic device comprising a processor and a memory, characterized in that: The processor runs the computer program or code stored in the memory to implement the cross-domain building extraction method according to any one of claims 1 to 4.

7. A computer-readable storage medium for storing a computer program or code, characterized in that: When the computer program or code is executed by a processor, the cross-domain building extraction method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Remote sensing image building extraction method fusing double-space attention features

    CN120198800A

  • Three-dimensional modeling method and apparatus

    WO2025138753A1