Precise building contour extraction method based on prior information and related device
By using multi-source prior information iterative reconstruction technology, the problem of balancing local details and global structure in building contour extraction of existing models has been solved, achieving high-precision and complete building contour generation that can adapt to complex scenarios.
Patent Information
- Application Number
- CN202511494349.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-20
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-10-20
AI Technical Summary
Existing discriminative models struggle to balance local details with global structure, resulting in geometric distortion of building outlines. The single-step inference process lacks inherent corrective capabilities and is ill-suited to addressing structural defects in initial predictions.
Multi-source prior information is introduced for iterative reconstruction, including local control prior information and global control prior information. Prior feature maps are generated through a pre-trained semantic segmentation network, edge detection algorithm and depth estimation model. Combined with a generative diffusion model for iterative optimization, an accurate building outline mask is formed.
It achieves geometric accuracy and structural integrity of building outlines, improves adaptability in complex scenarios, and generates high-precision outlines that are geometrically reasonable and structurally complete.
Smart Images

Figure CN120953632A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a method and related apparatus for accurate extraction of building outlines based on prior information. Background Technology
[0002] Buildings are important spatial carriers for human habitation and social production. The accurate extraction of their outlines and boundaries has significant application value in fields such as urban planning, urban management, and emergency response. For example, in urban planning, high-precision building outline data can be used for land use analysis, floor area ratio calculation, and 3D urban modeling. In urban management, building boundary information can support the monitoring of illegal buildings and the optimization of infrastructure layout. In emergency response scenarios, rapid and accurate building extraction is crucial for disaster assessment (such as earthquake damage analysis and flood inundation range delineation).
[0003] The rapid development of deep learning technology has provided strong technical support for building extraction tasks. Semantic segmentation of remote sensing images aims to assign a corresponding semantic category label to each pixel in the image. Compared with traditional low-level feature extraction methods, semantic segmentation can directly obtain pixel-level semantic information, providing important support for image-based intelligent analysis. Semantic segmentation methods based on convolutional neural networks (CNNs), such as FCN, Unet, and DeepLabV3, have made significant progress in building extraction tasks. Compared with traditional methods, deep learning models can automatically learn high-dimensional features and achieve efficient pixel-level classification in an end-to-end training framework. Huang et al. proposed an improved DeconvNet, adding upsampling and dense connection operations to the deconvolution layer to improve building extraction performance. Maggiori et al. designed a two-stage network to comprehensively solve the problems of building recognition and accurate localization. Shao et al. proposed the BRRNet network, using residual modules to repair the generated prediction map to improve the final segmentation accuracy. Yang et al. proposed a crack segmentation network based on the DeepLabv3+ architecture, which can improve the detection accuracy of cracks at different scales.
[0004] However, existing discriminative models in the current technology are difficult to balance local details and global structure, resulting in geometric distortion of building outlines; the single-step reasoning process lacks inherent correction capabilities and is difficult to cope with the structural defects of the initial prediction. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of the prior art. This invention provides a method and related device for accurate extraction of building outlines based on prior information, which achieves significant beneficial effects in terms of geometric accuracy, structural integrity and adaptability to complex scenes of building outlines.
[0006] To address the aforementioned technical problems, embodiments of the present invention provide a method for accurate extraction of building outlines based on prior information, the method comprising: Obtain remote sensing image data, perform prior information mining processing on the remote sensing image data for contour extraction, and obtain the contour extraction prior information corresponding to the remote sensing image data, wherein the contour extraction prior information includes local control prior information and global control prior information. The remote sensing image data and the prior information for contour extraction are input into the building contour extraction inference model. The remote sensing image data is then iteratively reconstructed using the prior information for contour extraction as a guiding factor in the building contour extraction inference model. The model outputs the precise contour mask corresponding to the building in the remote sensing image data.
[0007] Optionally, the prior information mining process for contour extraction of the remote sensing image data, to obtain prior information for contour extraction corresponding to the remote sensing image data, includes: Local control prior information mining processing is performed on the remote sensing image data to extract contours, thereby obtaining the local control prior information corresponding to the remote sensing image data. The local control prior information includes structure position perception prior information, boundary detail enhancement prior information, and spatial geometric constraint prior information. Global control prior information mining processing is performed on the remote sensing image data to extract contours, thereby obtaining global control prior information corresponding to the remote sensing image data. The global control prior information includes semantic guidance prior information.
[0008] Optionally, the local control prior information mining process for contour extraction of the remote sensing image data to obtain the local control prior information corresponding to the remote sensing image data includes: The remote sensing image data is input into a pre-trained semantic segmentation network model. The remote sensing image data is initially segmented within the semantic segmentation network model to form a coarse segmentation mask. The coarse segmentation mask is then used as prior information for the structure location awareness class. The remote sensing image data is processed by multiple edge detection algorithms with complementary dimensions to form multiple edge detection maps. These multiple edge detection maps are used as prior information for boundary detail enhancement. The multiple edge detection algorithms with complementary dimensions include the Canny edge detection algorithm, the HED edge response detection algorithm, and the Sobel gradient detection algorithm. The pre-trained monocular depth estimation model performs pixel-by-pixel depth inference on the remote sensing image data, generates a pixel-by-pixel relative depth image, and uses the relative depth image as prior information for spatial geometric constraints.
[0009] Optionally, the global control prior information mining process for contour extraction of the remote sensing image data to obtain the global control prior information corresponding to the remote sensing image data includes: Based on a large-scale vision-language pre-trained model, the text prompts describing the image content in the remote sensing image data are encoded into a high-dimensional semantic embedding vector, and the semantic embedding vector is used as the prior information of the semantic guidance class.
[0010] Optionally, the step of using the prior information for contour extraction as a guiding factor to iteratively reconstruct the remote sensing image data in the building contour extraction inference model, and outputting the precise contour mask corresponding to the building in the remote sensing image data, includes: In the building contour extraction reasoning model, the local control prior information in the contour extraction prior information is used to perform local prior guidance processing to form local prior guidance conditions. In the building contour extraction reasoning model, global prior information from the contour extraction prior information is used for global prior guidance processing to form global prior guidance conditions. Based on the local and global prior guidance conditions, the remote sensing image data is iteratively reconstructed in the building contour extraction inference model to output the precise contour mask corresponding to the building in the remote sensing image data.
[0011] Optionally, the step of using local control prior information from the contour extraction prior information to perform local prior guidance processing in the building contour extraction inference model to form local prior guidance conditions includes: In the building outline extraction inference model, the structural position perception prior information, boundary detail enhancement prior information, and spatial geometric constraint prior information of local control prior information are spliced and fused in the channel dimension to form a spliced and fused multi-channel feature. The spliced and fused multi-channel features are input into a dedicated multi-scale feature extractor within the building contour extraction inference model to generate prior feature maps guided at different resolution levels. The prior feature map is fused with the features of the decoder itself through spatial adaptive normalization within the decoder of the backbone denoising network in the building contour extraction inference model to form local prior guiding conditions. These local prior guiding conditions can dynamically generate scaling parameters and offset parameters.
[0012] Optionally, the iterative reconstruction processing of the remote sensing image data based on the local prior guidance conditions and the global prior guidance conditions in the building contour extraction inference model, and the output of the precise contour mask corresponding to the building in the remote sensing image data, includes: The building contour extraction inference model obtains a coarse segmentation mask corresponding to the remote sensing image data from the input contour extraction prior information, and uses the coarse segmentation mask as the mask in the initial state. ; The local prior guiding conditions and the global prior guiding conditions are used as masks in the initial state. Guiding conditions for iterative reconstruction of the building outline extraction inference model; During iterative reconstruction processing, from the time step Initially, at a time step t, the mask of the building outline extraction inference model in its current state is... Using all guiding conditions as input, predict the mask of the previous time step with less noise. ; Until At that time, the building contour extraction inference model outputs the precise contour mask corresponding to the building in the remote sensing image data.
[0013] In addition, embodiments of the present invention also provide a device for accurate extraction of building outlines based on prior information, the device comprising: Information mining module: used to obtain remote sensing image data, perform prior information mining processing on the remote sensing image data for contour extraction, and obtain the contour extraction prior information corresponding to the remote sensing image data, wherein the contour extraction prior information includes local control prior information and global control prior information; Iterative Reconstruction Module: This module is used to input the remote sensing image data and the prior information for contour extraction into the building contour extraction inference model, and to perform iterative reconstruction processing on the remote sensing image data in the building contour extraction inference model with the prior information for contour extraction as the guiding factor, and output the accurate contour mask corresponding to the building in the remote sensing image data.
[0014] In addition, embodiments of the present invention also provide an electronic device, including a processor and a memory, wherein the processor runs a computer program or code stored in the memory to implement the method for accurate extraction of building outlines as described in any of the above embodiments.
[0015] In addition, embodiments of the present invention also provide a computer-readable storage medium for storing a computer program or code, which, when executed by a processor, implements the method for accurate extraction of building outlines as described above.
[0016] In this embodiment of the invention, to address the problem that existing discriminative models struggle to balance local details with global geometric structure, a technical approach of generative modeling to reconstruct contour optimization tasks is proposed. This approach no longer treats the task as an independent pixel-level classification problem, but transforms it into a conditional generation process guided by multi-source structured prior information such as depth maps and edge detection maps, using an initial coarse segmentation mask as a condition. External structural knowledge about the true boundaries of buildings is introduced into the optimization process, guiding the model to generate geometrically more reasonable and structurally more complete contours in areas where pixel-level information is ambiguous or conflicting, thereby overcoming the geometric distortion problem caused by over-reliance on local pixel features. To address the lack of built-in correction capabilities in existing single-step inference processes, a progressive iterative optimization mechanism based on a diffusion model is introduced, which can gradually refine a flawed initial prediction into a high-precision result with a complete structure and clear boundaries. This achieves significant beneficial effects in terms of geometric accuracy, structural integrity, and adaptability to complex scenes in building contours. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating the method for accurately extracting building outlines based on prior information in an embodiment of the present invention. Figure 2 This is a flowchart illustrating a method for accurately extracting building outlines based on prior information, according to another embodiment of the present invention. Figure 3 This is a schematic diagram of the structural composition of the building outline precision extraction device based on prior information in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structural composition of the electronic device in an embodiment of the present invention; Figure 5 This is a schematic diagram of the overall framework of the method for accurately extracting building outlines based on prior information in an embodiment of the present invention. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] Example 1, please refer to Figure 1 , Figure 1 This is a flowchart illustrating the method for accurately extracting building outlines based on prior information in an embodiment of the present invention.
[0021] like Figure 1 As shown, a method for accurately extracting building outlines based on prior information is provided, the method comprising: S101: Obtain remote sensing image data, perform prior information mining processing on the remote sensing image data for contour extraction, and obtain the contour extraction prior information corresponding to the remote sensing image data, wherein the contour extraction prior information includes local control prior information and global control prior information. In the specific implementation of this invention, the prior information mining process for contour extraction of the remote sensing image data to obtain the contour extraction prior information corresponding to the remote sensing image data includes: local control prior information mining process for contour extraction of the remote sensing image data to obtain the local control prior information corresponding to the remote sensing image data, wherein the local control prior information includes structure location awareness prior information, boundary detail enhancement prior information, and spatial geometric constraint prior information; and global control prior information mining process for contour extraction of the remote sensing image data to obtain the global control prior information corresponding to the remote sensing image data, wherein the global control prior information includes semantic guidance prior information.
[0022] Furthermore, the local control prior information mining process for contour extraction of the remote sensing image data to obtain the local control prior information corresponding to the remote sensing image data includes: inputting the remote sensing image data into a pre-trained semantic segmentation network model, performing preliminary segmentation processing on the remote sensing image data within the semantic segmentation network model to form a coarse segmentation mask, and using the coarse segmentation mask as structure location-aware prior information; performing edge detection processing on the remote sensing image data using multiple edge detection algorithms with complementary dimensions to form multiple edge detection maps, and using the multiple edge detection maps as boundary detail enhancement prior information, wherein the multiple edge detection algorithms with complementary dimensions include the Canny edge detection algorithm, the HED edge response detection algorithm, and the Sobel gradient detection algorithm; performing pixel depth inference on the remote sensing image data based on a pre-trained monocular depth estimation large model to generate a pixel-by-pixel relative depth image, and using the relative depth image as spatial geometric constraint prior information.
[0023] Furthermore, the global control prior information mining process for contour extraction of the remote sensing image data to obtain the global control prior information corresponding to the remote sensing image data includes: encoding the text prompts describing the image content in the remote sensing image data into a high-dimensional semantic embedding vector based on a large-scale vision-language pre-trained model, and using the semantic embedding vector as the semantic guidance class prior information.
[0024] Specifically, the first step is the mining of prior information for building outline extraction. This addresses the problem that relying solely on single image information makes it difficult to reconstruct complete and regular building outlines stably. Therefore, this embodiment introduces a multi-source external prior information guidance system. Its core lies not in simply stacking information, but in systematically classifying and organizing prior information according to function to ensure that various types of information can complement each other, providing comprehensive and accurate guidance for the subsequent diffusion optimization process. The specific prior information mainly includes the following four categories: structural location-aware prior information, boundary detail enhancement prior information, spatial geometric constraint prior information, and semantic guidance prior information. Among these, structural location-aware prior information, boundary detail enhancement prior information, and spatial geometric constraint prior information are defined as local control prior information; semantic guidance prior information is defined as global control prior information.
[0025] The significance of structure-location-aware prior information lies in its provision of a reliable initial state for the contour optimization process. It defines the approximate location, basic shape, and spatial distribution of buildings in the image, significantly reducing the search space for subsequent iterative optimizations and ensuring that the refinement process can efficiently and stably converge to a reasonable result. The extraction method involves using a pre-trained semantic segmentation network (such as DeepLabV3+) to perform preliminary segmentation of the original remote sensing image, generating a coarse segmentation mask. This coarse segmentation mask serves as the initial input for subsequent steps in this embodiment.
[0026] The significance of boundary detail enhancement prior information lies in its core function of introducing high-frequency gradient information from the image to directly address boundary blurring, breakage, and jaggedness issues present in the initial mask. It acts like a detailed "boundary map" for the model, guiding the optimization process to precisely "attach" the contour to locations in the image where pixel values change dramatically, thereby significantly improving the continuity of the contour and the accuracy of geometric details. In terms of extraction methods, three classic edge detection algorithms are used in parallel to extract boundary features from complementary dimensions: Canny edge map: extracting high-contrast, clear edges in the image using the Canny operator; HED edge response map: capturing multi-scale structural edges, including blurred and weakly textured edges, through a global nested edge detection (HED) network; and Sobel gradient map: calculating image gradients using the Sobel operator to enhance linear contours with clear directionality.
[0027] The significance of spatial geometric constraint prior information lies in its aim to introduce three-dimensional geometric information into two-dimensional remote sensing imagery to address the depth blurring problem caused by projection transformation. By providing pixel-level relative height information, it helps the model effectively distinguish between the main body of a building and its projected shadow, or between foreground buildings and background features, providing geometric constraints based on the real physical world for contour generation. The extraction method involves using a pre-trained monocular depth estimation model (such as Depth-Anything) to infer from the original imagery and generate a pixel-by-pixel relative depth map.
[0028] The significance of semantically guided prior information lies in its ability to provide global, target-level semantic confirmation for the entire optimization process at the highest level. In extreme cases such as blurred local features, missing textures, or severe occlusion, it ensures that the model's optimization direction remains anchored to the semantic concept of "building," thereby enhancing the model's robustness in complex scenes and effectively preventing erroneous contour generation. The extraction method involves using a large-scale vision-language pre-trained model (such as CLIP) to encode textual prompts describing image content (e.g., "an aerial photograph containing buildings") into a high-dimensional semantic embedding vector. This vector is subsequently injected as a global condition in each iteration of the building contour extraction inference model.
[0029] S102: Input the remote sensing image data and the prior information for contour extraction into the building contour extraction inference model, and use the prior information for contour extraction as a coordinating guide to perform iterative reconstruction processing on the remote sensing image data in the building contour extraction inference model, and output the accurate contour mask corresponding to the building in the remote sensing image data.
[0030] In a specific implementation of this invention, the step of iteratively reconstructing the remote sensing image data using the prior information of contour extraction as a guiding factor in the building contour extraction inference model, and outputting the precise contour mask corresponding to the building in the remote sensing image data, includes: performing local prior guidance processing using the local control prior information in the prior information of contour extraction in the building contour extraction inference model to form local prior guidance conditions; performing global prior guidance processing using the global control prior information in the prior information of contour extraction in the building contour extraction inference model to form global prior guidance conditions; and performing iterative reconstruction processing of the remote sensing image data in the building contour extraction inference model based on the local prior guidance conditions and the global prior guidance conditions, and outputting the precise contour mask corresponding to the building in the remote sensing image data.
[0031] Furthermore, the local prior guidance processing using the local control prior information in the building contour extraction inference model to form local prior guidance conditions includes: in the building contour extraction inference model, splicing and fusing the structural position-aware prior information, boundary detail enhancement prior information, and spatial geometric constraint prior information of the local control prior information in the channel dimension to form spliced and fused multi-channel features; inputting the spliced and fused multi-channel features into a dedicated multi-scale feature extractor in the building contour extraction inference model to generate prior feature maps guided at different resolution levels; and fusing the prior feature maps with the features of the decoder itself in the decoder of the backbone denoising network in the building contour extraction inference model through spatial adaptive normalization to form local prior guidance conditions, wherein the local prior guidance conditions can dynamically generate scaling parameters and offset parameters.
[0032] Furthermore, the iterative reconstruction processing of the remote sensing image data based on the local and global prior guidance conditions in the building contour extraction inference model, and the output of the precise contour mask corresponding to the building in the remote sensing image data, includes: the building contour extraction inference model obtaining a coarse segmentation mask corresponding to the remote sensing image data from the input contour extraction prior information, and using the coarse segmentation mask as the mask in the initial state. The local prior guidance conditions and the global prior guidance conditions are used as masks in the initial state. The guiding conditions for iterative reconstruction using the building outline extraction inference model; during the iterative reconstruction process, from the time step... Initially, at a time step t, the mask of the building outline extraction inference model in its current state is... Using all guiding conditions as input, predict the mask of the previous time step with less noise. Until At that time, the building contour extraction inference model outputs the precise contour mask corresponding to the building in the remote sensing image data.
[0033] Specifically, the building contour extraction inference model will be based on a generative diffusion model, reconstructing the building contour optimization task into a multi-prior-guided iterative denoising and reconstruction process; the framework of the building contour extraction inference model is based on conditional diffusion-based iterative optimization; the framework treats a coarse segmentation mask with geometric defects as the initial "noisy state," and through a reverse denoising process, at multiple time steps ( The process iterates step by step to reconstruct a geometrically accurate and structurally complete refined mask. In each denoising step, the model's prediction is guided by the synergistic influence of local and global control prior information, thereby ensuring that the optimization process converges in the correct direction.
[0034] Among them, the decoupled multi-prior guidance adopts a decoupled guidance mechanism in order to effectively integrate prior information of different modalities and scales and avoid feature conflicts and interference that may be caused by direct fusion of multiple information. This mechanism draws on the structural ideas of ControlNet, and its components include: (1) a fixed backbone denoising network: the parameters of this network are frozen during the training process, and it is responsible for performing the core denoising and reconstruction task. Its stable structure ensures the basic capabilities of the generation process; (2) multiple parallel trainable control modules: each control module is responsible for receiving one or a class of prior information. Information (such as edge maps, depth maps, etc.); These modules are trainable, and their task is to encode the input prior information into control signals that can guide the backbone network; (3) Zero convolution connector: The control signals output by the control module are injected into the corresponding layer of the backbone denoising network through the zero convolution layer; This mechanism ensures that in the early stage of training, the prior information will not interfere with the backbone network, and its guidance strength will be adaptively learned as training progresses; The decoupled architecture ensures that the generation capability of the backbone network is not destroyed, and at the same time, the influence of each prior information can be learned and controlled independently and stably.
[0035] The implementation of local prior guidance, for local control prior information describing spatial details (including coarse segmentation mask, three edge maps, and depth map), the guidance process is mainly achieved through the following three steps: (1) the above local prior information is spliced and fused in the channel dimension; (2) the spliced multi-channel features are input into a dedicated multi-scale feature extractor to generate feature maps that can be guided at different resolution levels; (3) in the decoder part of the backbone denoising network, the above-extracted prior features are fused with the features of the decoder itself through a spatially adaptive denormalization (FDN) module; this module can dynamically generate scaling and shifting parameters according to the spatial distribution of the prior feature map, thereby performing pixel-by-pixel fine modulation of the decoder features to achieve accurate local contour guidance.
[0036] The implementation of global prior guidance is achieved through a cross-attention mechanism for high-level global control prior information. The extracted text semantic embedding vectors are injected into the cross-attention layer of the backbone denoising network. This mechanism enables the model to perceive the global semantic context throughout the denoising process, ensuring that the generated contours always conform to the high-level semantic concept of "building".
[0037] Inference and Reconstruction Workflow: The building contour extraction inference model obtains a coarse segmentation mask corresponding to the remote sensing image data from the input contour extraction prior information, and uses the coarse segmentation mask as the initial mask. Then, the Canny edge map, HED edge map, Sobel gradient map, depth map, and text semantic embedding vector are used as guiding conditions; the guiding conditions include local prior guiding conditions and global prior guiding conditions, that is, the local prior guiding conditions and global prior guiding conditions are used as masks in the initial state. Guiding conditions for iterative reconstruction using a building outline extraction inference model; during the iterative reconstruction process, starting from the time step... Initially, at a time step t, the mask of the building outline extraction inference model in its current state is... Using all guiding conditions as input, predict the mask of the previous time step with less noise. Repeat the iterative reconstruction until... At that time, the building outline extraction inference model outputs the precise outline mask corresponding to the building in the remote sensing image data.
[0038] For details, please refer to... Figure 5 The accurate extraction of building outlines mainly consists of two parts. The first part is a prior information mining mechanism for building outline extraction, which aims to automatically extract a variety of prior information with clear structural semantics from the input remote sensing image to enhance the model's ability to perceive and express building outline features. The second part is a multi-prior collaborative guided building outline extraction inference model, which takes the initial segmentation mask as input and gradually optimizes the spatial morphology of the mask through multiple rounds of diffusion-style reverse denoising iteration to achieve fine reconstruction of building outlines.
[0039] In this embodiment of the invention, to address the problem that existing discriminative models struggle to balance local details with global geometric structure, a technical approach of generative modeling to reconstruct contour optimization tasks is proposed. This approach no longer treats the task as an independent pixel-level classification problem, but transforms it into a conditional generation process guided by multi-source structured prior information such as depth maps and edge detection maps, using an initial coarse segmentation mask as a condition. External structural knowledge about the true boundaries of buildings is introduced into the optimization process, guiding the model to generate geometrically more reasonable and structurally more complete contours in areas where pixel-level information is ambiguous or conflicting, thereby overcoming the geometric distortion problem caused by over-reliance on local pixel features. To address the lack of built-in correction capabilities in existing single-step inference processes, a progressive iterative optimization mechanism based on a diffusion model is introduced, which can gradually refine a flawed initial prediction into a high-precision result with a complete structure and clear boundaries. This achieves significant beneficial effects in terms of geometric accuracy, structural integrity, and adaptability to complex scenes in building contours.
[0040] Example 2, please refer to Figure 2 , Figure 2 This is a flowchart illustrating a method for accurately extracting building outlines based on prior information, according to another embodiment of the present invention.
[0041] like Figure 2 As shown, a method for accurately extracting building outlines based on prior information is provided, the method comprising: S201: Obtain remote sensing image data, perform prior information mining processing on the remote sensing image data for contour extraction, and obtain the contour extraction prior information corresponding to the remote sensing image data, wherein the contour extraction prior information includes local control prior information and global control prior information. S202: In the building outline extraction reasoning model, the structural position perception prior information, boundary detail enhancement prior information, and spatial geometric constraint prior information of local control prior information are spliced and fused in the channel dimension to form a spliced and fused multi-channel feature. S203: Input the spliced and fused multi-channel features into a dedicated multi-scale feature extractor within the building contour extraction inference model to generate prior feature maps guided at different resolution levels; S204 uses spatial adaptive normalization to fuse the prior feature map with the features of the decoder itself in the decoder of the backbone denoising network within the building contour extraction inference model, forming a local prior guiding condition. The local prior guiding condition can dynamically generate scaling parameters and offset parameters.
[0042] For details on the specific implementation of Example 2, please refer to the above examples, which will not be repeated here.
[0043] Example 3, please refer to Figure 3 , Figure 3 This is a schematic diagram of the structural composition of the building outline precision extraction device based on prior information in an embodiment of the present invention.
[0044] like Figure 3 As shown, a device for accurately extracting building outlines based on prior information is provided, the device comprising: Information mining module 301: used to obtain remote sensing image data, perform prior information mining processing on the remote sensing image data for contour extraction, and obtain the contour extraction prior information corresponding to the remote sensing image data, wherein the contour extraction prior information includes local control prior information and global control prior information. In the specific implementation of this invention, the prior information mining process for contour extraction of the remote sensing image data to obtain the contour extraction prior information corresponding to the remote sensing image data includes: local control prior information mining process for contour extraction of the remote sensing image data to obtain the local control prior information corresponding to the remote sensing image data, wherein the local control prior information includes structure location awareness prior information, boundary detail enhancement prior information, and spatial geometric constraint prior information; and global control prior information mining process for contour extraction of the remote sensing image data to obtain the global control prior information corresponding to the remote sensing image data, wherein the global control prior information includes semantic guidance prior information.
[0045] Furthermore, the local control prior information mining process for contour extraction of the remote sensing image data to obtain the local control prior information corresponding to the remote sensing image data includes: inputting the remote sensing image data into a pre-trained semantic segmentation network model, performing preliminary segmentation processing on the remote sensing image data within the semantic segmentation network model to form a coarse segmentation mask, and using the coarse segmentation mask as structure location-aware prior information; performing edge detection processing on the remote sensing image data using multiple edge detection algorithms with complementary dimensions to form multiple edge detection maps, and using the multiple edge detection maps as boundary detail enhancement prior information, wherein the multiple edge detection algorithms with complementary dimensions include the Canny edge detection algorithm, the HED edge response detection algorithm, and the Sobel gradient detection algorithm; performing pixel depth inference on the remote sensing image data based on a pre-trained monocular depth estimation large model to generate a pixel-by-pixel relative depth image, and using the relative depth image as spatial geometric constraint prior information.
[0046] Furthermore, the global control prior information mining process for contour extraction of the remote sensing image data to obtain the global control prior information corresponding to the remote sensing image data includes: encoding the text prompts describing the image content in the remote sensing image data into a high-dimensional semantic embedding vector based on a large-scale vision-language pre-trained model, and using the semantic embedding vector as the semantic guidance class prior information.
[0047] Specifically, the first step is the mining of prior information for building outline extraction. This addresses the problem that relying solely on single image information makes it difficult to reconstruct complete and regular building outlines stably. Therefore, this embodiment introduces a multi-source external prior information guidance system. Its core lies not in simply stacking information, but in systematically classifying and organizing prior information according to function to ensure that various types of information can complement each other, providing comprehensive and accurate guidance for the subsequent diffusion optimization process. The specific prior information mainly includes the following four categories: structural location-aware prior information, boundary detail enhancement prior information, spatial geometric constraint prior information, and semantic guidance prior information. Among these, structural location-aware prior information, boundary detail enhancement prior information, and spatial geometric constraint prior information are defined as local control prior information; semantic guidance prior information is defined as global control prior information.
[0048] The significance of structure-location-aware prior information lies in its provision of a reliable initial state for the contour optimization process. It defines the approximate location, basic shape, and spatial distribution of buildings in the image, significantly reducing the search space for subsequent iterative optimizations and ensuring that the refinement process can efficiently and stably converge to a reasonable result. The extraction method involves using a pre-trained semantic segmentation network (such as DeepLabV3+) to perform preliminary segmentation of the original remote sensing image, generating a coarse segmentation mask. This coarse segmentation mask serves as the initial input for subsequent steps in this embodiment.
[0049] The significance of boundary detail enhancement prior information lies in its core function of introducing high-frequency gradient information from the image to directly address boundary blurring, breakage, and jaggedness issues present in the initial mask. It acts like a detailed "boundary map" for the model, guiding the optimization process to precisely "attach" the contour to locations in the image where pixel values change dramatically, thereby significantly improving the continuity of the contour and the accuracy of geometric details. In terms of extraction methods, three classic edge detection algorithms are used in parallel to extract boundary features from complementary dimensions: Canny edge map: extracting high-contrast, clear edges in the image using the Canny operator; HED edge response map: capturing multi-scale structural edges, including blurred and weakly textured edges, through a global nested edge detection (HED) network; and Sobel gradient map: calculating image gradients using the Sobel operator to enhance linear contours with clear directionality.
[0050] The significance of spatial geometric constraint prior information lies in its aim to introduce three-dimensional geometric information into two-dimensional remote sensing imagery to address the depth blurring problem caused by projection transformation. By providing pixel-level relative height information, it helps the model effectively distinguish between the main body of a building and its projected shadow, or between foreground buildings and background features, providing geometric constraints based on the real physical world for contour generation. The extraction method involves using a pre-trained monocular depth estimation model (such as Depth-Anything) to infer from the original imagery and generate a pixel-by-pixel relative depth map.
[0051] The significance of semantically guided prior information lies in its ability to provide global, target-level semantic confirmation for the entire optimization process at the highest level. In extreme cases such as blurred local features, missing textures, or severe occlusion, it ensures that the model's optimization direction remains anchored to the semantic concept of "building," thereby enhancing the model's robustness in complex scenes and effectively preventing erroneous contour generation. The extraction method involves using a large-scale vision-language pre-trained model (such as CLIP) to encode textual prompts describing image content (e.g., "an aerial photograph containing buildings") into a high-dimensional semantic embedding vector. This vector is subsequently injected as a global condition in each iteration of the building contour extraction inference model.
[0052] Iterative Reconstruction Module 302: is used to input the remote sensing image data and the prior information for contour extraction into the building contour extraction inference model, and to perform iterative reconstruction processing on the remote sensing image data in the building contour extraction inference model with the prior information for contour extraction as the guiding factor, and output the accurate contour mask corresponding to the building in the remote sensing image data.
[0053] In a specific implementation of this invention, the step of iteratively reconstructing the remote sensing image data using the prior information of contour extraction as a guiding factor in the building contour extraction inference model, and outputting the precise contour mask corresponding to the building in the remote sensing image data, includes: performing local prior guidance processing using the local control prior information in the prior information of contour extraction in the building contour extraction inference model to form local prior guidance conditions; performing global prior guidance processing using the global control prior information in the prior information of contour extraction in the building contour extraction inference model to form global prior guidance conditions; and performing iterative reconstruction processing of the remote sensing image data in the building contour extraction inference model based on the local prior guidance conditions and the global prior guidance conditions, and outputting the precise contour mask corresponding to the building in the remote sensing image data.
[0054] Furthermore, the local prior guidance processing using the local control prior information in the building contour extraction inference model to form local prior guidance conditions includes: in the building contour extraction inference model, splicing and fusing the structural position-aware prior information, boundary detail enhancement prior information, and spatial geometric constraint prior information of the local control prior information in the channel dimension to form spliced and fused multi-channel features; inputting the spliced and fused multi-channel features into a dedicated multi-scale feature extractor in the building contour extraction inference model to generate prior feature maps guided at different resolution levels; and fusing the prior feature maps with the features of the decoder itself in the decoder of the backbone denoising network in the building contour extraction inference model through spatial adaptive normalization to form local prior guidance conditions, wherein the local prior guidance conditions can dynamically generate scaling parameters and offset parameters.
[0055] Furthermore, the iterative reconstruction processing of the remote sensing image data based on the local and global prior guidance conditions in the building contour extraction inference model, and the output of the precise contour mask corresponding to the building in the remote sensing image data, includes: the building contour extraction inference model obtaining a coarse segmentation mask corresponding to the remote sensing image data from the input contour extraction prior information, and using the coarse segmentation mask as the mask in the initial state. The local prior guidance conditions and the global prior guidance conditions are used as masks in the initial state. The guiding conditions for iterative reconstruction using the building outline extraction inference model; during the iterative reconstruction process, from the time step... Initially, at a time step t, the mask of the building outline extraction inference model in its current state is... Using all guiding conditions as input, predict the mask of the previous time step with less noise. Until At that time, the building contour extraction inference model outputs the precise contour mask corresponding to the building in the remote sensing image data.
[0056] Specifically, the building contour extraction inference model will be based on a generative diffusion model, reconstructing the building contour optimization task into a multi-prior-guided iterative denoising and reconstruction process; the framework of the building contour extraction inference model is based on conditional diffusion-based iterative optimization; the framework treats a coarse segmentation mask with geometric defects as the initial "noisy state," and through a reverse denoising process, at multiple time steps ( The process iterates step by step to reconstruct a geometrically accurate and structurally complete refined mask. In each denoising step, the model's prediction is guided by the synergistic influence of local and global control prior information, thereby ensuring that the optimization process converges in the correct direction.
[0057] Among them, the decoupled multi-prior guidance adopts a decoupled guidance mechanism in order to effectively integrate prior information of different modalities and scales and avoid feature conflicts and interference that may be caused by direct fusion of multiple information. This mechanism draws on the structural ideas of ControlNet, and its components include: (1) a fixed backbone denoising network: the parameters of this network are frozen during the training process, and it is responsible for performing the core denoising and reconstruction task. Its stable structure ensures the basic capabilities of the generation process; (2) multiple parallel trainable control modules: each control module is responsible for receiving one or a class of prior information. Information (such as edge maps, depth maps, etc.); These modules are trainable, and their task is to encode the input prior information into control signals that can guide the backbone network; (3) Zero convolution connector: The control signals output by the control module are injected into the corresponding layer of the backbone denoising network through the zero convolution layer; This mechanism ensures that in the early stage of training, the prior information will not interfere with the backbone network, and its guidance strength will be adaptively learned as training progresses; The decoupled architecture ensures that the generation capability of the backbone network is not destroyed, and at the same time, the influence of each prior information can be learned and controlled independently and stably.
[0058] The implementation of local prior guidance, for local control prior information describing spatial details (including coarse segmentation mask, three edge maps, and depth map), the guidance process is mainly achieved through the following three steps: (1) the above local prior information is spliced and fused in the channel dimension; (2) the spliced multi-channel features are input into a dedicated multi-scale feature extractor to generate feature maps that can be guided at different resolution levels; (3) in the decoder part of the backbone denoising network, the above-extracted prior features are fused with the features of the decoder itself through a spatially adaptive denormalization (FDN) module; this module can dynamically generate scaling and shifting parameters according to the spatial distribution of the prior feature map, thereby performing pixel-by-pixel fine modulation of the decoder features to achieve accurate local contour guidance.
[0059] The implementation of global prior guidance is achieved through a cross-attention mechanism for high-level global control prior information. The extracted text semantic embedding vectors are injected into the cross-attention layer of the backbone denoising network. This mechanism enables the model to perceive the global semantic context throughout the denoising process, ensuring that the generated contours always conform to the high-level semantic concept of "building".
[0060] Inference and Reconstruction Workflow: The building contour extraction inference model obtains a coarse segmentation mask corresponding to the remote sensing image data from the input contour extraction prior information, and uses the coarse segmentation mask as the initial mask. Then, the Canny edge map, HED edge map, Sobel gradient map, depth map, and text semantic embedding vector are used as guiding conditions; the guiding conditions include local prior guiding conditions and global prior guiding conditions, that is, the local prior guiding conditions and global prior guiding conditions are used as masks in the initial state. Guiding conditions for iterative reconstruction using a building outline extraction inference model; during the iterative reconstruction process, starting from the time step... Initially, at a time step t, the mask of the building outline extraction inference model in its current state is... Using all guiding conditions as input, predict the mask of the previous time step with less noise. Repeat the iterative reconstruction until... At that time, the building outline extraction inference model outputs the precise outline mask corresponding to the building in the remote sensing image data.
[0061] For details, please refer to... Figure 5 The accurate extraction of building outlines mainly consists of two parts. The first part is a prior information mining mechanism for building outline extraction, which aims to automatically extract a variety of prior information with clear structural semantics from the input remote sensing image to enhance the model's ability to perceive and express building outline features. The second part is a multi-prior collaborative guided building outline extraction inference model, which takes the initial segmentation mask as input and gradually optimizes the spatial morphology of the mask through multiple rounds of diffusion-style reverse denoising iteration to achieve fine reconstruction of building outlines.
[0062] In this embodiment of the invention, to address the problem that existing discriminative models struggle to balance local details with global geometric structure, a technical approach of generative modeling to reconstruct contour optimization tasks is proposed. This approach no longer treats the task as an independent pixel-level classification problem, but transforms it into a conditional generation process guided by multi-source structured prior information such as depth maps and edge detection maps, using an initial coarse segmentation mask as a condition. External structural knowledge about the true boundaries of buildings is introduced into the optimization process, guiding the model to generate geometrically more reasonable and structurally more complete contours in areas where pixel-level information is ambiguous or conflicting, thereby overcoming the geometric distortion problem caused by over-reliance on local pixel features. To address the lack of built-in correction capabilities in existing single-step inference processes, a progressive iterative optimization mechanism based on a diffusion model is introduced, which can gradually refine a flawed initial prediction into a high-precision result with a complete structure and clear boundaries. This achieves significant beneficial effects in terms of geometric accuracy, structural integrity, and adaptability to complex scenes in building contours.
[0063] In this embodiment, the corresponding experimental scheme and result analysis are given as follows: Experimental Setup: To test the effectiveness and generalization ability of the above embodiments, systematic experiments were conducted on three representative publicly available remote sensing building datasets (WHU, INRIA, and Massachusetts). Several mainstream discriminative semantic segmentation models were selected as comparison methods, including the classic FCN, Unet, DeepLabV3+, and advanced Transformer-based SegFormer and hybrid architecture CBRNet. The experiments used the F1 coefficient (…). The pixel accuracy and geometric structure index (SSIM) are used to comprehensively evaluate the results from two dimensions: pixel accuracy and geometric structure, respectively. The details of the three experimental data are as follows: (1) The satellite dataset II (East Asia) provided by the WHU Buildings dataset is used to check the performance of the network architecture proposed in the above embodiments; the satellite dataset II covers an area of 550 square kilometers in East Asia, with a spatial resolution of 2.7m, and contains a total of 17,388 patches of 512×512 pixels; (2) The Inria (Inria Aerial Image Labeling Dataset) dataset contains multiple remote sensing aerial images that capture diverse land features in urban and rural areas. It consists of 180 training sets and 180 test sets, with a resolution of 0.3m and a single sample size of 5000×5000 pixels; (3) The Massachusetts dataset consists of 151 aerial images of the Boston area, each image is 1500×1500 pixels, with a spatial resolution of 1m. The entire dataset covers an area of approximately 340 square kilometers and includes various scenes such as urban areas, suburbs, and rural areas.
[0064] For the multi-class annotation information in the three datasets, non-building areas were uniformly classified as background, and all images and their corresponding labels were cropped into 512×512 pixel patches. Finally, various prior information samples were generated based on remote sensing images. In order to effectively learn the distribution characteristics of buildings with different shapes and areas, 80% of the samples were used for training, and the remaining 20% were used as the test set to ensure the diversity and representativeness of the training and test data.
[0065] Quantitative evaluation: The quantitative evaluation results of the methods used in the above embodiments and all comparative methods are shown in Table 1. Through the analysis of these results, the following two key conclusions can be drawn: Table 1. Quantitative evaluation results of various methods on three building recognition datasets (the best results are bolded, and the second-best results are underlined).
[0066] First, the method described in the above embodiments consistently and significantly outperforms all existing discriminative models in both pixel-level accuracy (F1) and structural similarity (SSIM) dimensions, establishing a new technological advantage. This advantage remains robust on datasets with varying resolutions and scene complexities. For example, even on the highly challenging high-resolution building extraction dataset INRIA, the method described in the above embodiments achieves an F1 score of 91.07%, representing a significant improvement of 6.81% compared to the second-best SegFormer (84.26%). Simultaneously, the SSIM metric also improves from 79.04% to 86.00%. On the other two datasets, the method described in the above embodiments also achieved the best F1 and SSIM scores. This result strongly demonstrates that the core technology—reconstructing contour optimization from "single-step discrimination" to "iterative generation"—is successful. The contrasting method, due to its inherent "pattern recognition" mechanism, sacrifices the ability to perceive high-frequency boundary information in pursuit of generalization. In contrast, the method described in the above embodiments, through progressive denoising and multiple prior guidance, can actively repair and reconstruct these fine geometric structures, thereby achieving a significant and simultaneous improvement in both evaluation dimensions.
[0067] Second, the methods described in the above embodiments demonstrate excellent generalization ability and stability on diverse datasets. From medium-resolution WHU satellite imagery to high-resolution, complex INRIA aerial imagery, and even the small and densely built Massachusetts dataset, the methods described in the above embodiments consistently perform best. In particular, on the Massachusetts dataset, traditional methods generally achieve F1 scores below 81% due to issues such as small targets, shadows, and occlusion, while the methods described in the above embodiments still achieve 85.68%. This fully demonstrates that the multi-source prior collaborative guidance system designed in the above embodiments is not an artificial adjustment for specific data, but a universal and robust technical framework. By fusing complementary information such as edges, depth, and semantics, it effectively overcomes the limitations of a single image information source and maintains a high level of performance in various complex remote sensing scenarios.
[0068] In summary, comprehensive quantitative experimental results confirm that the technical solution proposed in the above embodiments has established a new and leading technical standard for the accurate extraction of building outlines from remote sensing images.
[0069] This invention provides a computer-readable storage medium storing a computer program. When executed by a processor, this program implements the method for accurately extracting building outlines according to any of the above embodiments. The computer-readable storage medium includes, but is not limited to, any type of disk (including floppy disks, hard disks, optical disks, CD-ROMs, and magneto-optical disks), ROM (Read-Only Memory), RAM (Random Access Memory), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory, magnetic cards, or optical cards. In other words, the storage device includes any medium that stores or transmits information in a readable form by a device (e.g., a computer, a mobile phone), and can be a read-only memory, a disk, or an optical disk, etc.
[0070] This invention also provides a computer application running on a computer, which is used to execute the building outline accurate extraction method of any of the above embodiments.
[0071] also, Figure 4 This is a schematic diagram of the structural composition of the electronic device in an embodiment of the present invention.
[0072] This invention also provides an electronic device, such as... Figure 4 As shown. The electronic device includes components such as a processor 402, a memory 403, an input unit 404, and a display unit 405. Those skilled in the art will understand that... Figure 4 The structural components of the illustrated electronic device do not constitute a limitation on all devices and may include more or fewer components than illustrated, or combine certain components. Memory 403 can be used to store application program 401 and various functional modules. Processor 402 runs application program 401 stored in memory 403, thereby performing various functional applications and data processing of the device. Memory can be internal memory or external memory, or both. Internal memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, or random access memory. External memory may include hard disks, floppy disks, ZIP disks, USB flash drives, magnetic tapes, etc. The memory disclosed in this invention includes, but is not limited to, these types of memory. The memory disclosed in this invention is only an example and not a limitation.
[0073] Input unit 404 is used to receive signal input and user-input keywords. Input unit 404 may include a touch panel and other input devices. The touch panel can collect user touch operations on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel) and drive the corresponding connection device according to a pre-set program; other input devices may include, but are not limited to, one or more of physical keyboards, function keys (such as play control buttons, power buttons, etc.), trackballs, mice, joysticks, etc. Display unit 405 can be used to display user-input information or information provided to the user, as well as various menus of the terminal device. Display unit 405 may be in the form of a liquid crystal display, organic light-emitting diode, etc. Processor 402 is the control center of the terminal device, connecting various parts of the entire device through various interfaces and lines, performing various functions and processing data by running or executing software programs and / or modules stored in memory 403, and calling data stored in memory.
[0074] As one embodiment, the electronic device includes: one or more processors 402, a memory 403, and one or more application programs 401, wherein the one or more application programs 401 are stored in the memory 403 and configured to be executed by the one or more processors 402, and the one or more application programs 401 are configured to perform the building outline accurate extraction method corresponding to any of the above embodiments.
[0075] In this embodiment of the invention, to address the problem that existing discriminative models struggle to balance local details with global geometric structure, a technical approach of generative modeling to reconstruct contour optimization tasks is proposed. This approach no longer treats the task as an independent pixel-level classification problem, but transforms it into a conditional generation process guided by multi-source structured prior information such as depth maps and edge detection maps, using an initial coarse segmentation mask as a condition. External structural knowledge about the true boundaries of buildings is introduced into the optimization process, guiding the model to generate geometrically more reasonable and structurally more complete contours in areas where pixel-level information is ambiguous or conflicting, thereby overcoming the geometric distortion problem caused by over-reliance on local pixel features. To address the lack of built-in correction capabilities in existing single-step inference processes, a progressive iterative optimization mechanism based on a diffusion model is introduced, which can gradually refine a flawed initial prediction into a high-precision result with a complete structure and clear boundaries. This achieves significant beneficial effects in terms of geometric accuracy, structural integrity, and adaptability to complex scenes in building contours.
[0076] Furthermore, the above provides a detailed description of a method and related apparatus for accurate extraction of building outlines based on prior information provided by the embodiments of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for accurately extracting building outlines based on prior information, characterized in that, The method includes: Obtain remote sensing image data, perform prior information mining processing on the remote sensing image data for contour extraction, and obtain the contour extraction prior information corresponding to the remote sensing image data, wherein the contour extraction prior information includes local control prior information and global control prior information. The remote sensing image data and the prior information for contour extraction are input into the building contour extraction inference model. The remote sensing image data is then iteratively reconstructed using the prior information for contour extraction as a guiding factor in the building contour extraction inference model. The model outputs the precise contour mask corresponding to the building in the remote sensing image data.
2. The method for accurately extracting building outlines according to claim 1, characterized in that, The prior information mining process for contour extraction of the remote sensing image data, to obtain prior information for contour extraction corresponding to the remote sensing image data, includes: Local control prior information mining processing is performed on the remote sensing image data to extract contours, thereby obtaining the local control prior information corresponding to the remote sensing image data. The local control prior information includes structure position perception prior information, boundary detail enhancement prior information, and spatial geometric constraint prior information. Global control prior information mining processing is performed on the remote sensing image data to extract contours, thereby obtaining global control prior information corresponding to the remote sensing image data. The global control prior information includes semantic guidance prior information.
3. The method for accurately extracting building outlines according to claim 2, characterized in that, The local control prior information mining process for contour extraction of the remote sensing image data, to obtain the local control prior information corresponding to the remote sensing image data, includes: The remote sensing image data is input into a pre-trained semantic segmentation network model. The remote sensing image data is initially segmented within the semantic segmentation network model to form a coarse segmentation mask. The coarse segmentation mask is then used as prior information for the structure location awareness class. The remote sensing image data is processed by multiple edge detection algorithms with complementary dimensions to form multiple edge detection maps. These multiple edge detection maps are used as prior information for boundary detail enhancement. The multiple edge detection algorithms with complementary dimensions include the Canny edge detection algorithm, the HED edge response detection algorithm, and the Sobel gradient detection algorithm. The pre-trained monocular depth estimation model performs pixel-by-pixel depth inference on the remote sensing image data, generates a pixel-by-pixel relative depth image, and uses the relative depth image as prior information for spatial geometric constraints.
4. The method for accurately extracting building outlines according to claim 2, characterized in that, The global control prior information mining process for contour extraction of the remote sensing image data, to obtain the global control prior information corresponding to the remote sensing image data, includes: Based on a large-scale vision-language pre-trained model, the text prompts describing the image content in the remote sensing image data are encoded into a high-dimensional semantic embedding vector, and the semantic embedding vector is used as the prior information of the semantic guidance class.
5. The method for accurately extracting building outlines according to claim 1, characterized in that, The step of iteratively reconstructing the remote sensing image data using the prior information of the contour extraction in the building contour extraction inference model, and outputting the precise contour mask corresponding to the building in the remote sensing image data, includes: In the building contour extraction reasoning model, the local control prior information in the contour extraction prior information is used to perform local prior guidance processing to form local prior guidance conditions. In the building contour extraction reasoning model, global prior information from the contour extraction prior information is used for global prior guidance processing to form global prior guidance conditions. Based on the local and global prior guidance conditions, the remote sensing image data is iteratively reconstructed in the building contour extraction inference model to output the precise contour mask corresponding to the building in the remote sensing image data.
6. The method for accurately extracting building outlines according to claim 5, characterized in that, The local prior information in the contour extraction prior information of the building contour extraction inference model is used for local prior guidance processing to form local prior guidance conditions, including: In the building outline extraction inference model, the structural position perception prior information, boundary detail enhancement prior information, and spatial geometric constraint prior information of local control prior information are spliced and fused in the channel dimension to form a spliced and fused multi-channel feature. The spliced and fused multi-channel features are input into a dedicated multi-scale feature extractor within the building contour extraction inference model to generate prior feature maps guided at different resolution levels. The prior feature map is fused with the features of the decoder itself through spatial adaptive normalization within the decoder of the backbone denoising network in the building contour extraction inference model to form local prior guiding conditions. These local prior guiding conditions can dynamically generate scaling parameters and offset parameters.
7. The method for accurately extracting building outlines according to claim 5, characterized in that, The iterative reconstruction processing of the remote sensing image data based on the local and global prior guidance conditions in the building contour extraction inference model, outputting the precise contour mask corresponding to the building in the remote sensing image data, includes: The building contour extraction inference model obtains a coarse segmentation mask corresponding to the remote sensing image data from the input contour extraction prior information, and uses the coarse segmentation mask as the mask in the initial state. ; The local prior guiding conditions and the global prior guiding conditions are used as masks in the initial state. Guiding conditions for iterative reconstruction of the building outline extraction inference model; During iterative reconstruction processing, from the time step Initially, at a time step t, the mask of the building outline extraction inference model in its current state is... Using all guiding conditions as input, predict the mask of the previous time step with less noise. ; Until At that time, the building contour extraction inference model outputs the precise contour mask corresponding to the building in the remote sensing image data.
8. A device for accurately extracting building outlines based on prior information, characterized in that, The device includes: Information mining module: used to obtain remote sensing image data, perform prior information mining processing on the remote sensing image data for contour extraction, and obtain the contour extraction prior information corresponding to the remote sensing image data, wherein the contour extraction prior information includes local control prior information and global control prior information; Iterative Reconstruction Module: This module is used to input the remote sensing image data and the prior information for contour extraction into the building contour extraction inference model, and to perform iterative reconstruction processing on the remote sensing image data in the building contour extraction inference model with the prior information for contour extraction as the guiding factor, and output the accurate contour mask corresponding to the building in the remote sensing image data.
9. An electronic device comprising a processor and a memory, characterized in that, The processor runs a computer program or code stored in the memory to implement the method for accurately extracting building outlines as described in any one of claims 1 to 7.
10. A computer-readable storage medium for storing computer programs or code, characterized in that, When the computer program or code is executed by a processor, the method for accurately extracting building outlines as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Point cloud data disordered splicing method and system based on ground laser scanning equipment
CN117808673A
Multi-task consistency segmentation network method for cartilage segmentation based on deep shape prior model
CN120355913A
Priori structure rule constrained building contour optimization method, medium and equipment
CN120429927A
Cited By
Condition-controlled vector building sample data generation method and device
CN121639953A
Remote sensing image segmentation method based on cross-modal feature fusion and fine granularity compensation
CN121685950A