Confidence map post-processing method and system for building regularized mapping

By combining the confidence map post-processing method of CNN and Transformer, the challenges brought by mesoscale diversity and complex scenarios of building boundary extraction are solved, and the boundary extraction with higher accuracy and regularity is achieved, which is suitable for diverse building scenarios.

CN120070482APending Publication Date: 2025-05-30WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510050390.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing building boundary extraction methods have challenges in facing scale diversity and complex scenarios, resulting in problems of omissions, boundary blur and error accumulation.

Method used

The confidence map post-processing method combined with convolutional neural network (CNN) and transformer (Transformer) is adopted to achieve refined extraction and optimization of building boundaries through the confidence map reconstruction module and the building boundary regularization module.

Benefits of technology

It significantly improves the overall accuracy and regularity of building boundary extraction, is suitable for buildings of different scales and shapes, reduces the dependence on professional knowledge, and continuously optimizes model performance through data reflow mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070482A_ABST
    Figure CN120070482A_ABST
Patent Text Reader

Abstract

The invention discloses a confidence map post-processing method for building regularization mapping. The method comprises the following steps: firstly, designing and automatically generating a building regularization drawing sample set for training according to building characteristics in a high-resolution remote sensing image; secondly, constructing a confidence map post-processing method in combination with the advantages of Transform and a convolutional neural network, so as to improve the feature fusion capability of a local region and an overall scene of a confidence map and ensure the integrity of building extraction; thirdly, a generative adversarial module is introduced, and the fine degree of the building boundary is enhanced through a self-adaptive dichotomy process; and finally, verifying the validity of the method on a plurality of aviation and satellite data sets, and comparing and analyzing with the prior art. According to the method, the overall precision and regularity of building boundary extraction are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and remote sensing image processing, and specifically to a confidence map post - processing method and system for regular building mapping. Background Art

[0002] In application scenarios such as urban planning, urban management, and emergency response, accurately extracting building boundaries is crucial. With the acceleration of the global urbanization process, the demand for high - precision building distribution information has become increasingly urgent. During the inference process of traditional semantic segmentation networks, a specific threshold is set to perform binary classification on the confidence map output by the network to achieve building extraction. However, due to the influence of building size differences and the probability distribution error of deep models, this crude binary classification method may cause problems such as building omission and inaccurate positioning of target boundaries, thus hindering subsequent applications such as precise building mapping, automatic single - entityization, and precise statistics.

[0003] With the development of remote sensing technology and the ability to obtain high - resolution images, how to efficiently extract buildings with complete shapes and clear and regular boundaries from complex high - resolution remote sensing data has become a research hotspot. Although existing deep - learning - based methods have improved the effect of building segmentation to a certain extent, there are still challenges when facing scale diversity and complex scenarios. For example, changes in building scale may cause small buildings to be ignored; moreover, the probability distribution characteristics of the output of semantic segmentation models may also lead to the ellipse - like phenomenon of building right - angle boundaries, affecting the quality of segmentation results. In addition, the manifestations of buildings in high - resolution images vary due to factors such as roof materials, shapes, sizes, as well as imaging lighting conditions, occlusion, and shooting angles, which also increase the extraction difficulty.

[0004] Existing methods are mainly divided into two categories: The first category focuses on enhancing the representation ability of building semantic segmentation networks, especially by introducing boundary optimization modules to enhance the fine description of building boundaries. Although these methods improve the segmentation accuracy, they have limitations in terms of flexibility and adaptability. The second category of methods directly optimizes the binary - classified building mask or converts it into a polygon after vectorization and then optimizes it. However, the results of these post - processing methods are limited by the accuracy of semantic segmentation, and due to error accumulation, boundary extraction deviation may occur. At the same time, these methods often fail to specifically optimize the fuzzy or highly uncertain areas at the building edges, resulting in insufficient details of the extracted building boundaries.

[0005] In addition, traditional non-machine learning algorithms use morphological operations and curve optimization to optimize building boundaries. For example, the alpha shape algorithm is used to obtain the initial building boundary, and the dominant direction and line segments of the building are extracted through the Hough transform and combined with corner points to obtain a regular boundary through energy minimization. In recent years, with the development of deep learning technology, significant progress has been made in the fine extraction of buildings. For example, RegGAN uses a multi-scale discriminator to generate more realistic building roof boundaries; a coarse-to-fine network gradually refines building boundaries at different scales; the building features are decoupled into themes and boundaries, mimicking the human way of identifying buildings to achieve accurate boundary extraction. However, in the case of lack of diversity in training data, these methods often perform unstably in the target domain and fail to fully address the omission and misclassification problems brought about by the process from the probability value of model inference to the building contour.

[0006] In summary, the current building boundary extraction methods face many challenges, including but not limited to omissions caused by scale changes, boundary blurring caused by probability distribution characteristics, and error accumulation in post-processing optimization. Therefore, there is an urgent need to develop a new method that can overcome the above problems and provide a more accurate and complete building boundary extraction solution to meet the growing needs of urban management and disaster monitoring. Summary of the Invention

[0007] In view of this, it is necessary to develop a method for optimizing building boundaries that can overcome the above problems. The present invention proposes a confidence map post-processing method for building regularization mapping. This network uses an image-assisted confidence map for binary segmentation and combines the advantages of Transformer and CNN to improve the integrity of building extraction. At the same time, a generative adversarial module is designed to achieve adaptive binary classification of the confidence map, thereby improving the fineness of the boundary.

[0008] The present invention relates to a confidence map post-processing method for building regularization mapping, aiming to improve the extraction accuracy of building boundaries in remote sensing images by combining the convolutional neural network (CNN) and the transformer (Transformer) in deep learning technology, as well as the generative adversarial network (GAN). The following are the main steps of this technical method:

[0009] 1. Data Preparation:

[0010] Obtain high-resolution remote sensing images and use models such as FCN+ResNet, DeepLabv3+ResNet, or SegFormer to generate boundary confidence maps. Construct a sample set, including the original images, ground truths, and corresponding confidence maps.

[0011] 2. Network Construction and Training:

[0012] Design a post - processing network that includes a confidence map reconstruction module and a building boundary regularization module.

[0013] The confidence map reconstruction module adopts a dual - branch encoder structure. One branch is a CNN for capturing local details, and the other branch is a Transformer for perceiving global context information.

[0014] The building boundary regularization module introduces a GAN structure to optimize the expression of building boundaries.

[0015] 3. Boundary optimization:

[0016] The input image and the confidence map are respectively fed into the CNN and Transformer encoders for feature extraction. The extracted feature vectors are combined by weighted fusion, and the weights are calculated through a fully - connected layer. The fused features are passed to the residual decoder and transformed into a binary building mask to ensure that the generated building boundaries are more realistic and regular.

[0017] 4. Post - processing and manual interaction correction:

[0018] If the building boundary extraction result output by the model does not reach the set accuracy threshold, it is necessary to improve the model through manual re - interpretation and annotation.

[0019] To achieve the above objectives, the technical method of the present invention is as follows:

[0020] Step 1: Obtain high - resolution remote sensing images, and through a semantic segmentation model, predict the respective boundary confidence maps, which together with the original images and ground truths form a sample set.

[0021] Step 2: Combine Transformer and CNN to construct a confidence map post - processing network for building regularization mapping. This confidence map post - processing network includes a confidence map reconstruction module for capturing multi - scale features and a building boundary regularization module for optimizing building boundaries. Use the sample set in Step 1 to train the network to obtain a confidence map post - processing model.

[0022] Step 3: Use a building segmentation model to segment the buildings on the remote sensing image to be processed to obtain a building confidence map, and then infer the building confidence map through the trained confidence map post - processing model to obtain a fine building boundary extraction result.

[0023] Further, in step 1, first perform noise elimination and image quality enhancement on the original high-resolution remote sensing image, then use the FCN+ResNet or DeepLabv3+ResNet or SegFormer model to predict the respective boundary confidence maps, and finally correspond the original image, ground truth, and confidence Figure 1 one by one, and crop them to a certain size to generate a refined building extraction sample set containing n pairs of samples.

[0024] Further, the confidence map reconstruction module uses a dual-branch encoder to capture multi-scale features. One branch uses a CNN to capture local details, and the other branch uses a Transformer to perceive global context information. Then, the features extracted by the CNN and Transformer branches are weighted and fused, and the information after feature fusion is passed to the residual decoder to convert the fused features into the final binary building mask;

[0025] The generative adversarial network GAN structure is introduced into the building boundary regularization module. The generator G includes a dual-branch encoder and a residual decoder, which is used to generate realistic building masks, and the discriminator D is responsible for distinguishing the generated building masks from the real building boundaries.

[0026] Further, the specific processing process of the confidence map reconstruction module is as follows:

[0027] (A1) Input processing: Take the image and the building confidence map as inputs, and send them into the CNN and Transformer encoders respectively for local feature extraction and global semantic relationship capture;

[0028] (A2) Local feature extraction: Use the CNN encoder to extract local features of the building boundary through convolution and pooling operations;

[0029] (A3) Global semantic relationship capture: Use the self-attention mechanism of the Transformer encoder to obtain the global context information of the input image;

[0030] (A4) Feature vector generation: Embed the feature representations generated by the CNN and Transformer encoders into a sequence of feature vectors as the input for the subsequent fusion stage;

[0031] (A5) Feature weighted fusion: Calculate the fusion weights of the two feature branches through a fully connected layer, and perform weighted fusion on the features extracted by the CNN and Transformer;

[0032] (A6) Feature optimization: Use the backpropagation algorithm to update the weight parameters;

[0033] (A7) Output generation: Use the fused features as the input for the next module to provide comprehensive feature support for subsequent building boundary optimization.

[0034] Furthermore, the specific processing process of the building boundary optimization module is as follows:

[0035] (B1) Boundary generation by the generator: Use the multi-scale features generated by the confidence map reconstruction module as the input, and decode the features through a residual decoder to generate an initial building mask;

[0036] (B2) Boundary evaluation by the discriminator: Use the building mask generated by the generator and the real building boundary as the input, send them into the discriminator for true / false discrimination, and generate an adversarial loss signal based on the discriminator's judgment result to feedback to the generator, forcing the generator to continuously improve the mask quality;

[0037] (B3) Loss function optimization: Design a loss function, and optimize the parameters of the generator and the discriminator through the backpropagation algorithm to make the building mask generated by the generator gradually approach the real boundary.

[0038] Furthermore, the loss functions used during training include: Reconstruction loss: Calculate the binary cross-entropy loss between the generated mask and the real mask to ensure that the overall shape of the generated result is consistent with the real situation. Adversarial loss: Define the adversarial objective between the generator and the discriminator to constrain the boundary shape output by the generator to be more regular and real; Potts loss and Normalized Cut loss: Optimize the regularization result of the generated mask according to the weight relationship between nodes in the image to enhance the consistency with the real boundary; linearly combine the reconstruction loss, adversarial loss, and regularization loss as the final loss function.

[0039] Furthermore, the calculation formula for the reconstruction loss is as follows:

[0040] L G (G) = -E x,z [y · logG(x,z)] (1)

[0041] where L G (G) is the reconstruction loss, E is the expectation, G is the generator, the building contour confidence map constitutes the X domain, the ground truth building mask constitutes the Y domain, the remote sensing image constitutes the Z domain, x, y, z belong to X, Y, Z respectively, and G(x,z) is the output of the generator;

[0042] The calculation formula for the adversarial loss is as follows:

[0043] L GAN (G,D) = E x,z [log(1 - D(G(x,z)))] (2)

[0044] Among them, L GAN (G, D) is the adversarial loss, D is the discriminator, and D(G(x, z)) is the output of the discriminator.

[0045] Furthermore, the Potts loss is calculated as follows:

[0046]

[0047] Among them, E is the expectation, and G is the generator;

[0048] The NormalizedCut loss is calculated as follows:

[0049]

[0050] Among them, L Potts (G) is the Potts loss, and L ncut (G) is the NormalizedCut loss, S = G(x, z) is the generated k-channel mask, and the softmax function is used for each channel to obtain the confidence. S kT describes the transpose of the value of its k-th channel; W is the pairwise discontinuity cost matrix, which is used to measure the difference or connection strength between each pair of adjacent pixels in the image. Each element in this matrix represents the weight between two pixels and is calculated according to the gray-scale similarity metric; is the pairwise discontinuity cost matrix after over-normalization processing.

[0051] Furthermore, it also includes optimizing the confidence map post-processing model through a computer post-processing and manual interaction correction mechanism. Specifically: set an accuracy threshold a, automatically evaluate whether the building boundary extraction result meets the standard. If the prediction accuracy does not reach the threshold a, then these test data are fed back for manual re-interpretation and annotation to construct new labels, expand the existing dataset, or delete incorrect labels.

[0052] The present invention also provides a confidence map post-processing system for building regularized mapping, including:

[0053] A processor and a memory. The memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute a confidence map post-processing method for building regularized mapping as described in the above technical solution.

[0054] In summary, the present invention provides a confidence map post-processing method for building regularized mapping, which solves the challenges brought by the building scale diversity and improves the accuracy and detail integrity of the building boundary. In addition, through the data feedback mechanism, the algorithm can be continuously improved during the actual application process, providing strong support for fields such as urban planning and disaster monitoring.

[0055] The present invention verifies the effectiveness of the method on multiple aviation and satellite datasets and conducts a comparative analysis with the prior art. The method of the present invention significantly improves the overall accuracy and regularity of building boundary extraction, especially performs well in dealing with buildings of different scales and shapes, and is applicable to application scenarios such as urban planning, urban management, and disaster emergency response, providing strong technical support for accurate building mapping, automatic monomerization, and statistics. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0057] Figure 1 is a flowchart of an embodiment of the present invention;

[0058] Figure 2 is a schematic diagram of the model architecture of an embodiment of the present invention;

[0059] Figure 3 is a schematic diagram of the method for automatically constructing refined building extraction samples provided by an embodiment of the present invention;

[0060] Figure 4 is a schematic diagram of a refined building extraction sample set of an embodiment of the present invention;

[0061] Figure 5 is a comparison of the experimental results of the present invention on the INRIA dataset. Among them: (a) is the original image; (b) is the ground truth; (c) is the confidence map; (d) is the binarization with a confidence threshold of 0.5; (e) is the Otsu method; (f) is the MFCNN method; (g) is the method proposed by Zorzi et al.; (h) is the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0063] It should be understood that the schematic drawings are not drawn to scale. The flowcharts used in the present invention illustrate operations implemented according to some embodiments of the present invention. It should be understood that the operations in the flowchart may not be implemented in sequence, and steps without logical context relationships may be reversed or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of the present invention. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor systems and / or microcontroller systems.

[0064] The descriptions such as "first" and "second" involved in the embodiments of the present invention are only for implicit purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Therefore, the technical features defined with "first" and "second" may explicitly or implicitly include at least one such feature. "And / or" describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.

[0065] Referring to "embodiment" herein means that a specific feature, structure or characteristic described in connection with the embodiment can be included in at least one embodiment of the present invention. The appearance of this phrase at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0066] The main technical difficulties to be solved by the present invention:

[0067] 1. Construct a large-scale refined extraction dataset of buildings covering various terrains and scenes:

[0068] By collecting high-resolution remote sensing images from different countries and regions around the world, covering a rich variety of architectural styles and types, and accurately marking the buildings on these images, high-quality marked images are obtained. This dataset includes not only data from aerial sensors (such as the INRIA building dataset) and satellite sensors (such as the WHU satellite dataset II), but also pays special attention to the building distribution in the East Asian region. In addition, in order to ensure the diversity and representativeness of the sample library, we also manually annotated some unique building instance data. These images cover different time periods, seasons, lighting conditions, and shooting angles, thus ensuring the broad representativeness and diversity of the dataset.

[0069] 2. Adapt to multi-scale buildings and handle complex boundary details. By combining the advantages of convolutional neural networks (CNNs) and Transformers, the present invention effectively captures multi-scale features, solves the problem of small object recognition that is difficult to handle by traditional methods, and can meet the requirements of handling complex building boundary details. For buildings and boundary characteristics of various scales, the present invention ensures fine adjustment of their boundaries to make them more conform to the actual situation on the basis of maintaining the integrity of the buildings through the fusion of local feature extraction and global context understanding. At the same time, by introducing a generative adversarial network (GAN) structure, the building boundary is further optimized, overcoming problems such as building omission and inaccurate boundary positioning caused by a single threshold in traditional methods.

[0070] 3. Solve the confidence Figure 2 Classification and optimize the regularization of building boundaries. For the challenges in the confidence Figure 2 classification process, the present invention designs a special combination of loss functions, including reconstruction loss, adversarial loss, Potts loss, and Normalized Cut loss. These loss functions optimize the regularization results by constraining the geometric boundaries of buildings, significantly improving the accuracy and consistency of building boundaries. Through this innovative design, the present invention not only effectively solves the binary classification problem of building masks, but also makes the finally generated mask highly consistent with the actual boundary, achieving the purpose of optimizing the regularization boundary results.

[0071] 4. Extracting building boundaries generally requires professional knowledge training, and the cost of manual extraction and review is relatively high. The traditional process of extracting building boundaries usually relies on manual operations or semi-automatic tools of professionals, which is not only time-consuming and laborious, but also easily affected by subjective factors. To solve this problem, the automated method proposed by the present invention can efficiently and accurately extract building boundaries from remote sensing images, greatly reducing the dependence on professional knowledge. Through the training and application of the algorithm provided by the present invention, even people without relevant backgrounds can quickly start using the algorithm. More importantly, the present invention introduces a data feedback mechanism, enabling the algorithm to continuously adjust and optimize the model performance according to the feedback, further improving the accuracy and efficiency of recognition. This self-improving ability is crucial for the long-term maintenance and update of building information, and also significantly reduces the need and cost of manual intervention.

[0072] The specific implementation manner of the complete technical solution of the present invention is described as follows:

[0073] The present invention discloses a technical solution for optimizing building boundaries. By designing a confidence map post-processing method for building regularization mapping, combining multi-scale feature fusion and boundary optimization processing of the building confidence map, refined extraction of building boundaries is achieved. As Figure 1As shown in the figure, the present invention includes the generation of refined building extraction samples, the construction of a confidence map post-processing algorithm, and its training and application. The specific implementation methods are as follows:

[0074] 1. Generation of refined building extraction samples: As Figure 3 shown, an automated method is used to generate a dataset containing confidence maps of buildings with different qualities. These datasets are collected by a variety of aerospace and aviation sensors, covering different countries and regions around the world, and having rich and diverse architectural styles and types.

[0075] (1) Data source: The sample data comes from high-resolution remote sensing images, including the INRIA public dataset, which covers buildings in different countries and regions.

[0076] (2) Preprocessing: Noise elimination and image quality enhancement are performed on the original images.

[0077] (3) Confidence map generation: Multiple mainstream semantic segmentation models (such as FCN+ResNet, DeepLabv3+ResNet, SegFormer) are used to predict the remote sensing images to generate diverse building confidence maps to supplement training samples under different conditions.

[0078] (4) Final sample set: The original images, ground truths, and confidence Figure 1 maps are paired one by one and cropped to a size of 512×512 pixels to generate a refined building extraction sample set containing n pairs of samples, ensuring sample diversity and representativeness. As Figure 4 shown.

[0079] 2. Construction of the confidence map post-processing algorithm: The algorithm includes a confidence map reconstruction module and a building boundary optimization module. The former captures multi-scale features through a dual-branch encoder (CNN and Transformer), and the latter introduces a GAN structure to implement building boundary optimization processing. The two are jointly trained to ensure that the finally generated building boundaries are more realistic and accurate. As Figure 2 shown.

[0080] The confidence map reconstruction module captures multi-scale features through a dual-branch encoder: a convolutional neural network (CNN) and a Transformer. They are represented by G C and G T respectively. Among them, G C processes the image through stacked convolutional and pooling operations. The convolutional kernel slides over a window of a fixed size to detect different features E C (x,z) in the image, such as edges, textures, etc. The extraction of such local features helps to retain the details and textures of the building contours, making them clearer. G TThe self-attention mechanism is adopted, which can capture the global dependencies in the image. In this module, the input intensity image and confidence map are respectively embedded as sequences of feature vectors and input into G T In it, the output E T (x,z) = G T (x,z) integrates the semantic relationships of the entire image. G T Effectively learns the global context information of the image through the self-attention mechanism, enabling the model to better understand the overall structure of the building contour and helping to generate more accurate and holistic contours.

[0081] This design makes full use of the advantages of the two architectures: CNN is good at local feature extraction, such as edges, textures, etc.; while Transformer can capture global dependencies through the self-attention mechanism and understand the overall structure of the building boundary. Using Transformer as the encoder to gradually extract and refine features through a hierarchical method. At different network levels, the features are downsampled and transformed to form feature representations of different scales. Transformer as the encoder provides powerful feature extraction capabilities. Transformer includes several stacked Transformer Blocks, and each Transformer block includes an MLP layer. The calculation process of the MLP layer is described as:

[0082]

[0083] In the formula, Z 1 is the first linear transformation formula, X is the input, U is the weight matrix of the first linear mapping, b 1 is the bias term corresponding to U, A is the activation function, σ is the ReLU function, Y is the output, V is the weight matrix of the second linear mapping, b 2 is the bias term corresponding to V.

[0084] Weighted fusion is performed on the features extracted by the CNN and Transformer branches, and buildings of different sizes can be extracted. Specifically, the features E C (x,z) and E T (x,z) are concatenated, and then a fully connected layer is used to calculate the feature weights to achieve feature fusion. During the training process, the weight parameters are updated through the backpropagation algorithm, enabling the model to learn the optimal weights to fuse the features of the two modules. This process ensures that the model can both focus on local details and maintain an understanding of the overall scene, thereby improving the quality and visual effect of the generated results.

[0085] The information after feature fusion is passed to the residual decoder, which is responsible for converting the fused features into the final binary mask. The design of the residual decoder allows the network to retain more original information while reconstructing the image layer by layer, reducing information loss and ensuring the clarity and accuracy of the building boundaries.

[0086] The building boundary optimization module takes the multi-scale features extracted in the confidence map reconstruction module as input, and then the decoder generates the building mask. The discriminator D is responsible for distinguishing the generated building mask from the real building boundary. As the training progresses, the discriminator D becomes more sensitive, forcing the generator G to continuously improve the generation quality until the generated result is difficult to be distinguished as fake. During this process, the generator G and the discriminator D are updated synchronously. The generator G tries to generate a boundary closer to the real situation, while the discriminator D tries to correctly classify whether the boundary is real or not. Finally, the generator G is trained to be able to generate a building mask highly similar to the real situation.

[0087] In summary, the confidence map post-processing algorithm of the present invention mainly consists of three modules. Module 1 is the confidence map reconstruction module composed of CNN and Transformer. Module 2 is the building boundary optimization module based on the generative adversarial network. Module 3 is the decoder composed of the residual decoder. These three are connected in series and change the single loss function to a combination of multiple loss functions.

[0088] 2.1, Confidence Map Reconstruction Module

[0089] (1) Input processing: The image and the building confidence map are used as inputs and fed into two encoders, CNN and Transformer respectively, for local feature extraction and global semantic relationship capture.

[0090] (2) Local feature extraction: The CNN encoder is used to extract local features of the building boundary through convolution and pooling operations, including texture, edge information, etc.

[0091] (3) Global semantic relationship capture: The self-attention mechanism of the Transformer encoder is used to obtain the global context information of the input image and enhance the understanding of the overall structure.

[0092] (4) Feature vector generation: The feature representations generated by the CNN and Transformer encoders are respectively embedded into a sequence of feature vectors as the input for the subsequent fusion stage.

[0093] (5) Feature weighted fusion: The fusion weights of the two feature branches are calculated through a fully connected layer, and the features extracted by CNN and Transformer are weighted and fused.

[0094] (6) Feature Optimization: Use the backpropagation algorithm to update the weight parameters, enabling the model to automatically learn the optimal fusion method and better adapt to the building boundary extraction requirements at different scales.

[0095] (7) Output Generation: Use the fused features as the input for the next module, providing comprehensive feature support for subsequent building boundary optimization.

[0096] 2.2, Building Boundary Optimization Module

[0097] (1) Boundary Generation by the Generator: Use the features generated by the confidence map reconstruction module as the input, and decode the features through the residual decoder to generate the initial building mask.

[0098] (2) Boundary Evaluation by the Discriminator: Use the building mask generated by the generator and the real building boundary as the input, and send them into the discriminator for true / false discrimination. Based on the judgment result of the discriminator, generate an adversarial loss signal and feedback it to the generator to force the generator to continuously improve the mask quality.

[0099] (3) Loss Function Optimization: Reconstruction Loss: Calculate the binary cross-entropy loss between the generated mask and the real mask to ensure that the overall shape of the generated result is consistent with the real situation. Adversarial Loss: Define the adversarial objective between the generator and the discriminator to constrain the boundary shape output by the generator to be more regular and real. Potts Loss and Normalized Cut Loss: Optimize the regularization result of the generated mask according to the weight relationship between nodes in the image to enhance the consistency with the real boundary.

[0100] Linearly combine the reconstruction loss, adversarial loss, and regularization loss, and optimize the parameters of the generator and discriminator through the backpropagation algorithm to make the building mask generated by the generator gradually approach the real boundary.

[0101] Adopt the binary cross-entropy loss function, and by comparing the differences between the generated mask and the real mask pixel by pixel, force the generator to focus on the main structural areas of the building to ensure that the generated building mask is consistent with the real mask in terms of overall shape and boundary position.

[0102] L G (G) = -E x,z [y · logG(x,z)] (2)

[0103] where L G (G) is the reconstruction loss, E is the expectation, G is the generator, given the training samples and x i ∈ X, y i ∈ Y, z i∈Z, the building contour confidence map constitutes the X domain, the ground truth building mask constitutes the Y domain, the original remote sensing image constitutes the Z domain, and G(x,z) is the binary building mask output by the generator.

[0104] Through the mechanism of the Generative Adversarial Network (GAN), the mask generated by the generator is made sufficient to deceive the discriminator, thus approaching the real situation more closely in terms of geometric boundaries. The generator attempts to minimize the adversarial loss, while the discriminator attempts to maximize this loss, thereby forming a game between the generator and the discriminator to optimize the authenticity and detail performance of the generated mask.

[0105] L GAN (G,D) = E x,z [log(1 - D(G(x,z)))] (3)

[0106] Among them, L GAN (G,D) is the adversarial loss, D is the discriminator, and D(G(x,z)) is the output of the discriminator.

[0107] The Potts loss is introduced to impose geometric constraints on neighboring pixels, promoting the boundary smoothness of the generated mask while avoiding excessive small isolated areas. By calculating the discontinuity cost between adjacent pixels, the local structure of the generated mask is optimized.

[0108]

[0109] The NormalizedCut loss is introduced. Based on the graph segmentation theory, the global semantic information of the image is used to further optimize the regularity of the mask boundary. By cutting the graph structure formed by the image pixels, the connection strength between different regions is reduced while the consistency within the regions is increased.

[0110]

[0111] In equations (4) and (5), L Potts (G) is the Potts loss, and L ncut (G) is the NormalizedCut loss. S = G(x,z) is the k-channel mask generated by the network. The softmax function is used to obtain the confidence for each channel, and S kT describes the transpose of the value of its k-th channel. W is the pairwise discontinuity cost matrix, which is a matrix used to measure the difference or connection strength between each pair of adjacent pixels in the image. Each element in this matrix represents the weight between two pixels. For each pixel pair (i,j), the weight W i,j is calculated according to the gray-scale similarity metric, and this weight can reflect whether they should be assigned to the same class or different classes. is the pairwise discontinuity cost matrix that has been normalized. represents the discontinuity cost between adjacent pixels and is normalized; the normalization factor ensures the numerical stability of the loss function and avoids unstable outputs in extreme cases.

[0112] During the training process, the present invention linearly combines the adversarial loss, the regularization loss, and the reconstruction loss:

[0113] L(G,D) = αL GAN (G,D) + βL G (G) + γL Potts (G) + δL ncut (G) (6)

[0114] where α, β, γ, and δ are parameters.

[0115] 3. Training and generation of the confidence map post - processing method: The present invention is trained based on the generated refined building extraction dataset, using a building confidence map post - processing model for refined building extraction. After training, the confidence map is input into the building confidence map post - processing model for inference, and a refined building mask can be directly generated to meet the actual application requirements.

[0116] 4. The present invention also optimizes the building prediction model through a computer post - processing and manual interaction correction mechanism. Specifically, a precision threshold a is set to automatically evaluate whether the prediction results meet the standard. If the prediction accuracy does not reach the threshold a, these test data are fed back for manual re - judgment and annotation to construct new labels, expand the existing dataset, or delete incorrect labels. This process not only improves the accuracy of the model but also effectively controls the workload of manual secondary annotation by adjusting the threshold a, ensuring that the model continuously improves according to the feedback.

[0117] Experiments of the present invention:

[0118] To verify the effectiveness of the present invention, we conducted experiments on the INRIA building segmentation dataset respectively. These two datasets contain diverse building images from different countries around the world, with rich building styles and types.

[0119] The INRIA (Inria Aerial Image Labeling Dataset) dataset contains multiple remote - sensing aerial images that capture diverse ground features in urban and rural areas. It consists of 180 training sets and 180 test sets, with a resolution of 0.3 meters and a sample size of 5000×5000 pixels.

[0120] To verify the optimization ability of the present invention as a post - processing tool, two sets of segmentation confidence maps were obtained using the baseline methods DeepLabv3+ and Segformer semantic segmentation networks, and were respectively compared with the fixed - threshold method, MFCNN + morphological filtering used by Otsu, Xie, etc., and the GAN - based boundary optimization method proposed by Zorzi, etc. on three datasets. These methods were used to process the two sets of preliminary segmentation results to generate the final predicted building boundary map. Since the input of the present invention includes segmentation confidence maps, we adopted an activation function suitable for probability estimation in the output layer of the baseline method, and directly used the probability value output by the activation function as the confidence map. Through this processing method, the output of the model will be a confidence map representing the confidence of each pixel belonging to the building, rather than a hard binary result.

[0121] Figure 5 The comparative experimental results on the INRIA dataset are shown, from which it can be clearly seen that the present invention shows significant optimization effects on confidence maps with different accuracies. As Figure 5 shown, the baseline method and the Otsu method are both global binary methods. These methods overly rely on the segmentation ability of the segmentation model, resulting in low - confidence regions being misclassified as the background. Therefore, a single - threshold method is not suitable for the binaryzation of building confidence maps. In contrast, the present invention does not require special requirements for foreground sample selection and can extract buildings of different shapes more accurately. In addition, as Figure 5 shown in the DeepLabv3+ part of (h), even in the case where the contrast between the building and the surrounding background is low, the present invention can make accurate judgments in low - confidence regions and generate a binary result consistent with the actual building boundary, fully demonstrating the adaptability and robustness of the present invention. In contrast, it is difficult for a single - threshold method to handle such complex scenarios. By comparing Figure 5 (g) and (h), it can be found that the building boundary extracted by the present invention is more regular, the geometric shape is more realistic, and it has more accurate right - angle boundaries. In the area where buildings are dense, the present invention can correctly distinguish adjacent buildings and retain more detailed information. This further proves the advantages of the present invention in dealing with complex building segmentation tasks.

[0122] The embodiment of the present invention also provides a confidence - map post - processing system for building regularization mapping, including:

[0123] A processor and a memory. The memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute a confidence - map post - processing method for building regularization mapping as described in the above technical solution.

[0124] The above method has described the examples of the present invention in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims. All of these fall within the protection scope of the present invention.

Claims

1. A confidence map post-processing method for building regularization mapping, characterized in that: The steps include: Step 1: Obtain high-resolution remote sensing images and predict the boundary confidence maps of each image through the semantic segmentation model, which together with the original image and the true value constitute a sample set; Step 2: Combining Transformer and CNN to build a confidence map post-processing network for building regularization mapping, the confidence map post-processing network includes a confidence map reconstruction module for capturing multi-scale features and a building boundary regularization module for optimizing building boundaries; the network is trained using the sample set in step 1 to obtain a confidence map post-processing model; Step 3: Use the building segmentation model to segment the buildings on the remote sensing image to be processed to obtain a building confidence map, and then use the trained confidence map post-processing model to infer the building confidence map to obtain a detailed building boundary extraction result.

2. A confidence map post-processing method for building regularization mapping as claimed in claim 1, characterized in that: In step 1, the original high-resolution remote sensing image is first subjected to noise elimination and image quality enhancement, and then the FCN+ResNet or DeepLabv3+ResNet or SegFormer model is used to predict the respective boundary confidence maps. Finally, the original image, true value, and confidence map are matched one by one and cropped to a certain size to generate a building refined extraction sample set containing n pairs of samples.

3. The confidence map post-processing method for building regularization mapping according to claim 1, characterized in that: The confidence map reconstruction module uses a dual-branch encoder to capture multi-scale features. One branch uses CNN to capture local details, and the other branch uses Transformer to perceive global context information. The features extracted by the CNN and Transformer branches are then weighted fused. The information after feature fusion is passed to the residual decoder, which converts the fused features into the final binary building mask. The generative adversarial network (GAN) structure is introduced in the building boundary regularization module, where the generator G includes a dual-branch encoder and a residual decoder to generate realistic building masks, and the discriminator D is responsible for distinguishing the generated building masks from the real building boundaries.

4. A confidence map post-processing method for building regularization mapping as claimed in claim 1 or 3, characterized in that: The specific processing process of the confidence map reconstruction module is as follows: (A1) Input processing: The image and building confidence map are taken as input and sent to the CNN and Transformer encoders respectively to extract local features and capture global semantic relationships; (A2) Local feature extraction: The CNN encoder is used to extract local features of building boundaries through convolution and pooling operations; (A3) Global semantic relationship capture: The self-attention mechanism of the Transformer encoder is used to obtain the global context information of the input image; (A4) Feature vector generation: The feature representations generated by the CNN and Transformer encoders are embedded into feature vector sequences as input to the subsequent fusion stage; (A5) Weighted feature fusion: The fusion weights of the two feature branches are calculated through the fully connected layer, and the features extracted by CNN and Transformer are weightedly fused; (A6) Feature optimization: Use the back propagation algorithm to update the weight parameters; (A7) Output generation: The fused features are used as the input of the next module to provide comprehensive feature support for subsequent building boundary optimization.

5. The confidence map post-processing method for building regularization mapping as claimed in claim 3, characterized in that: The specific processing process of the building boundary optimization module is as follows: (B1) Boundary generation of the generator: The multi-scale features generated by the confidence map reconstruction module are taken as input, the features are decoded by the residual decoder and the initial building mask is generated; (B2) Boundary evaluation of the discriminator: The building mask generated by the generator and the real building boundary are used as input and sent to the discriminator for true and false identification. Based on the judgment result of the discriminator, an adversarial loss signal is generated and fed back to the generator, forcing the generator to continuously improve the mask quality; (B3) Loss function optimization: Design the loss function and optimize the parameters of the generator and discriminator through the back-propagation algorithm so that the building mask generated by the generator gradually approaches the real boundary.

6. A confidence map post-processing method for building regularization mapping as claimed in claim 5, characterized in that: The loss functions used in the training process include: reconstruction loss: calculating the binary cross entropy loss between the generated mask and the real mask to ensure that the overall shape of the generated result is consistent with the real situation, adversarial loss: defining the adversarial goal between the generator and the discriminator to constrain the boundary shape of the generator output to be more regular and realistic; Potts loss and Normalized Cut loss: according to the weight relationship between the nodes in the image, optimize the regularization result of the generated mask to enhance the consistency with the real boundary; the reconstruction loss, adversarial loss and regularization loss are linearly combined as the final loss function.

7. A confidence map post-processing method for building regularization mapping as claimed in claim 6, characterized in that: The calculation formula of reconstruction loss is as follows: L G (G)=-E x,z [y·logG(x,z)] (1) Among them, L G (G) is the reconstruction loss, E is the expectation, G is the generator, the building outline confidence map constitutes the X domain, the ground truth building mask constitutes the Y domain, the remote sensing image constitutes the Z domain, x, y, z belong to, Y, Z respectively, G(x,z) is the output of the generator; The formula for calculating the adversarial loss is as follows: L GAN (G,D)=Ex,z[log(1-D(G(x,z)))] (2) Among them, L GAN (G,D) is the adversarial loss, D is the discriminator, and D(G(x,z)) is the output of the discriminator.

8. The confidence map post-processing method for building regularization mapping according to claim 6, characterized in that: The calculation of Potts loss is as follows: Among them, E is the expectation and G is the generator; The NormalizedCut loss is calculated as follows: Among them, L Potts (G) is the Potts loss, L ncut (G) is the NormalizedCut loss, S = G(x, z) is the generated k-channel mask, and the softmax function is used to obtain the confidence for each channel. kT Describes the transpose of the value of its kth channel; W is a pairwise discontinuous cost matrix, a matrix used to measure the difference or connection strength between each pair of adjacent pixels in the image. Each element in this matrix represents the weight between two pixels, calculated based on the grayscale similarity measure; is the normalized pairwise discontinuous cost matrix.

9. The confidence map post-processing method for building regularization mapping according to claim 1, characterized in that: It also includes optimizing the confidence map post-processing model through computer post-processing and manual interactive correction mechanism. Specifically: setting an accuracy threshold a to automatically evaluate whether the building boundary extraction results meet the standards. If the prediction accuracy does not reach the threshold a, the test data will be returned for manual re-interpretation and annotation to build new labels, expand existing data sets or delete erroneous labels.

10. A confidence map post-processing system for building regularization mapping, characterized in that: include: A processor and a memory, the memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute a confidence map post-processing method for building regularization mapping as described in any one of claims 1 to 9.