Building change detection method fusing building explicit prior and multi-stage feature aggregation

By introducing explicit priors for edges and corners in building change detection, and combining multi-level feature fusion and interaction modules, the problems of ignoring edge and corner features and difficulty in associating high-level and low-level features are solved, thus achieving higher accuracy in building change detection.

CN120931944APending Publication Date: 2025-11-11CAPITAL NORMAL UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510815322.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies neglect corner features in building change detection, and it is difficult to correlate high-level features with low-level features, making it difficult to accurately extract changed buildings.

Method used

In the feature extraction stage, explicit priors of building edges and corners are introduced. Multi-level feature fusion is used to enhance the network's ability to extract edge information. At the output end, edge and corner information is used to optimize and change building edges. Multi-level feature interaction modules are combined to adaptively select features for fusion.

Benefits of technology

It improves the accuracy of building change detection, enhances the ability to segment the edges and corners of changed buildings, and makes the detection results more consistent with human perception, reducing edge smoothing and corner blurring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120931944A_ABST
    Figure CN120931944A_ABST
Patent Text Reader

Abstract

The invention discloses a building change detection method fusing building explicit prior and multi-stage feature aggregation, and belongs to the technical field of image processing. The method comprises the steps of obtaining a double-time-phase image at a to-be-detected position, inputting the obtained double-time-phase image into a trained building change detection network to obtain a changed building prediction map of the double-time-phase image, and determining a building change result according to the changed building prediction map. According to the method, in the feature extraction stage, the explicit prior of the edge of the building is introduced, and the edge information extraction capacity of the network is enhanced. Meanwhile, in the feature fusion stage, interaction of local-global information is promoted in a multi-level feature fusion mode, and the reusability of the multi-scale feature map is enhanced. According to the method, the edge and corner information is utilized again at the output end to enhance the optimization capability of the network on the changed building edge, and the building change detection precision is jointly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, specifically to a method for detecting building changes that integrates explicit building priors and multi-stage feature aggregation. Background Technology

[0002] Urbanization plays a vital role in promoting high-quality economic development and improving people's quality of life. As one of the most dynamic man-made structures in a city, the detection of changes in buildings is a major task in remote sensing imagery applications. Effectively monitoring changes in building distribution is crucial for urban cadastral map updates, land resource management, and natural disaster assessment. With the continuous improvement of satellite resolution and the development of automated information extraction technologies, ground features possess clearer edge contours and morphological characteristics. Improving the accuracy of change detection and reducing the detection of false change areas by extracting and learning the semantic information and practical features such as spatial structure and edges of changed buildings has become possible, enabling high-precision building change detection.

[0003] In the early stages, pixel-based and object-based methods were well-suited for shape detection (BCD) of low- and medium-resolution remote sensing images. However, with the development of sensor technology, satellite sensors can acquire richer spatial textures and structural features of the Earth's surface, making the BCD results of traditional methods insufficient for practical applications. In recent years, with the advancement of computer technology, deep learning has demonstrated powerful feature extraction and fusion capabilities in remote sensing image processing. In BCD, deep learning techniques aggregate low-level features to abstractly represent high-level features, and extract regions of interest through metric analysis or classification strategies. Convolutional neural networks (CNNs), primarily based on image convolution, and Transformer architectures, primarily based on self-attention, have demonstrated superior performance in change detection. Commonly used change detection structures are mainly single-branch networks and two-branch networks (Siamese architecture), with studies proving that two-branch networks have higher accuracy than single-branch networks.

[0004] CNN-based architectures extract features at different scales through multiple convolutions, aiming to incorporate global, background, and semantic information. They also extract change information by aggregating multi-scale features or interacting with peer features. In feature maps at different scales, high-level features focus on describing global information, capturing the global semantics of changing buildings in remote sensing images, but their ability to model building boundaries is poor. Low-level features describe local textures, spectra, and other information in remote sensing images, helping to determine details such as edges, but they cannot determine whether an edge is related to a changing building. Therefore, fusing high-level and low-level features—that is, jointly modeling the global scene semantics and local low-level detail textures of changing buildings—is crucial for improving the accuracy of building change detection. Due to the significant semantic gap between high-level and low-level features, it is difficult to correlate them, lacking the modeling of long-distance dependencies, which poses a significant challenge to the accurate extraction of changing buildings and edge segmentation.

[0005] Currently, the development of self-attention mechanisms is booming, and they are favored by researchers due to their ability to effectively model long-range dependencies. The Transformer architecture based on self-attention (SA) does not require downsampling and pooling operations during encoding, and it captures complex spatial transformations and long-range dependencies more effectively than general CNN architectures. ChangeFormer was the first to utilize Transformer blocks and extract bi-temporal features at different levels. While this method can capture global features at different scales, directly using the features extracted by the Transformer cannot highlight the dominant features at local scales, which is obviously detrimental to the recovery of local details in changing architectures. Unlike inserting SA mechanisms into CNNs, some researchers have attempted to combine CNNs and Transformers at the network architecture level, while preserving the local features of CNNs and the global features of Transformers. In terms of CD, TransUNetCD extracts multi-scale features from bi-temporal images using CNNs, and then uses Transformers to encode the CNN feature maps in blocks, associating them with rich global contextual information. Liu et al. enhanced the features extracted by CNNs and input them into the TR module to process long-range dependencies, modeling contextual information in bi-temporal images. Xu et al. used the PS-ViT module to establish global contextual relationships for features extracted by CNN modules.

[0006] In natural scenes, buildings possess regular lines and smooth boundary features. Identifying the semantic attributes of changing building boundary points, whether within the building itself or the background, is typically challenging. Since the accurate classification of these potential boundary points is highly correlated with the final recognition accuracy, improving the classification accuracy of these points is crucial for enhancing the overall accuracy of building change detection. Some researchers have improved the network's accuracy in recognizing changing building boundary points by designing boundary-aware modules and spatial detail capture branches during the feature extraction stage to capture more spatial details in the encoder segment; and by designing boundary decoders and edge refinement modules at the decoding end, and by designing boundary loss to enhance the network's ability to perceive and refine edge regions. While these methods effectively improve the detection accuracy of BCD by using the boundary contour as prior information for the network, they often directly transplant the boundary contours calculated based on feature pixel values ​​in the CV field. For buildings with regular lines and sharp angles, considering only the edge contours and ignoring the corner features is clearly insufficient. Summary of the Invention

[0007] In view of this, the purpose of this invention is to provide a building change detection method that integrates explicit building priors and multi-stage feature aggregation, in order to solve the problems of neglecting corner features during existing feature extraction and the difficulty in associating high-level features with low-level features, which leads to difficulties in accurately extracting changed buildings.

[0008] To achieve the above objectives, in the feature extraction stage, this invention enhances the network's ability to extract edge information by introducing explicit priors about building edges. Simultaneously, in the feature fusion stage, multi-level feature fusion is used to promote the interaction of local and global information, enhancing the reusability of multi-scale feature maps. At the output stage, the method further utilizes edge and corner information to enhance the network's optimization ability for changing building edges, collectively improving the accuracy of building change detection. Specifically, the method includes the following steps:

[0009] 1) Acquire dual-temporal images of the location to be detected, wherein the dual-temporal images are images of the same area acquired at two different time points;

[0010] 2) Input the acquired dual-temporal images into the trained building change detection network to obtain a predicted map of changed buildings in the dual-temporal images;

[0011] The building change detection network includes encoder feature extraction and decoder feature fusion;

[0012] The encoder feature extraction is used to extract multi-level features of explicit prior architecture from dual-temporal image fusion.

[0013] The decoder feature fusion includes a multi-level feature interaction module, which is used to adaptively select features at each level for fusion to highlight the structural information of the changing building at different scales. The original image size is restored by upsampling the deep features step by step, and edge and corner information is fused at the output head of the decoder to output the predicted map of the changing building.

[0014] 3) Determine the building change results based on the predicted building change map.

[0015] The method of this invention has the following advantages: It integrates explicit prior knowledge of buildings, such as edge and corner features, into the feature extraction and output modeling stages, enhancing the network's optimization ability for regular corners in changing buildings through edge and corner features. It effectively extracts shallow texture, spectral, and structural information, as well as deep semantic information, from dual-temporal remote sensing images during feature extraction. Furthermore, it utilizes deep semantic difference information to guide the location of changing buildings, and leverages multi-level information interaction to achieve multiple uses of deep semantic information and shallow texture and spectral information, thereby fully utilizing change features to alleviate inter-class similarity and intra-class difference problems in building change detection. Finally, edge and corner features are fused again to optimize the semantic segmentation results of the changed buildings output by the network, achieving extraction results that conform to the intuitive perception of surveying professionals.

[0016] Furthermore, the multi-level feature interaction module is used to adaptively select features at each level for fusion to highlight the structural information of changing buildings at different scales, including the following steps:

[0017] Align features from different stages to the current feature size using max pooling or interpolation operations;

[0018] Enhance the correlation between features from different channels by using ECA attention mechanism and adaptive grouping;

[0019] Extract features from each stage, calculate the importance of each stage feature, and select the necessary stage features for aggregation based on the importance.

[0020] The original group channels are superimposed, and the number of feature channels is adjusted to the number of channels in the current feature layer, and the current layer features are output.

[0021] Furthermore, the importance of features at each stage is calculated using the following formula:

[0022]

[0023] Where δ represents the feature layer weights obtained after the activation function, and i is the number of adaptive groups. Further, selecting the necessary stage features for aggregation based on importance includes: when the feature layer weights exceed a threshold, selecting stage features of spatial detail features for aggregation; otherwise, selecting stage features of contextual semantic features for aggregation.

[0024] This invention addresses the issue that high-dimensional features extracted through multiple downsampling steps typically contain high-level semantic information related to changing buildings. This information can often indicate the location of the changing building, but its ability to reconstruct the true shape of the changing building is insufficient. Meanwhile, low-dimensional features contain knowledge about the shape and structure of the changing building, but they cannot provide sufficient contextual information. To solve this problem, this invention employs a multi-level feature interaction module. This module adaptively selects appropriate features (detailed textures, global information) for fusion during the feature fusion stage. This multi-level feature fusion approach promotes the interaction of local and global information, enhancing the reusability of multi-scale feature maps.

[0025] Furthermore, edge and corner information is fused at the decoder's output head location to output a change-building prediction map, including:

[0026] The final stage upsampling features are input into the prediction head to generate change prediction maps R. c Direction prediction diagram R d and edge prediction graph R e ;

[0027] In the edge prediction branch, by analyzing the edge detection map R... e The corner detection Harris operator and Sigmoid activation function are used to extract corner weights, resulting in the edge enhancement map R. e' ;

[0028] In the change prediction graph branch, the change prediction graph R... c Performing multiplication of eight-direction operators yields the feature enhancement maps R of the change prediction map in eight directions. c' ;

[0029] Feature enhancement map R c' The direction prediction map R normalized by the Softmax function d Multiply and combine with edge enhancement map R e' Multiplying them together yields the edge enhancement result R. E ;

[0030] Edge enhancement map R e' Complement and Change Prediction Chart R c Multiply, then add to the edge enhancement result R E Add them together to obtain the final predicted building change map R.

[0031] Furthermore, feature enhancement maps R of the change prediction map in eight directions are obtained. c' And edge enhancement results R E The process is as follows:

[0032]

[0033] R E =R c' ×Softmax(R d )×R e' .

[0034] By employing a dual strategy of semantic segmentation of changing buildings and edge optimization, the output predicted map of changing buildings exhibits more regular edges and corners, conforming to human sensory perception. This method further enhances the network's ability to optimize the edges of changing buildings at the output stage using edge and corner information, collectively improving the accuracy of building change detection.

[0035] Furthermore, the encoder feature extraction for extracting multi-level features of explicit prior knowledge of buildings in dual-temporal image fusion includes:

[0036] The encoder feature extraction consists of a main dual-branch feature extraction network, an auxiliary dual-branch feature extraction network, and a semantic context aggregation module. The main dual-branch feature extraction network is used to extract multi-level features of dual-temporal images, the auxiliary dual-branch feature extraction network is used to extract multi-level features of edge maps of dual-temporal images, and the semantic context aggregation module is used to connect the same-level main and auxiliary branch features of a single temporal image to obtain multi-level features of explicit priors of buildings in dual-temporal image fusion.

[0037] The method of the present invention not only extracts multi-level features from dual-temporal images during feature extraction, but also extracts multi-level features from the edge maps of dual-temporal images. Furthermore, it connects the features of the same-level main dual-branch feature extraction network of a single temporal image with the features of the auxiliary dual-branch feature extraction network to enhance the learning of more image texture information by the main features.

[0038] Furthermore, the edge map of the dual-temporal image is obtained by extracting the edges and contours in the dual-temporal image using the Canny edge extraction algorithm.

[0039] Edge feature extraction aims to extract edge details in a scene. The Canny edge extraction algorithm is used to extract edges and contours from dual-temporal images, which can effectively segment homogeneous objects and enhance the representation ability of edge features during the feature extraction process.

[0040] Furthermore, the semantic context aggregation module connects the main and auxiliary branch features of the same level in a single temporal phase to obtain multi-level features of the explicit prior of the dual-temporal image fusion building, including: edge features are sampled to the same size as the main feature map through downsampling operation, the channels of the two features are reduced to 1 through 1×1 convolutional layers, and then the pixel values ​​of the two features are added together. Then, LeakReLU, 1×1 convolutional layers and sigmoid function are used to retain important features and eliminate irrelevant features. Finally, the edge features and the main network features are fused through feature multiplication and concatenation operations.

[0041] Since edge features contain a lot of noise, fusing them with feature maps from different levels can introduce irrelevant information and cause interference. To address this issue, a Semantic Context Aggregation Module (SCAM) was designed to enrich the edge representation capabilities of the main network using edge information.

[0042] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described in detail below with reference to the accompanying drawings. Attached Figure Description

[0043] Figure 1 The building change detection network structure diagram used in this invention;

[0044] Figure 2 The edge information enhancement output module structure diagram used in this invention;

[0045] Figure 3 The experimental results are shown in the figure obtained by qualitatively comparing the method of the present invention with other methods. Detailed Implementation

[0046] The technical solution of the present invention will be clearly and completely described below with reference to specific embodiments. However, those skilled in the art should understand that the embodiments described below are only for illustrating the present invention and should not be regarded as limiting the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0047] Example of a building change detection method integrating explicit building priors and multi-stage feature aggregation

[0048] This embodiment utilizes the Building Change Detection Network (IBEPNet), which integrates explicit building priors and multi-stage feature aggregation, to detect building changes. Its motivation and innovation mainly come from the following two aspects.

[0049] Explicit Prior Feature Fusion: Explicit prior features of buildings, such as corners and edges, are characteristics that easily distinguish buildings from other features in remote sensing imagery. Introducing these explicit features can effectively improve the accuracy of building change detection. Compared with traditional machine learning-based methods, deep learning has greatly improved the accuracy of building change detection, but its ability to constrain regular building edges is relatively weak, leading to smooth edges of changed buildings in the detection results. Buildings have linear, regular edges and sharp corner features. Utilizing these explicit prior features can optimize the edge segmentation results of changed buildings, ensuring that the detected changed buildings have regular edges and clear angular attributes. The building change detection network proposed in this invention is based on this discovery. It incorporates explicit prior knowledge of building edges and corners into the feature extraction and changed building segmentation process of the original image, which helps to enhance the segmentation ability of changed building edges and corners in the scene, effectively improving the segmentation ability of changed building edges during the building change detection process.

[0050] Feature transformation and fusion mechanisms: Utilizing feature transformation and fusion mechanisms can reduce semantic differences between features at different scales, enhance the reusability of features at different scales, and help improve the accuracy of change detection. In the process of building change detection, shallow features contain more texture, structural, and spectral information, while deep features mainly cover deep semantic information. There is a significant semantic gap between deep and shallow features. Feature transformation and fusion help reduce the semantic differences between deep and shallow features, promote the interaction of multi-level information, indicate the existence of changed buildings through deep features, and then restore the structural and spatial details of the changed buildings through shallow features. Compared with the traditional encoder-decoder network structure, the feature transformation and fusion mechanism can fuse multi-scale features, effectively alleviating the problems of inter-class similarity and intra-class differences in the building change detection process, and helping to detect changed buildings more accurately.

[0051] like Figure 1The diagram shows the architecture of the building change detection network used in this embodiment. IBEPNet (Building Change Detection Network) mainly consists of an encoder feature extraction stage and a decoder feature fusion and supervision stage. In the feature extraction stage, it comprises two dual-branch feature extraction networks. The main branch consists of a dual-branch network with shared weights, used to extract multi-level features, while an auxiliary branch network assists the network in learning more edge knowledge. The same-level main-auxiliary branch features of a single temporal phase are connected by SCAM (Semantic Context Aggregation Module) to enhance the main feature's learning of more image texture information. For the deep semantic information extracted by the network, cascaded operations are used to aggregate and enhance the difference information of the dual-temporal features. Simultaneously, MFSIM (Multi-Level Feature Interaction Module) is used to fuse multi-level difference features, highlighting the structural information of the changing building at different scales. In the decoder, the original image size is restored through progressive upsampling of deep difference features, using skip connections to associate same-level feature maps to supplement high-level semantic information down to low-level spatial information. At the decoder's output, edge and corner information are fused to enhance the output, resulting in a more regular and detailed predicted image of the changing building.

[0052] The method in this embodiment includes the following steps:

[0053] 1) Acquire dual-temporal images of the location to be detected, wherein dual-temporal images are images of the same area acquired at two different time points;

[0054] 2) Input the acquired dual-temporal images into the trained building change detection network to obtain a predicted map of changed buildings in the dual-temporal images;

[0055] The building change detection network includes encoder feature extraction and decoder feature fusion;

[0056] Encoder feature extraction is used to extract multi-level features of explicit prior architecture from dual-temporal image fusion.

[0057] The decoder feature fusion includes a multi-level feature interaction module, which is used to adaptively select features at each level for fusion to highlight the structural information of the changing buildings at different scales. The original image size is restored by upsampling the deep features step by step, and edge and corner information is fused at the output head of the decoder to output the predicted map of the changing buildings.

[0058] 3) Determine the building change results based on the building change prediction map.

[0059] Specifically, this embodiment employs an explicit building feature enhancement module to achieve feature extraction. Due to the progressive downsampling during feature extraction, the size of the feature map decreases, leading to the gradual loss of edge and structural information in the scene. The explicit building feature enhancement module proposed in this invention aims to enhance the network's ability to perceive edge features. By increasing the representation of edge features, it improves the ability to extract edge features during building change detection. This module mainly consists of two branches: ① edge feature extraction, and ② edge feature fusion.

[0060] ① Edge feature extraction.

[0061] Edge feature extraction aims to extract edge details in a scene. The Canny edge extraction algorithm is used to extract edges and contours from dual-temporal images, effectively segmenting homogeneous objects and enhancing the representational power of edge features during the feature extraction process. First, the images are color-converted from RGB images to grayscale images to facilitate gradient information calculation. The conversion method for 24-bit images is as follows:

[0062] Gray=R*0.299+G*0.587+B*0.114 (1)

[0063] In the formula, R, G, and B are the color depth values ​​of the red, green, and blue channels in a true-color image, respectively.

[0064] Gaussian filtering and non-maximum suppression are applied to the grayscale image to remove noise interference with edge detection. To smooth the image, a Gaussian filter is used for convolution, and the calculation method is as follows:

[0065]

[0066] In the above formula, σ is a constant of 1.4, and 2k+1 is the size of the convolution kernel. The size of the Gaussian convolution kernel will affect the edge detection effect. The larger the size, the lower the detector's sensitivity to noise, but the edge detection localization error will also increase. In order to extract edge information in the image more accurately, the convolution kernel size is 3 in this experiment, that is, k is set to 1.

[0067] After normalizing the above results, the pixel value corresponding to the center point (i,j) of the convolution kernel can be expressed as:

[0068]

[0069] Right now:

[0070]

[0071] Where m and n are the offsets in two directions of the convolution kernel, respectively.

[0072] Gradient calculations are performed on the Gaussian-filtered image. The Sobel kernel is used to calculate the first derivatives in the horizontal and vertical directions. The calculated gradient direction is always perpendicular to the edge, and the gradient directions are categorized into four types: vertical, horizontal, and the two diagonals. The calculation formula is as follows:

[0073]

[0074]

[0075] Among them, G x G y These are the first derivatives in the horizontal and vertical directions, respectively.

[0076] Since the Sobel operator detects relatively coarse edges, it is necessary to suppress pixels with insufficient gradients and retain only the maximum gradient to achieve accurate edge detection. In the experiment, nonmaximum suppression and dual thresholding were used to determine the edges. The high threshold and low threshold in the dual thresholding were set to 20 and 100, respectively.

[0077] ②Edge feature fusion.

[0078] Since edge features contain a lot of noise, fusing them with feature maps from different levels can introduce irrelevant information. To address this issue, this embodiment designs a Semantic Context Aggregation Module (SCAM) to enrich the edge representation capabilities of the main network using edge information. The model's input consists of multi-level features and edge features extracted by the main network. Edge features are downsampled to the same size as the main feature map. The two features are then processed through 1×1 convolutional layers to reduce the number of channels to 1, followed by pixel summation. Then, LeakReLU, 1×1 convolutional layers, and the sigmoid function are used to retain important features and eliminate irrelevant ones. Finally, feature multiplication and concatenation operations are used to fuse the edge features and main network features. The calculation details are as follows:

[0079] F es '=Conv1×1(f e )+Conv1×1(f s (8)

[0080] F es ”=Sigmoid(Conv1×1(LeakReLU(F es '))) (9)

[0081] F S =Conv1×1(Cat(f e F es ”,fs ))(10)

[0082] Among them, f e For edge features, f s The main network features are used to extract edge detail features from the scene using the Canny algorithm. The part number is the dot product operation, and Cat() refers to the cascade operation.

[0083] This embodiment uses a multi-scale feature semantic interaction module (MFSIM) to adaptively select appropriate features (detailed textures, global information) for fusion.

[0084] This embodiment considers that high-dimensional features extracted through multiple downsampling operations typically contain high-level semantic information related to the changing building. This information can often indicate the location of the changing building, but its ability to reconstruct the true shape of the changing building is insufficient. Meanwhile, low-dimensional features contain knowledge about the shape and structure of the changing building, but they cannot provide sufficient contextual information. To address this issue, this embodiment proposes a multi-level feature interaction module (MFSIM). MFSIM can adaptively select appropriate features (detailed texture, global information) for fusion. Specifically, MFSIM first aligns features from different stages to the current feature size using max pooling and interpolation operations, i.e., F... c ∈R C×H×W Subsequently, the correlation between features from different channels is enhanced through ECA attention mechanism and adaptive grouping. Adaptive grouping divides the three groups of features into group i. Extract the features of each stage of the i-th group respectively, and calculate the importance of each stage feature using the following formula.

[0085]

[0086] Here, δ represents the feature layer weights obtained after the activation function. If δ is large, lower spatial detail features are prioritized; if δ is small, the network focuses more on the aggregation of contextual semantic features. After the above calculation, the original grouped channels are superimposed, and the number of feature channels is adjusted to the current feature layer channel number. The output current layer features can refine low-level features while aggregating deep semantic information. In this experiment, i is set to 4, that is, the features are divided into 4 groups for multi-level feature interaction.

[0087] This embodiment employs edge information enhancement to make the output of the model's changed buildings have more regular edges and corners, conforming to human sensory perception. Specifically, because the progressive upsampling process loses some spatial detail information, many edge smoothing and blurring phenomena exist in the detected changed buildings. To solve this problem, this embodiment uses an edge feature optimization module, employing a dual strategy of changed building semantic segmentation and edge optimization. The module structure diagram is shown below. Figure 2 As shown:

[0088] Specifically, the last stage upsampling feature F last ∈R 32×256×256 The input is fed into the prediction head to generate change prediction maps R. c Direction prediction diagram R d and edge detection map R e .

[0089] In the edge prediction branch, by adjusting R e Corner weights R are extracted using the Harris operator and Sigmoid activation function for corner detection. e' In Harris corner detection, the sliding kernel size is 3, and the corner response value threshold k is set to 0.04.

[0090] In the change prediction map branch, the edges of buildings in each direction are refined using a combined 8-directional operator. The change prediction map is multiplied by the eight-directional operator, which divides 360° into eight quadrants, each occupying 45 degrees. The weights for the eight directions are: top left, left, bottom left, bottom, bottom right, right, top right, and top. These weights are then multiplied with the direction prediction results to obtain the semantic center inverse direction for each of the eight directions. This inverse direction is then multiplied with the weights for the eight directions, enhancing the features of the building change prediction map in these eight directions. This process can be represented as:

[0091]

[0092] Among them, e i Weights for 8 directions, each e i Both are 3×3 matrices, with the weight in the top left direction being... The weight of the upward direction is The weights for the other six directions follow the same pattern. Then, R... c' The direction prediction map R normalized by the Softmax function d Multiply by the image and then multiply by the edge enhancement map to obtain the edge enhancement result R. E This can be described as the following process:

[0093] R E =R c' ×Softmax(R d)×R e' (15)

[0094] Then, calculate R. e' The complement, i.e., 1-R e' Subsequently with R c Multiplication enhances the semantic segmentation capabilities of varied architectures. This is then combined with R... E The summations yield the final enhanced building change detection map R, which enhances the detection capability of changing buildings from both the edges and interiors, thereby increasing the precision of the method. In other words:

[0095] R C =R c ×(1-R e' (16)

[0096] R = R E +R C (17)

[0097] By analyzing the edge detection map R ed The Harris corner detector is used to generate corner points in the prediction map. In the Harris corner detection, the size of the sliding kernel is 3, and the corner response value threshold k is set to 0.04.

[0098] The detected corner points are then normalized by sequentially applying the Sigmoid function and adding it to the original edge detection result, thereby enhancing the weight of the corner points on the edge.

[0099] To address the imbalance problem where the number of pixels occupied by changing buildings is much smaller than that of unchanged areas, and the insufficient boundary optimization capability, this invention uses a hybrid loss function to address both the classification imbalance problem and the edge optimization problem of changing buildings. This loss function consists of three parts: focal loss, structural similarity loss, and L... SSIM and edge loss L edge Among them, the focus loss L focal This is used to address class imbalance problems; structural similarity loss focuses more on structural information; edge loss considers building edge features by calculating the root mean square error between the predicted segment and the true boundary. Specifically:

[0100] L total =aL focal +bL SSIM +cL edge (18)

[0101] Where a, b, and c are coefficients for the three losses, used to adjust the weights of different losses; in this embodiment, they are set to 1, 0.5, and 0.5, respectively. SSIM L edge and L focal They are defined as follows:

[0102]

[0103] In this context, α and γ are used to adjust the weights for positive and negative classifications and reduce the loss of easily classified samples, making the network pay more attention to difficult-to-classify samples. α and γ are set to 0.25 and 2, respectively, p is the probability that the model predicts a positive case, and y is the sample label value.

[0104] SSIM initially evaluates image inpainting quality by capturing structural information from two images. To better evaluate the global structural information of predicted and labeled instances with varying structures, this embodiment incorporates this into the loss function to learn the structural information of the labels. Let X = {x i i = 1, 2, ..., n 2} and Y = {y i i = 1, 2, ..., n 2 Let} be two n×n cropping blocks, whose pixel values ​​come from the predicted probability map and the labeled map GT, respectively. The relationship between X and Y can be defined as:

[0105]

[0106] Where, μ x μ y and σ x σ y These are the mean and standard deviation of x and y, respectively, and σ xy It is the covariance, C1=C2=0.001 2 Used to avoid division by zero.

[0107] This embodiment utilizes the Laplacian bidirectional line detection operator to extract the edge features of P and GT respectively, in order to reduce the edge difference between P and GT. The bidirectional detection operators are respectively and The root mean square error is used to calculate the difference between the two, thereby optimizing the P-edge. Specifically, it can be expressed as follows:

[0108]

[0109] in and These are the labeled horizontal and vertical edges, as well as the predicted horizontal and vertical edge maps. After extracting the edges using two directional operators, the edge differences are calculated using the root mean square error.

[0110] The method in this embodiment has significant advantages over other methods:

[0111] 1) Experimental data.

[0112] The LEVIR-CD and CD_GZ datasets were selected as the comparative datasets for this invention. The LEVIR-CD dataset includes images from different seasons between 2002 and 2018, covering different building types such as residential and commercial areas, which can effectively verify the robustness of the model in detecting false changes caused by seasonal differences in imaging. The CD_GZ dataset images were taken in Guangzhou, China. Compared with the LEVIR-CD dataset, the shapes and sizes of buildings in this dataset vary greatly, and the changes in buildings are complex and diverse. At the same time, high-rise buildings have greater displacement due to viewpoint projection, which poses a great challenge to the accurate detection of changing buildings. In the experiment, data augmentation operations were performed on both datasets, including flipping, image rotation, random cropping, and noise addition. Each image patch was cropped to a size of 256×256 pixels, and the training, validation, and test sets were divided in a 7:1:2 ratio. The final number of images in the training, validation, and test sets of the LEVIR-CD dataset was 7120, 1024, and 2028, respectively, while the number of images in the training, validation, and test sets of the CD_GZ dataset was 8463, 1209, and 2418, respectively.

[0113] 2) Comparison method.

[0114] The proposed encoder and decoder of IBEPNet (Building Change Detection Network) are applicable to various encoder-decoder-based two-branch Siamese networks similar to UNet. To verify the effectiveness and robustness of the proposed method, and to broadly validate its effectiveness, this embodiment categorizes the comparison networks into three classes: ① Traditional fully convolutional neural network (CNN) based methods: FC-Siam-Diff, FC-Siam-Conc, MSCANet, RDPNet, GeSANet, and CGNet; ② Transformer based methods: BIT; and CNN and Transformer combined methods: ACABFNet and MSCANet.

[0115] In the experiments, the proposed IBEPNet was implemented in the PyTorch framework and trained on an NVIDIA GeForce RTX 3090. The model optimizer used in the experiments was AdamW, with an initial learning rate of 5e-4, weight decay of 0.0025, and the learning rate adjustment strategy employed the Cosine Annealing Warm Restarts algorithm, where the parameter T0 was set to 15, and T... mult The value is set to 2. The method proposed in this embodiment was trained for 200 epochs on the LEVIR-CD and CD_GZ datasets respectively, with a batch size of 8. The network eventually reached convergence.

[0116] 3) Evaluation indicators.

[0117] To evaluate the accuracy of building change detection, this embodiment uses several common quantitative evaluation metrics, including precision (Pre), recall, F1 score (F1), and overall accuracy (OA) as evaluation standards. Furthermore, it utilizes the intersection-over-union (IoU) and mean intersection-over-union (mIoU) metrics commonly used in semantic segmentation to evaluate the network's performance. The semantic segmentation ability of the model is evaluated using the IoU of the changed buildings and the overall average mIoU between the changed buildings and the unchanged areas; the higher both are, the better the model's performance. These metrics are defined as follows:

[0118]

[0119]

[0120] Where TP represents the number of true positives, FP represents the number of false positives, FN represents the number of false negatives, and TN represents the number of true negatives.

[0121] 4) Quantitative comparison experiment.

[0122] This section compares the results of different methods with those of IBEPNet. Table 1 quantitatively describes the performance of several methods in building change detection on the LEVIR-CD dataset. The network proposed in this embodiment achieves the best detection accuracy across all five accuracy metrics. The model in this embodiment also has the highest IoU and MIoU scores. Since these values ​​are related to the number of correctly detected true change pixels, it is speculated that this may be due to the addition of boundary information to the network, which leads to the accurate detection of corner pixels of changed buildings.

[0123] Table 1 shows the comparison results of the proposed method with other methods on the LEVIR-CD dataset.

[0124]

[0125] Table 2 quantitatively describes the performance of several methods in building change detection on the CD_GZ dataset. Overall, the method in this embodiment achieves the best results across the four accuracy metrics. However, the Pre and Recall scores differ significantly among the models, indicating a serious problem of missed detections in each model, meaning that the ability of each method to identify real changes needs improvement. The model in this embodiment improves the absolute accuracy of Recall by more than 5%, demonstrating high robustness even under low resolution and non-orthophoto imaging conditions.

[0126] Table 2 shows the comparison results of the proposed method with other methods on the CD-GZ dataset.

[0127]

[0128] 5) Qualitative comparative experiment.

[0129] like Figure 3 As shown, the results of different methods are visually compared with those of IBEPNet. Through this comparison on the two datasets, the proposed method demonstrates better edge optimization capabilities and fewer corner errors, indicating that corner information is effectively utilized and enhanced. Furthermore, regarding edge optimization of changing buildings, the proposed method produces more regular boundaries for the detected changing buildings, better conforming to human visual perception.

[0130] This invention presents a building change detection method that integrates explicit prior information about buildings and multi-stage feature aggregation. In the feature extraction stage, this method enhances the network's ability to extract edge information by introducing explicit prior information about building edges. Simultaneously, in the feature fusion stage, multi-level feature fusion is used to promote the interaction of local and global information, enhancing the reusability of multi-scale feature maps. At the output end, this method further utilizes edge and corner information to enhance the network's optimization ability for changing building edges, collectively improving the accuracy of building change detection.

[0131] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection claimed by the present invention.

Claims

1. A method for detecting building changes that integrates explicit building priors and multi-stage feature aggregation, characterized in that, Includes the following steps: 1) Acquire dual-temporal images of the location to be detected, wherein the dual-temporal images are images of the same area acquired at two different time points; 2) Input the acquired dual-temporal images into the trained building change detection network to obtain a predicted map of changed buildings in the dual-temporal images; The building change detection network includes encoder feature extraction and decoder feature fusion; The encoder feature extraction is used to extract multi-level features of explicit prior architecture from dual-temporal image fusion. The decoder feature fusion includes a multi-level feature interaction module, which is used to adaptively select features at each level for fusion to highlight the structural information of the changing building at different scales. The original image size is restored by upsampling the deep features step by step, and edge and corner information is fused at the output head of the decoder to output the predicted map of the changing building. 3) Determine the building change results based on the predicted building change map.

2. The building change detection method according to claim 1, characterized in that, The multi-level feature interaction module is used to adaptively select features at each level for fusion to highlight the structural information of changing buildings at different scales, including the following steps: Align features from different stages to the current feature size using max pooling or interpolation operations; Enhance the correlation between features from different channels by using ECA attention mechanism and adaptive grouping; Extract features from each stage, calculate the importance of each stage feature, and select the necessary stage features for aggregation based on the importance. The original group channels are superimposed, and the number of feature channels is adjusted to the number of channels in the current feature layer, and the current layer features are output.

3. The building change detection method according to claim 2, characterized in that, The importance of features at each stage is calculated using the following formula: Where δ represents the feature layer weights obtained after the activation function, and i is the number of adaptive groupings.

4. The building change detection method according to claim 3, characterized in that, The selection of the required stage features for aggregation based on importance includes: when the weight of the feature layer exceeds the threshold, the stage features of spatial detail features are selected for aggregation; otherwise, the stage features of contextual semantic features are selected for aggregation.

5. The building change detection method according to claim 4, characterized in that, The edge and corner information is fused at the decoder's output head to output a change-building prediction map, including: The final stage upsampling features are input into the prediction head to generate change prediction maps R. c Direction prediction diagram R d and edge prediction graph R e ; In the edge prediction branch, by analyzing the edge detection map R... e The corner detection Harris operator and Sigmoid activation function are used to extract corner weights, resulting in the edge enhancement map R. e' ; In the change prediction graph branch, the change prediction graph R... c Performing multiplication of eight-direction operators yields the feature enhancement maps R of the change prediction map in eight directions. c' ; Feature enhancement map R c' The direction prediction map R normalized by the Softmax function d Multiply and combine with edge enhancement map R e' Multiplying them yields the edge enhancement result R. E ; Edge enhancement map R e' Complement and Change Prediction Chart R c Multiply, then combine with the edge enhancement result R E Add them together to obtain the final predicted building change map R.

6. The building change detection method according to claim 5, characterized in that, The feature enhancement maps R of the change prediction map in 8 directions are obtained. c' And edge enhancement results R E The process is as follows: R E =R c' ×Softmax(R d )×R e' 。 7. The building change detection method according to claim 1, characterized in that, During the stepwise upsampling of deep features to restore the original image size, skip connections are used to associate sibling feature maps to supplement high-level semantic information with low-level spatial information.

8. The building change detection method according to claim 1, characterized in that, The encoder feature extraction is used to extract multi-level features of explicit architectural priors for dual-temporal image fusion, including: The encoder feature extraction consists of a main dual-branch feature extraction network, an auxiliary dual-branch feature extraction network, and a semantic context aggregation module. The main dual-branch feature extraction network is used to extract multi-level features of dual-temporal images, the auxiliary dual-branch feature extraction network is used to extract multi-level features of edge maps of dual-temporal images, and the semantic context aggregation module is used to connect the same-level main and auxiliary branch features of a single temporal image to obtain multi-level features of explicit priors of buildings in dual-temporal image fusion.

9. The building change detection method according to claim 8, characterized in that, The edge map of the dual-temporal image is obtained by extracting the edges and contours in the dual-temporal image using the Canny edge extraction algorithm.

10. The building change detection method according to claim 9, characterized in that, The semantic context aggregation module connects the main and auxiliary branch features of the same level in a single temporal phase to obtain the multi-level features of the explicit prior of the building in dual-temporal image fusion. These features include: edge features are sampled to the same size as the main feature map through downsampling operations; the two features are reduced to 1 channel through 1×1 convolutional layers; the pixel values ​​of the two features are then added; and then LeakReLU, 1×1 convolutional layers, and the Sigmoid function are used to retain important features and eliminate irrelevant features; finally, the edge features and the main network features are fused through feature multiplication and concatenation operations.

Citation Information

Cited By

  • Remote sensing image classification method and system based on change perception and space-time fusion

    CN121505370A

  • A remote sensing image classification method and system based on change perception and spatio-temporal fusion

    CN121505370B