Building change detection method and system based on twinborn change attention residual network

By adopting a twin change attention residual network based on ResNet and UNet in building change detection, the problem of insufficient detection efficiency and accuracy of traditional methods in complex environments is solved, and more efficient and accurate change detection is achieved.

CN120147249AInactive Publication Date: 2025-06-13AEROSPACE INFORMATION RES INST CAS

Patent Information

Application Number
CN202510211686.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional building change detection methods are difficult to adapt to rapid changes in complex urban environments, and deep learning methods require a large amount of labeled data and have poor interpretability, resulting in insufficient detection efficiency and accuracy.

Method used

A twin change attention residual network based on ResNet and UNet is adopted, and a twin structure and change attention residual module with shared weights are used to process remote sensing images in combination with geometric registration and radiation correction to realize building change detection.

Benefits of technology

It improves the accuracy and efficiency of building change detection, reduces the amount of parameters, enhances the sensitivity to changing areas, and supports pre-training and multi-task learning of building semantic segmentation samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147249A_ABST
    Figure CN120147249A_ABST
Patent Text Reader

Abstract

The invention provides a building change detection method and system based on a twinborn change attention residual network. The method comprises the following steps: acquiring original remote sensing images of two stages of research areas; performing geometric registration and radiation correction on the original remote sensing image in sequence to obtain two time-phase feature maps; and inputting the two time phase characteristic graphs into a twinborn change attention residual network based on ResNet and UNet to obtain a change probability graph. The scheme provided by the invention has higher robustness and accuracy in a building change detection task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of high-precision change detection of buildings in remote sensing images, and particularly relates to a method and a system for building change detection based on a twin change attention residual network. Background Art

[0002] With the acceleration of the urbanization process, the changes of buildings are frequent and rapid, and traditional change detection methods often struggle to adapt to this change. High-resolution remote sensing images provide rich spatial information. However, due to the diverse shapes of buildings, complex environments, and the influence of factors such as lighting and weather during image acquisition, the difficulty of automatically extracting building changes has increased significantly, and a large amount of noise and redundant information (such as lighting changes, shadows, cloud interference, etc.) has been introduced, increasing the difficulty of change detection. At the same time, the acceleration of the urbanization process requires change detection methods to have a higher degree of automation and real-time performance. Therefore, developing an efficient and accurate building change detection method that can make full use of the information in high-resolution remote sensing images to timely identify and analyze building changes is not only of great significance for urban management, planning, and environmental monitoring but also provides new technical means for research in related fields.

[0003] With the application of high-resolution remote sensing images, traditional change detection methods have shown deficiencies in complex urban environments. Existing solutions mainly include pixel-based change detection, object-based change detection, and deep learning methods. Pixel-based detection relies on pixel value differences and is easily affected by noise and lighting changes; object-based methods improve the adaptability through image segmentation and feature comparison, but have a high computational complexity; while deep learning techniques, such as convolutional neural networks, although able to automatically extract features and achieve high-precision detection, require a large amount of labeled data and have poor interpretability.

[0004] Traditionally, the change area is extracted by an object-oriented segmentation method, and then a neural network is used to determine whether each change block is a building change block to output a pixel-level building change detection result. Traditionally, a deep learning semantic segmentation network is used to extract buildings from two-phase images respectively, and then the extraction results are combined with simulated samples to train a building change detection network to achieve building change detection. Such two-stage methods have problems of slow speed and error accumulation. The end-to-end building change detection network takes two or more phases of images as input and directly outputs a change map, avoiding the problem of error accumulation and being faster and more efficient than the two-stage method.

[0005] The most representative in the single-stage deep learning change detection network is the Siamese network. The Siamese network is a convolutional neural network with a dual-path weight sharing structure and has been widely used in the field of image change detection based on neural networks. It is proposed to use a shared weight network combined with a contrastive loss function to achieve change detection. The VGG network is used to mine feature maps, and then the PCA transformation is performed on the multi-temporal feature maps to mine change information. The Onera Satellite Change Detection (OSCD) dataset is made based on Sentinel-2 satellite images, and a convolutional neural network based on early fusion and Siamese network is proposed for change detection of image patches. PPCNET is proposed, and a change block branch is introduced on the basis of the Siamese network. The PGA-Siames network is proposed, and a change residual module and a collaborative attention module are introduced into the Siamese network, improving the WHU building change detection accuracy.

[0006] The above methods mainly focus on modifying the network structure based on the Siamese network to improve the network's learning ability for change detection tasks, without considering improving change detection accuracy through sample data augmentation, nor making full use of building segmentation samples to train the network. Summary of the Invention

[0007] To solve the above technical problems, the present invention proposes a technical solution for a building change detection method based on a Siamese change attention residual network to solve the above technical problems.

[0008] The first aspect of the present invention discloses a building change detection method based on a Siamese change attention residual network, and the method includes:

[0009] Step S1, obtaining the original remote sensing images of two periods of the research area;

[0010] Step S2, successively performing geometric registration and radiometric correction on the original remote sensing images to obtain two-temporal feature maps;

[0011] Step S3, inputting the two-temporal feature maps into a Siamese change attention residual network based on ResNet and UNet to obtain a change probability map.

[0012] According to the method of the first aspect of the present invention, in the step S3, the two branches of the Siamese change attention residual network based on ResNet and UNet share weights.

[0013] According to the method of the first aspect of the present invention, in the step S3, the inputting the two-temporal feature maps into a Siamese change attention residual network based on ResNet and UNet to obtain a change probability map includes:

[0014] Input the first image of the two - phase feature map into the first branch of the Siamese change attention residual network based on ResNet and UNet to obtain five first - fused feature maps of different scales, specifically including:

[0015] Input the first image of the two - phase feature map into a network with the first 5 - layer encoding layer of ResNet34 as the network backbone for downsampling to obtain four downsampling intermediate result feature maps and a downsampling result semantic feature map; pass the downsampling result semantic feature map through a convolution once to obtain a convolutional semantic feature map;

[0016] Input the convolutional semantic feature map into a network of the UNet high - low layer fusion structure based on 5 - layer bilinear interpolation upsampling + convolution method for upsampling to obtain four upsampling intermediate result feature maps and a segmentation probability map, that is, the extracted building probability map;

[0017] Perform high - low layer feature map fusion on the first image and the four downsampling intermediate result feature maps and the four upsampling intermediate result feature maps and the segmentation probability map to obtain five first - fused feature maps of different scales;

[0018] Input the second image of the two - phase feature map into the first branch of the Siamese change attention residual network based on ResNet and UNet to obtain five second - fused feature maps of different scales;

[0019] Input the five first - fused feature maps of different scales and the five second - fused feature maps of different scales into the change attention residual structure to obtain five third - fused feature maps and four fourth - fused feature maps of different scales;

[0020] Add the five third - fused feature maps of different scales and the five fourth - fused feature maps of different scales to obtain five fifth - fused feature maps of different scales;

[0021] Stitch the five fifth - fused feature maps of different scales to obtain a change probability map.

[0022] According to the method of the first aspect of the present invention, in step S3, the step of inputting the first image of the two - phase feature map into a network with the first 5 - layer encoding layer of ResNet34 as the network backbone for downsampling to obtain four downsampling intermediate result feature maps and a downsampling result semantic feature map includes:

[0023] Perform a 7×7 convolution with a stride of 2 on the first figure, reducing the size of the original image by half and increasing the number of channels from 3 to 64; then pass through two residual connection layers, each time reducing the size of the image by half; finally, pass through two more residual connection layers, each time reducing the size of the image by half and doubling the number of channels, ultimately obtaining a downsampled result semantic feature map with a size of 1 / 32 and 512 channels for the first figure, as well as four downsampled intermediate result feature maps with different scales and numbers of channels.

[0024] According to the method of the first aspect of the present invention, in step S3, the network of the UNet high-low layer fusion structure with bilinear interpolation upsampling + convolution includes:

[0025] Apply bilinear interpolation upsampling + convolution to replace the transposed convolution in the original UNet high-low layer fusion structure network.

[0026] According to the method of the first aspect of the present invention, in step S3, inputting the five first fusion feature maps and second fusion feature maps with different scales into the variable attention residual structure to obtain five third fusion feature maps and fourth fusion feature maps with different scales includes:

[0027] The variable attention residual structure uses the difference between the five first fusion feature maps and second fusion feature maps with different scales as the variable residual path and the fusion of the five first fusion feature maps and second fusion feature maps with different scales as the variable feature path to obtain five third fusion feature maps and fourth fusion feature maps with different scales.

[0028] According to the method of the first aspect of the present invention, in step S3, the variable feature path:

[0029] First, the five first fusion feature maps and second fusion feature maps with different scales are respectively concatenated and fused by channels, and then a 3x3 convolution is performed on the concatenated feature map to obtain five fourth fusion feature maps with different scales and the same number of channels as before concatenation.

[0030] The second aspect of the present invention discloses a building change detection system based on a twin variable attention residual network, and the system includes:

[0031] A first processing module, configured to obtain the original remote sensing images of two study areas;

[0032] A second processing module, configured to perform geometric registration and radiometric correction on the original remote sensing images in sequence to obtain two-temporal feature maps;

[0033] A third processing module, configured to input the two-temporal feature maps into a twin variable attention residual network based on ResNet and UNet to obtain a change probability map.

[0034] In a third aspect of the present invention, an electronic device is disclosed. The electronic device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the steps in a building change detection method based on a twin change attention residual network according to any one of the first aspects of the present disclosure are implemented.

[0035] In a fourth aspect of the present invention, a computer-readable storage medium is disclosed. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps in a building change detection method based on a twin change attention residual network according to any one of the first aspects of the present disclosure are implemented.

[0036] In summary, for the solution proposed by the present invention, first, by combining the residual connection of ResNet and the high and low level feature fusion of UNet, the deep feature extraction ability of ResNet and the multi-level feature fusion advantage of UNet are fully utilized. Moreover, the residual connection effectively alleviates the gradient vanishing problem of the deep network, ensuring the stable transmission of feature information. The skip connection mechanism of UNet improves the richness and accuracy of feature representation by fusing low-level detail information and high-level semantic information. Second, the twin structure with shared weights makes the network process the input image pair more consistently, avoids feature deviation caused by inconsistent weights, greatly reduces the number of parameters, enhances the sensitivity to the changed area, and thus improves the accuracy of change detection. In addition, the change attention residual module is introduced, and by combining the change residual path and the change feature path, the detection ability for the changed area is enhanced. Description of the Drawings

[0037] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0038] Figure 1 It is a flowchart of a building change detection method based on a twin change attention residual network according to an embodiment of the present invention;

[0039] Figure 2 It is a structural diagram of a building change detection system based on a twin change attention residual network according to an embodiment of the present invention;

[0040] Figure 3 It is a structural diagram of an electronic device according to an embodiment of the present invention. Detailed Embodiments

[0041] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0042] The first aspect of the present invention discloses a building change detection method based on a twin change attention residual network. Figure 1 As shown in the flowchart of a building change detection method based on a twin change attention residual network according to an embodiment of the present invention, Figure 1 as shown, the method includes:

[0043] Step S1, obtaining the original remote sensing images of two periods of the research area;

[0044] Step S2, sequentially performing geometric registration and radiometric correction on the original remote sensing images to obtain two-temporal feature maps;

[0045] Step S3, inputting the two-temporal feature maps into a twin change attention residual network based on ResNet and UNet to obtain a change probability map.

[0046] In step S2, the original remote sensing images are sequentially subjected to geometric registration and radiometric correction to obtain two-temporal feature maps.

[0047] Specifically, the original remote sensing images cannot be directly used. There are not only some problems such as geometric errors, radiometric errors, and dull brightness, but also problems such as the unification of multi-source data. Remote sensing image preprocessing generally refers to the process of performing a series of processing and correction on the original remote sensing data before analyzing and applying the remote sensing images. Its main purpose is to improve the quality, accuracy, and applicability of the images, etc. The original remote sensing images are sequentially subjected to processing such as geometric registration and radiometric correction to obtain available two-temporal feature maps. The geometric registration method is implemented by manually selecting control points and using the quadratic polynomial correction method. The geometric registration of the two-period images is completed by manually selecting control points and using the quadratic polynomial correction method, and the offset is controlled within one pixel. First, the polynomial coefficients are calculated using the control points between the image to be registered and the reference image; the polynomial coefficient calculation formulas are shown in Equations (1) and 4.2, where (x src , y src ) are the coordinates of the image to be registered, (x ref , y ref ) are the coordinates of the reference image, (A 1 , B 1 , C 1, D 1 , E 1 , F 1 ), (A 2 , B 2 , C 2 , D 2 , E 2 , F 2 ) are the coefficients of the quadratic polynomial to be solved.

[0048]

[0049] By selecting corresponding control points on the image to be registered and the reference image, several groups of (x src , y src ) and (x ref , y ref ) point pairs can be obtained. Since there are 12 unknowns to be solved and two sets of equations can be listed for a pair of corresponding points, at least 6 pairs of control points are required to calculate the polynomial coefficients. Usually, more than 6 pairs of control points are collected to achieve redundant observations and improve the calculation accuracy.

[0050] After the polynomial is solved, the position of the corresponding point on the pre-registration image can be calculated according to the coordinates of each registered image point. Let the registered coordinates be (x new , y new ). Just substitute these coordinates into (x ref , y ref ) in equations (1) and (2) to obtain the position on the pre-registration image.

[0051] The coordinates calculated in the previous step are often not integers, while the coordinates of the original image are discrete integer values. At this time, interpolation methods need to be used to sample the pixel values at these non-integer coordinates. Bilinear interpolation is an interpolation method widely used in the field of image processing. The specific calculation process is as follows.

[0052] Round the coordinate values x and y up and down to obtain the four corner coordinates:

[0053] (x min , y min ), (x min , y max ), (x max , y min ), (x max , y max )

[0054] For example, the four corner coordinates of (11.3, 12.4) are (11, 12), (11, 13), (12, 12), (12, 13)

[0055] Interpolate linearly twice in the X-axis direction according to the values on the four corner coordinates to obtain the values of the two points (x, y min ), (x, y max ).

[0056] Interpolate linearly in the Y-axis direction using the values of the two points (x, y min ), (x, y max ) to obtain the value of the point (x, y).

[0057] The above process involves two linear interpolations, so it is also called bilinear interpolation.

[0058] The calculation formula for linear interpolation is shown in Equation (3), which shows the steps of linearly interpolating the point (x, y mim , y min ), (x max , y min ) to the point (x, y min ).

[0059]

[0060] Since the four interpolation points in this article are obtained by rounding up and down the x and y coordinates, we have y max = y min + 1, x max = x min + 1, that is, y max - y min = 1, x max - x min = 1. According to this relationship, simplify Equation (3) and combine the processes of two linear interpolations in the X-axis direction and one linear interpolation in the Y-axis direction to obtain the calculation formula for interpolating the pixel value at the (x, y) coordinate position using the four corner coordinates, as shown in Equation (4).

[0061] f(x, y) = (1 + y min - y)(1 + x min - x)f(x min y min ) +

[0062] (y - y min )(1 + x min - x)f(x min , y max ) +

[0063] (1 + y min - y)(x - x min )f(x max , y min ) +

[0064] (y - y min)(x - x min )f(x max , y max )(4)

[0065] Histogram matching is used for radiation correction. The basic idea is based on the gray - level distribution of the reference image, and each pixel in the image to be matched is subjected to gray - level mapping so that the gray - level distribution of the new image is basically the same as that of the reference image. The specific calculation steps are as follows:

[0066] Input: Image I to be matched f , reference image I g

[0067] Output: Image I after histogram matching matched

[0068]

[0069]

[0070] In step S3, the two - phase feature maps are input into a Siamese change attention residual network based on ResNet and UNet to obtain a change probability map. Different from ordinary convolutional neural networks that take a single image as input, this structure takes a pair of images as input. The input image pair is gradually passed from the lower layer to the higher layer of the network in parallel. Since the parameters such as convolutional layers used in the two parallel network paths are shared, it is called a Siamese network structure. The Siamese change attention residual network adopts the structure of ResNet in the down - sampling part and the high - low layer fusion structure of UNet in the up - sampling part.

[0071] In some embodiments, in step S3, the two branches of the Siamese change attention residual network based on ResNet and UNet share weights.

[0072] The inputting the two - phase feature maps into a Siamese change attention residual network based on ResNet and UNet to obtain a change probability map includes:

[0073] Input the first image of the two - phase feature maps into the first branch of the Siamese change attention residual network based on ResNet and UNet to obtain five first - fused feature maps of different scales, specifically including:

[0074] Input the first image of the two-temporal-phase feature maps into a network with the first 5 encoding layers of ResNet34 as the network backbone for downsampling to obtain four downsampling intermediate result feature maps and a downsampling result semantic feature map; separately perform convolution on the downsampling result semantic feature map once to change the number of channels to 320 while keeping the scale size unchanged, concentrate useful feature information, eliminate redundant feature information, and at the same time reduce model parameters to obtain a convolutional semantic feature map;

[0075] The first layer of ResNet34 is an initial convolutional layer, and the latter 4 layers are residual connection layers. Each residual connection layer consists of multiple residual units, and each residual unit contains two 3×3 convolutional layers and a shortcut connection. The introduced residual connection shortcut connection solves the problems of gradient disappearance and network degradation in deep networks, enabling the network to be effectively trained to deeper levels.

[0076] Input the convolutional semantic feature map into a network of the UNet high-low layer fusion structure based on 5-layer bilinear interpolation upsampling + convolution method for upsampling, gradually changing the number of channels from the initial 320 to 160, 96, 64, 48, and 1 in turn for four upsampling intermediate result feature maps and a segmentation probability map, that is, the extracted building probability map;

[0077] Fuse the high and low layer feature maps of the first image and the four downsampling intermediate result feature maps and the four upsampling intermediate result feature maps and the segmentation probability map to obtain five different-scale first fusion feature maps of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32;

[0078] Input the second image of the two-temporal-phase feature maps into the first branch of the Siamese change attention residual network based on ResNet and UNet to obtain five different-scale second fusion feature maps;

[0079] Input the five different-scale first fusion feature maps and the second fusion feature maps into the change attention residual structure to obtain five different-scale third fusion feature maps and fourth fusion feature maps;

[0080] Add the five different-scale third fusion feature maps and the five different-scale fourth fusion feature maps to obtain five different-scale fifth fusion feature maps of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32;

[0081] Stitch the five different-scale fifth fusion feature maps to obtain a change probability map.

[0082] The step of inputting the first image of the two-temporal-phase feature maps into a network with the first 5 encoding layers of ResNet34 as the network backbone for downsampling to obtain four downsampling intermediate result feature maps and a downsampling result semantic feature map includes:

[0083] Perform a 7×7 convolution with a stride of 2 on the first figure, reducing the size of the original image by half and increasing the number of channels from 3 to 64; then pass through two residual connection layers, each time reducing the size of the image by half; finally, pass through two more residual connection layers, each time reducing the size of the image by half and doubling the number of channels, ultimately obtaining the downsampling result semantic feature map of the first figure with a size of 1 / 32 and 512 channels, as well as four downsampling intermediate result feature maps with different scales and numbers of channels. After the feature map size is halved, it is possible to make the feature map channels deeper while occupying the same video memory, and the deepening of the number of channels is beneficial for the network to learn more rich semantic information features. This set of semantic feature maps contains rich building feature information of different scales.

[0084] The network of the UNet high-low layer fusion structure with bilinear interpolation upsampling + convolution includes:

[0085] Apply bilinear interpolation upsampling + convolution to replace the transposed convolution in the original UNet high-low layer fusion structure network. Because in actual use, it is found that the upsampling implemented by the transposed convolution often causes the final segmentation result to show a significant patch effect. The method of first using bilinear interpolation to expand the feature map size by 2 times and then using a conventional convolution to convolve the expanded feature map can well avoid the patch effect. Therefore, in this embodiment, this method is used to replace the transposed convolution scheme of the original UNet high-low layer fusion structure in the network upsampling stage.

[0086] Input the five first fusion feature maps and second fusion feature maps of different scales into the variable attention residual structure to obtain five third fusion feature maps and fourth fusion feature maps of different scales, including:

[0087] The variable attention residual structure uses the difference between the five first fusion feature maps and second fusion feature maps of different scales as the variable residual path and the fusion of the five first fusion feature maps and second fusion feature maps of different scales as the variable feature path to obtain five third fusion feature maps and fourth fusion feature maps of different scales.

[0088] The variable feature path:

[0089] First, the five first fusion feature maps and second fusion feature maps of different scales are respectively concatenated and fused by channel, and then the concatenated feature map is convolved with a 3x3 convolution to obtain five fourth fusion feature maps of different scales with the same number of channels as before concatenation.

[0090] Compared with the change residual structure, the change attention residual structure of this embodiment has the following improvements: (1) The fusion of two-phase feature maps is achieved by means of convolution after channel-wise splicing of feature maps. Compared with the direct addition method of the original change residual structure, this method allows the network to learn richer feature map fusion patterns; (2) The channel attention mechanism is used to weight the fused feature maps, enabling the network to assign higher weights to important features.

[0091] Splicing the five fifth fusion feature maps of different scales to obtain a change probability map includes:

[0092] The five fifth fusion feature maps of different scales are successively upsampled from small to large and then spliced by channel, and finally a single-channel feature map with the same resolution as the original image is obtained. This feature map passes through the Sigmoid activation function to obtain a change probability map with a value range of (0, 1). The original UNet network backbone part can continue to output the semantic segmentation map of the building, enabling the network structure in this paper to support the pre-training of building semantic segmentation samples and multi-task learning.

[0093] In summary, for the solution proposed by the present invention, first, by combining the residual connection of ResNet and the high-low layer feature fusion of UNet, the deep feature extraction ability of ResNet and the multi-level feature fusion advantage of UNet are fully utilized. Moreover, the residual connection effectively alleviates the gradient disappearance problem of the deep network, ensuring the stable transmission of feature information; while the skip connection mechanism of UNet improves the richness and accuracy of feature representation by fusing low-level detail information and high-level semantic information. Second, the twin structure with shared weights makes the network's processing of input image pairs more consistent, avoids feature deviation caused by inconsistent weights, greatly reduces the number of parameters, and enhances the sensitivity to the changed regions, thereby improving the accuracy of change detection. In addition, the change attention residual module is introduced, and the detection ability for the changed regions is enhanced through the combination of the change residual path and the change feature path. In addition to improving the twin network structure, it is also considered to fully utilize the by-product in the process of making building change detection samples: "building segmentation samples" to improve the network detection ability through pre-training with segmentation samples, and to simulate building change detection samples with building segmentation samples to expand the change detection data set, so as to achieve the ability to improve the accuracy.

[0094] The second aspect of the present invention discloses a building change detection system based on a twin change attention residual network. Figure 2 It is a structural diagram of a building change detection system based on a twin change attention residual network according to an embodiment of the present invention; as Figure 2 shown, the system 100 includes:

[0095] The first processing module is configured to obtain the original remote sensing images of the study area in two phases;

[0096] The second processing module is configured to perform geometric registration and radiometric correction on the original remote sensing images in sequence to obtain two-temporal feature maps;

[0097] The third processing module is configured to input the two-temporal feature maps into a Siamese change attention residual network based on ResNet and UNet to obtain a change probability map.

[0098] According to the system of the second aspect of the present invention, the second processing module 102 is specifically configured as follows. The original remote sensing images cannot be directly used. There are not only some problems such as geometric errors, radiometric errors, and dull brightness, but also problems such as the unification of multi-source data. Remote sensing image preprocessing usually refers to the process of performing a series of processing and correction on the original remote sensing data before analyzing and applying the remote sensing images. Its main purpose is to improve the quality, accuracy, and applicability of the images, etc. Perform processing such as geometric registration and radiometric correction on the original remote sensing images in sequence to obtain available two-temporal feature maps. The geometric registration method is implemented by manually selecting control points and using the quadratic polynomial correction method. The geometric registration of the two-phase images is completed by manually selecting control points and using the quadratic polynomial correction method, and the offset is controlled within one pixel. First, solve the polynomial coefficients using the control points between the image to be registered and the reference image; the polynomial coefficient calculation formulas are shown in Equations (1) and 4.2, where (x src , y src ) are the coordinates of the image to be registered, (x ref , y ref ) are the coordinates of the reference image, (A 1 , B 1 , C 1 , D 1 , E 1 , F 1 ), (A 2 , B 2 , C 2 , D 2 , E 2 , F 2 ) are the quadratic polynomial coefficients to be solved.

[0099]

[0100] By selecting homologous control points on the image to be registered and the base image, several groups of (x src , y src ) and (x ref , y ref) Point pairs. Since there are 12 unknowns to be solved, and two sets of equations can be listed for a pair of homologous points, at least 6 pairs of control points are required to solve the polynomial coefficients. Usually, more than 6 pairs of control points are collected to achieve redundant observations and improve the solution accuracy.

[0101] After the polynomial solution is completed, the position of each registered image point on the pre-registration image can be calculated according to its coordinates. Let the registered coordinates be (x new , y new ). Just substitute these coordinates into (x ref , y ref ) in Equations (1) and (2) to obtain the position on the pre-registration image.

[0102] The coordinates calculated in the previous step are often not integers, while the coordinates of the original image are discrete integer values. At this time, interpolation methods need to be used to sample the pixel values at these non-integer coordinates. Bilinear interpolation is an interpolation method widely used in the field of image processing. The specific calculation process is as follows.

[0103] Round the coordinate values x and y up and down to obtain the four corner coordinates:

[0104] (x min , y min ), (x min , y max ), (x max , y min ), (x max , y max )

[0105] For example, the four corner coordinates of (11.3, 12.4) are (11, 12), (11, 13), (12, 12), (12, 13)

[0106] According to the values at the four corner coordinates, linearly interpolate twice in the X-axis direction to obtain the values of the two points (x, y min ), (x, y max )

[0107] In the Y-axis direction, linearly interpolate using the values of the two points (x, y min ), (x, y max ) to obtain the value of the point (x, y).

[0108] The above process involves two linear interpolations, so it is also called bilinear interpolation.

[0109] The calculation formula of linear interpolation is shown in Equation (3). This formula shows the use of (x min , y min ), (x max , y min)Two - point pairs (x, y min )Steps for linear interpolation of points.

[0110]

[0111] Since the four interpolation points in this article are obtained by rounding the x and y coordinates up and down, so there is y max = y min + 1, x max = x min + 1, that is, y max - y min = 1, x max - x min = 1. Simplify Equation (3) according to this relationship, and combine the processes of two - time interpolation in the X - axis direction and one - time interpolation in the Y - axis direction, the calculation formula for interpolating the pixel value at the (x, y) coordinate position using the four - corner coordinates can be obtained, as shown in Equation (4).

[0112] f(x, y)=(1 + y min - y)(1 + x min - x)f(x min y min )+

[0113] (y - y min )(1 + x min - x)f(x min , y max )+

[0114] (1 + y min - y)(x - x min )f(x max , y min )+

[0115] (y - y min )(x - x min )f(x max , y max )(4)

[0116] The radiation correction adopts the histogram - matching algorithm. The basic idea is to use the gray - level distribution of the reference image as the basis, and perform gray - level mapping on each pixel in the image to be matched, so that the gray - level distribution of the new image is basically the same as that of the reference image. The specific calculation steps are as follows:

[0117] Input: Image I to be matched f , reference image I g

[0118] Output: Image I after histogram matching matched

[0119]

[0120]

[0121] For the system according to the second aspect of the present invention, the third processing module 103 is specifically configured such that the two branches of the Siamese change attention residual network based on ResNet and UNet share weights.

[0122] The inputting the two-temporal-phase feature maps into the Siamese change attention residual network based on ResNet and UNet to obtain the change probability map includes:

[0123] Inputting the first map of the two-temporal-phase feature maps into the first branch of the Siamese change attention residual network based on ResNet and UNet to obtain five first fusion feature maps with different scales, specifically including:

[0124] Inputting the first map of the two-temporal-phase feature maps into a network with the first 5 encoding layers of ResNet34 as the network backbone for downsampling to obtain four downsampling intermediate result feature maps and a downsampling result semantic feature map; respectively passing the downsampling result semantic feature map through a convolution once to change the number of channels to 320, with the scale size unchanged, concentrating useful feature information, eliminating redundant feature information, and at the same time reducing model parameters to obtain a convolutional semantic feature map;

[0125] The first layer of ResNet34 is an initial convolutional layer, and the latter 4 layers are residual connection layers. Each residual connection layer is composed of multiple residual units, and each residual unit contains two 3×3 convolutional layers and a shortcut connection. The introduced residual connection shortcut connection solves the problems of gradient disappearance and network degradation in deep networks, enabling the network to be effectively trained to deeper levels.

[0126] Inputting the convolutional semantic feature map into a network of the UNet high-low layer fusion structure based on 5-layer bilinear interpolation upsampling + convolution method for upsampling, gradually changing the number of channels, from the initial 320 to four upsampling intermediate result feature maps and a segmentation probability map of 160, 96, 64, 48, 1 in sequence, that is, the extracted building probability map;

[0127] Fusing the first map and the four downsampling intermediate result feature maps with the four upsampling intermediate result feature maps and the segmentation probability map for high-low layer feature map fusion to obtain five first fusion feature maps with scales of 1 / 2, 1 / 4, 1 / 8, 1 / 16, 1 / 32;

[0128] Inputting the second map of the two-temporal-phase feature maps into the first branch of the Siamese change attention residual network based on ResNet and UNet to obtain five second fusion feature maps with different scales;

[0129] Input the first fusion feature maps and the second fusion feature maps of the five different scales into the change attention residual structure to obtain the third fusion feature maps and the fourth fusion feature maps of the five different scales;

[0130] Add the third fusion feature maps of the five different scales to the fourth fusion feature maps of the five different scales to obtain the fifth fusion feature maps of the five different scales of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32;

[0131] Stitch the fifth fusion feature maps of the five different scales to obtain the change probability map.

[0132] The input of the first map of the two-temporal-phase feature maps into the network with the first 5 encoding layers based on ResNet34 as the network backbone for downsampling to obtain four downsampling intermediate result feature maps and the downsampling result semantic feature map includes:

[0133] Perform convolution on the first map with a kernel size of 7×7 and a stride of 2 to reduce the size of the original image by half, and increase the number of channels from 3 to 64; then pass through two residual connection layers, each time reducing the size of the image by half; finally, pass through two more residual connection layers, each time reducing the size of the image by half and doubling the number of channels. Eventually, obtain the downsampling result semantic feature map with a size of 1 / 32 and 512 channels for the first map, and at the same time, there are four downsampling intermediate result feature maps with different scales and numbers of channels. After the feature map size is halved, it can make the feature map channels deeper while occupying the same video memory, and the deepening of the number of channels is beneficial for the network to learn more rich semantic information features. This set of semantic feature maps contains rich building feature information of different scales.

[0134] The network of the bilinear interpolation upsampling + convolution-based UNet high-low layer fusion structure includes:

[0135] Apply bilinear interpolation upsampling + convolution to replace the transposed convolution in the original UNet high-low layer fusion structure network. Because it is found in actual use that the upsampling implemented by the transposed convolution often causes the final segmentation result to show a significant patch effect. The method of first using bilinear interpolation to expand the feature map size by 2 times and then using conventional convolution to convolve the expanded feature map can well avoid the patch effect. Therefore, in this embodiment, this method is used to replace the transposed convolution scheme of the original UNet high-low layer fusion structure in the network upsampling stage.

[0136] The input of the first fusion feature maps and the second fusion feature maps of the five different scales into the change attention residual structure to obtain the third fusion feature maps and the fourth fusion feature maps of the five different scales includes:

[0137] The variable attention residual structure uses the difference between the first fusion feature maps and the second fusion feature maps at five different scales as the variable residual path and the fusion of the first fusion feature maps and the second fusion feature maps at five different scales as the variable feature path to obtain the third fusion feature maps and the fourth fusion feature maps at five different scales.

[0138] The variable feature path:

[0139] First, the first fusion feature maps and the second fusion feature maps at five different scales are respectively fused by channel concatenation, and then the concatenated feature maps are subjected to 3x3 convolution to obtain the fourth fusion feature maps at five different scales with the same number of channels as before concatenation.

[0140] Compared with the variable residual structure, the variable attention residual structure of this embodiment has the following improvements: (1) Using the method of convolution after channel concatenation of feature maps to realize the fusion of two-phase feature maps. Compared with the direct addition method of the original variable residual structure, this method allows the network to learn richer feature map fusion modes; (2) Using the channel attention mechanism to weight the fused feature maps, enabling the network to give higher weights to important features.

[0141] The five different-scale fifth fusion feature maps are concatenated to obtain a change probability map, including:

[0142] The five different-scale fifth fusion feature maps are upsampled from small to large and then concatenated by channel, and finally a single-channel feature map with the same resolution as the original image is obtained. This feature map passes through the Sigmoid activation function to obtain a change probability map with a value range of (0,1). The original UNet network backbone part can continue to output the semantic segmentation map of the building, enabling the network structure in this paper to support the pre-training of building semantic segmentation samples and multi-task learning.

[0143] The third aspect of the present invention discloses an electronic device. The electronic device includes a memory and a processor. When the processor executes the computer program stored in the memory, it implements the steps in a building change detection method based on a twin variable attention residual network according to any one of the first aspects disclosed in the present invention.

[0144] Figure 3 It is a structural diagram of an electronic device according to an embodiment of the present invention, as Figure 3As shown, the electronic device includes a processor, a memory, a communication interface, a display screen, and an input device connected via a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the electronic device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be achieved through WIFI, a carrier network, near field communication (NFC), or other technologies. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the electronic device, or an external keyboard, touchpad, or mouse, etc.

[0145] Those skilled in the art can understand that Figure 3 the structure shown in [the figure] is only the structural diagram of the part related to the technical solution of the present disclosure, and does not constitute a limitation on the electronic device to which the solution of this application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0146] The fourth aspect of the present invention discloses a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps in a building change detection method based on a twin change attention residual network according to any one of the first aspects disclosed in the present invention are implemented.

[0147] Please note that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combinations of these technical features do not conflict, they should be considered to be within the scope described in this specification. The above embodiments only represent several implementation manners of this application. Their descriptions are relatively specific and detailed, but they should not be construed as a limitation on the scope of the invention patent. It should be pointed out that for those of ordinary skill in the art, without departing from the concept of this application, several deformations and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of the patent of this application should be subject to the appended claims.

Claims

1. A building change detection method based on twin change attention residual network, characterized in that: The method comprises: Step S1, obtaining original remote sensing images of the two phases of study areas; Step S2, performing geometric registration and radiation correction on the original remote sensing image in sequence to obtain a two-phase characteristic map; Step S3: input the two-phase feature graph into the twin change attention residual network based on ResNet and UNet to obtain a change probability graph.

2. According to claim 1, a building change detection method based on twin change attention residual network is characterized in that: In step S3, the two branches of the twin variation attention residual network based on ResNet and UNet share weights.

3. According to claim 2, a building change detection method based on twin change attention residual network is characterized in that: In step S3, the two-phase feature map is input into a twin change attention residual network based on ResNet and UNet to obtain a change probability map, including: The first image of the two-phase feature map is input into the first branch of the twin change attention residual network based on ResNet and UNet to obtain the first fusion feature maps of five different scales, including: The first image of the two-phase feature map is input into a network based on the first 5 encoding layers of ResNet34 as the network skeleton for downsampling, and four downsampled intermediate result feature maps and downsampled result semantic feature maps are obtained; the downsampled result semantic feature maps are respectively convolved once to obtain convolution semantic feature maps; The convolution semantic feature maps are respectively input into the UNet high-low layer fusion structure network based on 5-layer bilinear interpolation upsampling + convolution for upsampling, and four upsampling intermediate result feature maps and segmentation probability maps are obtained, namely, the extracted building probability maps; The first image and four down-sampling intermediate result feature maps are fused with the four up-sampling intermediate result feature maps and the segmentation probability map to obtain five first fused feature maps of different scales; The second image of the two-phase feature map is input into the first branch of the twin variation attention residual network based on ResNet and UNet to obtain the second fused feature maps of five different scales; Inputting the first fused feature maps and the second fused feature maps of the five different scales into the variation attention residual structure to obtain the third fused feature maps and the fourth fused feature maps of the five different scales; Adding the third fused feature maps of the five different scales to the fourth fused feature maps of the five different scales to obtain fifth fused feature maps of the five different scales; The five fifth fusion feature maps of different scales are spliced ​​to obtain a change probability map.

4. According to claim 3, a building change detection method based on twin change attention residual network is characterized in that: In step S3, the first image of the two-phase feature map is input into a network based on the first 5 coding layers of ResNet34 as the network skeleton for downsampling, and four downsampled intermediate result feature maps and downsampled result semantic feature maps are obtained, including: The first image is convolved with a 7×7 convolution and a stride of 2 to halve the size of the original image and increase the number of channels from 3 to 64; then it passes through two residual connection layers to halve the size of the image each time; finally, it passes through two residual connection layers to halve the size of the image each time and double the number of channels, and finally obtains a downsampled semantic feature map with a size of 1 / 32 and a channel number of 512 of the first image, as well as four downsampled intermediate feature maps with different scales and channel numbers.

5. According to claim 3, a building change detection method based on twin change attention residual network is characterized in that: In the step S3, the network of the UNet high-low layer fusion structure of the bilinear interpolation upsampling + convolution method includes: Apply bilinear interpolation upsampling + convolution to replace the transposed convolution in the original UNet high-low layer fusion structure network.

6. According to claim 3, a building change detection method based on twin change attention residual network is characterized in that: In step S3, the first fused feature maps and the second fused feature maps of the five different scales are input into the variable attention residual structure to obtain the third fused feature maps and the fourth fused feature maps of the five different scales, including: The changing attention residual structure uses the difference of the first fused feature map and the second fused feature map of five different scales as the changing residual path and the fusion of the first fused feature map and the second fused feature map of five different scales as the changing feature path to obtain the third fused feature map and the fourth fused feature map of five different scales.

7. A building change detection method based on twin change attention residual network according to claim 6, characterized in that: In step S3, the change characteristic path: First, the first fused feature maps and the second fused feature maps of five different scales are spliced ​​and fused by channel respectively, and then a 3x3 convolution is performed on the spliced ​​feature maps to obtain the fourth fused feature maps of five different scales with the same number of channels as before splicing.

8. A building change detection system based on twin change attention residual network, characterized in that: The system comprises: The first processing module is configured to obtain original remote sensing images of the two-phase study area; The second processing module is configured to sequentially perform geometric registration and radiation correction on the original remote sensing image to obtain a two-phase characteristic map; The third processing module is configured to input the two-phase feature maps into a twin change attention residual network based on ResNet and UNet to obtain a change probability map.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps in a building change detection method based on a twin change attention residual network according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the steps in the building change detection method based on the twin change attention residual network described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Remote sensing image ground object change detection method and device

    CN116385881A

  • Building change detection method and system based on twinborn Unet model

    CN117036941A

  • Remote sensing image change detection method fusing twinborn coding and decoding and attention mechanism

    CN117953369A

Cited By

  • High-resolution optical remote sensing image building change detection method, system and equipment based on texture frequency domain perception and medium

    CN121095770A