Unsupervised image change detection method and device, medium and program product

Through the unsupervised image change detection network, the problem of high error detection rate of remote sensing images under view angle deviation is solved, and high-precision change detection of alignment and detection of heterogeneous images is realized, which improves detection performance.

CN120279407APending Publication Date: 2025-07-08NANJING VOCATIONAL UNIV OF IND TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510305064.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the case of viewing angle deviation, the error detection rate of remote sensing image change detection is high, and there is a lack of explicit constraints on unchanged areas, which makes it difficult to filter out false differences caused by environmental interference and is not robust enough.

Method used

Unsupervised image change detection network is adopted, including a bilateral spatial transformation difference network and a change map iterative correction network. The spatial offset parameters between images are learned through the bidirectional spatial transformation network, combined with the bilateral feature encoding-decoder and bilinear sampler to align images and features, and the change map is used to iteratively correct the network to optimize the binary change map to achieve effective alignment and change detection of images.

Benefits of technology

Effectively aligned multi-source heterogeneous images, suppress structural distortion and non-essential changes, greatly improve the change detection performance of unregistered images, and improve the accuracy and robustness of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279407A_ABST
    Figure CN120279407A_ABST
Patent Text Reader

Abstract

The invention provides an unsupervised image change detection method, an unsupervised image change detection device, a medium and a program product, and designs an image change detection network which is an unsupervised learning network and comprises a bilateral spatial transformation difference network and a change graph iteration correction network. The method comprises the following steps: learning a spatial offset parameter between double-time-phase images by using a bidirectional spatial transformation network, extracting depth features of the images and reconstructing the images by using a bilateral feature coder-decoder, and further generating an initial binary change graph; and completing iterative optimization of the change graph in the change graph iterative correction network. According to the method, a'registration-detection 'combined learning process is constructed, the difference of multi-source heterogeneous images in a space-feature domain is continuously reduced, effective alignment of the heterogeneous images is realized, meanwhile, structural distortion and non-essential change are inhibited, the change detection performance of unregistered images is greatly improved, and the detection precision of the heterogeneous images is improved. The problem of high-precision detection of image change under the condition of visual angle deviation is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image change detection, and specifically relates to an unsupervised image change detection method, device, medium, and program product. Background Art

[0002] With the rapid development of hardware technology and imaging devices, a variety of remote sensing imaging devices have emerged, and various types of remote sensing image data have been generated. These image data have huge differences in spatial resolution, spectral characteristics, imaging conditions, and imaging mechanisms. As an important remote sensing application, remote sensing image change detection refers to an image analysis technology that uses multi-temporal remote sensing images covering the same surface area and other auxiliary data to determine and analyze changes in the surface and objects of interest. By using the information on the changes in the state of ground objects obtained from multi-temporal images, resources can be better managed and utilized.

[0003] Currently, the mainstream change detection methods all require that the two multi-temporal remote sensing images to be compared have been accurately registered. Such as the graph-based multivariate regression method (Graph-based Image Regression and Markov Random Field, GIRMRF) [Sun Y, Lei L, Tan X, et al. Structured Graph Based Image Regression for Unsupervised Multimodal Change Detection [J]. ISPRS Journal of Photogrammetry and Remote Sensing, 2022, 185: 16–31.] and the change detection method based on sparse-constrained adaptive structure consistency [Sun Y, Lei L, Guan D, et al. Sparse-Constrained Adaptive Structure Consistency Based Unsupervised Image Regression for Heterogeneous Remote-Sensing Change Detection [J]. IEEE Transactions on Geoscience and Remote Sensing, 2022, 60: 1–14.] (Sparse-Constrained Adaptive Structure Consistency, SCASC).

[0004] However, since the acquisition of remote sensing images usually involves multiple moments, multiple angles, and differences in imaging devices, etc., even for images of the same area acquired at different time points, there are offsets in the positions of ground objects. This indicates that they may not be spatially aligned, resulting in inconsistent pixel positions of the same geographical location in the two images. Such inconsistencies will cause a large number of non-essential changes, thus directly affecting the accuracy of change detection. On the other hand, remote sensing images with perspective deviations are affected by weather, lighting, seasonal changes, and noise pollution during acquisition, which also makes the change detection problem more complex. To address this problem, the Bilateral Difference Network (BDNN) defines a change mask covering the changed area, designs a global difference convolutional network to compare the global features of multi-temporal images, and omits the registration process. However, the BDNN method lacks explicit constraints on unchanged areas (such as the cross-reconstruction constraint in the present invention), resulting in difficulty in filtering out false differences caused by environmental interference. At the same time, since this method directly follows the paradigm of global feature difference comparison and does not explicitly model the geometric transformation relationship between images, it has insufficient robustness to perspective offsets and non-rigid deformations.

[0005] In addition, with the progress of remote sensing technology, the resolution of corresponding remote sensing images has been improved; compared with past low-resolution images, there will be more situations where tall buildings are difficult to align in multiple perspectives, further exacerbating the misdetection problem of the above-mentioned change detection algorithms. Summary of the Invention

[0006] Aiming at the deficiencies in the prior art, the present invention provides an unsupervised image change detection method, device, medium, and program product to solve the problem of high misdetection rate in image change detection in the presence of perspective deviations.

[0007] The present invention achieves the above technical objectives through the following technical means.

[0008] An unsupervised image change detection method uses the following image change detection network, including:

[0009] A bilateral spatial transformation difference network, which is internally provided with a bidirectional spatial transformation network, a bilateral feature encoding-decoder, and a bilinear sampler, where:

[0010] The bidirectional spatial transformation network is used to learn the spatial offset parameters between the two-temporal images X and Y;

[0011] The bilateral feature encoding-decoder is used to extract the respective deep features of the two-temporal images X and Y, and generate corresponding reconstructed images based on their respective deep features and

[0012] The bilinear sampler aligns the image and features in the spatial domain based on the spatial offset parameters;

[0013] Based on the aligned depth features, an initial feature difference map D is produced, and an initial binary change map P is generated based on the initial feature difference map D 0 ;

[0014] The change map iterative correction network iteratively optimizes the binary change map based on the aligned image.

[0015] Furthermore, in the bidirectional spatial transformation network, the spatial coordinate offset B between the two-phase images is obtained through the localization network ctr :

[0016] B ctr = Lnet([X; Y], S ctr )

[0017] where [X; Y] represents the channel-wise concatenation of images X and Y, Lnet(·) represents the localization network, and S ctr is the control point;

[0018] The localization network has two layers, where:

[0019] B1 = Tanh(MaxPool2d 5×5 (Conv 25×25 ([X; Y])))

[0020] B2 = Tanh(MaxPool2d 5×5 (Conv 5×5 (B1)))

[0021] B ctr = Lnet([X; Y], S ctr ) = Linear(AdaptiveAvgPool 24×24 (B2))

[0022] where B1 represents the output of the first layer, B2 represents the output of the second layer, Conv 25×25 (·) represents a 25×25 convolution, MaxPool2d 5×5 (·) represents a 5×5 max pooling, Tanh(·) represents the Tanh activation function, Conv 5×5 (·) represents a 5×5 convolution, AdaptiveAvgPool 24×24 (·) represents a 24×24 adaptive average pooling, and Linear(·) represents a linear fully connected layer;

[0023] Based on the control point S ctr and the spatial coordinate offset B ctr, learn the spatial offset parameters between the dual-phase images X and Y:

[0024]

[0025]

[0026] In the formula, represents the control points for the alignment transformation from image X to image Y, represents the control points for the alignment transformation from image Y to image X, represents the spatial transformation function for the alignment of image X to image Y, and the subscript represents the learnable parameters in the positioning network, is the grid sampling point after the alignment transformation on image X, and Grid(·) represents the grid generator, where the sampling point coordinates are generated by the TPS interpolation method, is the grid sampling point after the alignment transformation on image Y, represents the spatial transformation function for the alignment of image Y to image X.

[0027] Furthermore, the bilateral feature encoding-decoder includes a bilateral feature encoder and a bilateral feature decoder, where:

[0028] The bilateral feature encoder is composed of two encoders E1(·) and E2(·) with the same structure. Each encoder has L cascaded convolutional layers, which extract the depth features F X and F Y from images X and Y respectively:

[0029]

[0030] In the formula, θ1 and θ2 are the learnable parameters in the encoders E1(·) and E2(·) respectively, are the convolutional kernel weights and biases of the l-th layer in the encoder E1(·) respectively, are the convolutional kernel weights and biases of the l-th layer in the encoder E2(·) respectively, σ(·) is the activation function, and "*" represents the convolution operation, represents the cascading of 1 to L convolutional layers;

[0031] The bilateral feature decoder is composed of two decoders D1(·) and D2(·) with the same structure. Each decoder is composed of a cascade of multiple convolutional layers, and respectively decodes through the depth features F X and F Y to obtain the reconstructed images and

[0032]

[0033] Where, ζ1 and ζ2 are the learnable parameters in the decoders D1(·) and D2(·) respectively;

[0034] The bilinear sampler aligns the and aligned images and features:

[0035]

[0036] Where, S(·) represents the differentiable bilinear sampler, X ′ is the image aligned from image X to image Y, Y ′ is the image aligned from image Y to image X, F X ′ is the feature aligned from feature F X to feature F Y aligned, F Y ′ is the feature aligned from feature F Y to feature F X aligned.

[0037] Furthermore, in the encoder, there are sequentially 5×5 convolution, 1×1 convolution, 5×5 convolution, 1×1 convolution. After each convolution, it is activated by the Tanh activation function, and finally the depth feature is output after layer normalization;

[0038] In the decoder, there are 4 consecutive 1×1 convolutions, and after each convolution, it is activated by the Tanh activation function.

[0039] Furthermore, in the bilateral spatial transformation difference network, there is the following objective function:

[0040]

[0041] Where (i,j) represents the spatial coordinates of the pixel point, F Y (i,j), F X (i,j), X(i,j), M2(i,j), Y(i,j), M1(i,j) respectively represent F Y , F X , X, M2, Y, M1 at the coordinate point (i,j), ‖·‖2 represents the 2-norm, and λ is the weight coefficient.

[0042] Furthermore, the initial feature difference map D is obtained by the following formula:

[0043]

[0044] Where, D(i,j), F X ′(i, j) respectively represent D and F X ′ The value at the coordinate point (i, j);

[0045] The initial binary change map is obtained by performing pixel segmentation on the initial feature difference map D using an automatic multi-threshold segmentation algorithm:

[0046] P 0 = AMT(D)

[0047] In the formula, AMT(·) represents the automatic multi-threshold segmentation algorithm;

[0048] The change iteration correction network is composed of multiple K correction modules. Each correction module includes a feature extraction network and a discriminant network. In the k-th correction module:

[0049] The feature extraction network has a symmetric bilateral network structure. One side of the network structure is as follows: First, in the first layer, two parallel branches are used to extract image features using 5×5 convolution and 3×3 convolution respectively. After convolution of the two branches, they are activated by the Tanh activation function. Then, the output results of the two branches are concatenated according to the feature channel dimension. The second to the third layers are two consecutive 1×1 convolutions, and each convolution is activated by the Tanh activation function. The fourth layer first undergoes 1×1 convolution and then is activated by the Sigmoid activation function. Finally, it passes through the convolutional block attention module and L2 normalization in sequence;

[0050] The two symmetric sides of the feature extraction network respectively input the image X ′ and the image Y, and the Euclidean distance is calculated between the obtained outputs to obtain the feature difference map

[0051] In the discriminant network, according to the feature difference map Output a binary classification result:

[0052]

[0053] In the formula, is the probability change map output by the k-th correction module, represents the discriminant network, φ is the learnable parameter in the discriminant network, Tanh(·) represents the Tanh activation function, Conv 1×1 (·) represents 1×1 convolution, Conv 3×3 (·) represents 3×3 convolution;

[0054] For Perform binarization processing using the following dynamic selection threshold strategy to obtain the binary change map P output by the k-th correction module k :

[0055]

[0056] The dynamic selection threshold strategy is as follows:

[0057] Set a group of thresholds, and perform a round of binarization processing on the change map in turn based on each threshold ;

[0058] In each round of binarization, calculate the Euclidean distance between the binarized change map obtained in the current round and the output result of the previous round. If:

[0059] ① The Euclidean distance ≥ the distance judgment parameter Fbefore, then use the output result of the previous round as the output of this round;

[0060] ② The Euclidean distance < the distance judgment parameter Fbefore, then use the binarized change map obtained in this round as the output of this round, and update the value of the distance judgment parameter Fbefore to the Euclidean distance;

[0061] Where P k-1 is used as the initial binarized change map for comparison with the first round, and the initial value of Fbefore is given manually;

[0062] Until all rounds of iteration are completed, output the binarized change map.

[0063] Furthermore, the objective function of the k-th correction module is:

[0064]

[0065] In the formula, are the learnable parameters in the feature extraction network, α is the parameter for controlling the boundary of the change map, ω k-1 is the dynamic weight, P k-1 (i, j), ω k-1 (i, j) respectively represent the values of P k-1 , ω k-1 at the coordinate point (i, j); the expression of the dynamic weight ω k-1 is:

[0066]

[0067] In the formula, n and γ are weight parameters;

[0068] The above objective function is optimized by the following alternating iteration method:

[0069]

[0070] Among them, during the optimization When dealing with sub - problems, the back - propagation algorithm is used to optimize the learnable parameters in the feature extraction network Obtain the required feature difference map after optimization

[0071] During the optimization When dealing with sub - problems, fix the optimized Iteratively update the learnable parameters φ of the discriminant network to obtain the optimal solution

[0072] A computer device includes a memory and a processor;

[0073] The memory is used to store computer programs;

[0074] The processor is used to execute the computer program and implement the above - mentioned unsupervised image change detection method when executing the computer program.

[0075] A computer - readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to execute the above - mentioned unsupervised image change detection method.

[0076] A computer program product includes a computer program, and when the computer program is executed by a processor, the above - mentioned unsupervised image change detection method is implemented.

[0077] The beneficial effects of the present invention are as follows:

[0078] (1) The present invention provides an unsupervised image change detection method, device, medium, and program product. Among them, an image change detection network is designed, which is an unsupervised learning network, including a bilateral spatial transformation difference network and a change map iterative correction network. The two sub - networks construct a "registration - detection" joint learning process, continuously narrow the differences between multi - source heterogeneous images in the spatial - feature domain, achieve effective alignment of heterogeneous images, and at the same time suppress structural distortion and non - essential changes, greatly improving the change detection performance of unregistered images and solving the problem of high - precision image change detection in the case of perspective deviation.

[0079] (2) The two - way spatial transformation network designed by the present invention can adaptively learn the spatial coordinate transformation relationship between unregistered images by establishing a two - way transformation learning process of the images before and after the event.

[0080] (3) The present invention constructs an objective function composed of spatial transformation parameters, feature difference minimization, and cross - reconstruction constraints, aiming to learn the spatial offset or geometric transformation parameters of the unchanged area from multi - source heterogeneous remote sensing images in an unsupervised manner and obtain double - temporal image features with consistent distribution and spatial alignment, and then obtain the surface changes in the double - temporal remote sensing images by comparing the features.

[0081] (4) The designed bilateral spatial transformation difference network of the present invention realizes the consistency of pixels in the spatial domain and the feature domain by using cross-reconstruction and feature difference minimization, which is convenient for the calculation of the feature difference map and the initial binary change map. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] Figure 1 It is the structural diagram of the image change detection network of the present invention;

[0083] Figure 2 It is the structural diagram of the bidirectional spatial transformation network in the present invention;

[0084] Figure 3 It is the structural diagram of the bilateral feature encoder-decoder in the present invention;

[0085] Figure 4 It is the structural diagram of the correction module in the present invention;

[0086] Figure 5 They are the dual-temporal images of the Beijing dataset and the reference map of the change difference between the two;

[0087] Figure 6 For Figure 5 the dataset shown, it is the verification result of the change maps generated by each detection method;

[0088] Figure 7 They are the dual-temporal images of the UVA dataset and the reference map of the change difference between the two;

[0089] Figure 8 For Figure 7 the dataset shown, it is the verification result of the change maps generated by each detection method;

[0090] Figure 9 They are the dual-temporal images of the G1 dataset and the reference map of the change difference between the two;

[0091] Figure 10 For Figure 9 the dataset shown, it is the verification result of the change maps generated by each detection method;

[0092] Figure 11 They are the dual-temporal images of the G2 dataset and the reference map of the change difference between the two;

[0093] Figure 12 For Figure 11 the dataset shown, it is the verification result of the change maps generated by each detection method;

[0094] Figure 13 They are the dual-temporal images of the Gloucester dataset;

[0095] Figure 14 ForFigure 13 The dataset shown, the verification results of the change maps generated by each detection method, and the reference maps;

[0096] Figure 15 Is the double - temporal image of the Tianhe dataset;

[0097] Figure 16 For Figure 15 The dataset shown, the verification results of the change maps generated by each detection method, and the reference maps;

[0098] Figure 17 Is the double - temporal image of the Shuguang dataset;

[0099] Figure 18 For Figure 17 The dataset shown, the verification results of the change maps generated by each detection method, and the reference maps;

[0100] Figure 19 Is the double - temporal image of the Sardinia dataset;

[0101] Figure 20 For Figure 19 The dataset shown, the verification results of the change maps generated by each detection method, and the reference maps. Detailed implementation manners

[0102] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.

[0103] I. Technical solution

[0104] For the double - temporal remote - sensing images (input images) that need to perform image change detection, denote the pre - event image as The post - event image as In the formula, H, W, C X (or C Y ) are respectively the three dimensions of the input image, that is, height, width, and number of channels.

[0105] As Figure 1 Shown is the image change detection network designed by the present invention, which is used to perform corresponding image change detection on the above - mentioned double - temporal remote - sensing images. The detection network mainly includes ① bilateral spatial transformation difference network and ② change map iterative correction network.

[0106] 1. Bilateral spatial transformation difference network

[0107] As Figure 1As shown in the figure, the bilateral spatial transformation difference network specifically includes ① the bidirectional spatial transformation network, ② the bilateral feature encoding-decoder, and ③ the bilinear sampler.

[0108] 1.1, Bidirectional Spatial Transformation Network

[0109] As Figure 2 shown in the figure is the structural diagram of the bidirectional spatial transformation network, where:

[0110] (1) Given a set of initial control point sets The initial control points are evenly distributed in the image space; for example, when d = 2, there are four control points {(x i , y j ) | i, j ∈ {1, 2}} and are distributed at the four corners of the image. The input images X and Y pass through the localization network to predict the spatial coordinate offsets

[0111] B ctr = Lnet([X; Y], S ctr )

[0112] In the formula, [X; Y] represents concatenating images X and Y along the channel dimension, and Lnet(·) represents the localization network, through which the set B of spatial coordinate offsets of each control point is obtained ctr .

[0113] There are two layers in the localization network, where:

[0114] 1) For the input image in the first layer, perform 25×25 convolution (stride = 11, dilation = 1) and 5×5 max pooling (MaxPool2d, stride = 2) in sequence, and then activate with the Tanh activation function. The obtained result is input to the second layer;

[0115] 2) For the output of the first layer in the second layer, perform 5×5 convolution (stride = 1) and 5×5 max pooling (MaxPool2d, stride = 2) in sequence, and then activate with the Tanh activation function;

[0116] 3) Finally, perform 24×24 adaptive average pooling (AdaptiveAvgPool) on the output of the second layer and output through a linear fully connected layer (Linear, number of control points C = 4) to obtain the spatial coordinate offset B ctr .

[0117] The above process is expressed in the form of an expression as:

[0118] B1 = Tanh(MaxPool2d 5×5 (Conv25×25 ([X; Y])))

[0119] B2 = Tanh(MaxPool2d 5×5 (Conv 5×5 (B1)))

[0120] B ctr = Lnet([X; Y], S ctr ) = Linear(AdaptiveAvgPool 24×24 (B2))

[0121] In the formula, B1 represents the output of the first layer, B2 represents the output of the second layer, Conv 25×25 (·) represents a 25×25 convolution, MaxPool2d 5×5 (·) represents a 5×5 max pooling, Tanh(·) represents the Tanh activation function, Conv 5×5 (·) represents a 5×5 convolution, AdaptiveAvgPool 24×24 (·) represents a 24×24 adaptive average pooling, Linear(·) represents a linear fully connected layer.

[0122] (2) Based on the initial control point S ctr and the spatial coordinate offset B ctr , respectively obtain the corrected coordinates of each control point on images X and Y:

[0123]

[0124] In the formula, represents the alignment transformation from image X to image Y, and the set of corrected control point coordinates after transformation; represents the alignment transformation from image Y to image X, and the set of control point coordinates after transformation.

[0125] (3) Obtain the grid sampling points after the alignment transformation on images X and Y. Taking the alignment transformation from image X to image Y as an example:

[0126] Set a number of sampling points on the original image X. The sampling points are evenly distributed on the surface of image X in a grid pattern, and the set formed by the sampling points is denoted as G;

[0127] After the alignment transformation of image X to image Y, based on the control points adopt the Thin Plate Spines (TPS) interpolation method to obtain the coordinates of other pixel points (except control points) after the alignment transformation of image X, and correspondingly obtain the coordinates of each grid sampling point after the alignment transformation:

[0128]

[0129] In the formula, represents the spatial transformation function for aligning image X to image Y, and the subscript represents the learnable parameters in the positioning network. is the set of grid sampling points after the alignment transformation on image X. Grid(·) represents the grid generator, that is, the sampling point coordinates are generated using the TPS interpolation method.

[0130] Similarly, for the alignment transformation of image Y to image X, the corresponding grid sampling points are:

[0131]

[0132] In the formula, is the set of grid sampling points after the alignment transformation on image Y. represents the spatial transformation function for aligning image Y to image X.

[0133] 1.2, Bilateral Feature Encoding - Decoder

[0134] As Figure 3 shown is the structural diagram of the bilateral feature encoding - decoder, including a bilateral feature encoder and a bilateral feature decoder, where:

[0135] (1) Bilateral Feature Encoder

[0136] The bilateral feature encoder consists of two encoders E1(·) and E2(·) with the same structure. Each encoder has L cascaded convolutional layers. E1(·) and E2(·) input different images respectively, thus obtaining their respective corresponding outputs. The depth features of images X and Y are extracted using the above - mentioned bilateral feature encoder:

[0137]

[0138] In the formula, and are the depth features extracted from images X and Y respectively. E1(·) and E2(·) are the two encoders respectively, and θ1 and θ2 are the learnable parameters in encoders E1(·) and E2(·) respectively. are the convolutional kernel weights and biases of the l - th layer in encoder E1(·) respectively. are the convolutional kernel weights and biases of the l - th layer in encoder E2(·) respectively. σ(·) is the activation function, and "*" represents the convolution operation. represents cascading 1 to L convolutional layers.

[0139] As Figure 3As shown, for the encoder E1(·) or E2(·), in this embodiment, L = 4, and the activation function σ(·) is the Tanh activation function. Specifically:

[0140] 1) The first layer is a 5×5 convolution (padding = 2, number of output features 16), followed by activation by the Tanh activation function;

[0141] 2) The second layer is a 1×1 convolution (number of output features 16), followed by activation by the Tanh activation function;

[0142] 3) The third layer is a 5×5 convolution (padding = 2, number of output features 32), followed by activation by the Tanh activation function;

[0143] 4) The fourth layer is a 1×1 convolution (number of output features 32), followed by activation by the Tanh activation function and layer normalization to obtain the output of the encoder.

[0144] The above process is expressed in the form of an expression as:

[0145] F 11 = Tanh(Conv 5×5 (X or Y))

[0146] F 12 = Tanh(Conv 1×1 (F 11 ))

[0147] F 13 = Tanh(Conv 5×5 (F 12 ))

[0148] F 14 = Tanh(Conv 1×1 (F 13 ))

[0149] E 1或2 (X or Y) = LN(F 14 )

[0150] In the formula, F 11 , F 12 , F 13 , F 14 are the outputs of the above first layer, second layer, third layer, and fourth layer respectively, and LN(·) represents layer normalization.

[0151] (2) Bilateral Feature Decoder

[0152] The bilateral feature decoder consists of two decoders D1(·) and D2(·) with the same structure. Each decoder is composed of multiple cascaded convolutional layers. The above bilateral feature decoder is used to process the depth feature FX and F Y are decoded to obtain the reconstructed images accordingly and

[0153]

[0154] where D1(·) and D2(·) are two decoders respectively, and ζ1 and ζ2 are the learnable parameters in the decoders D1(·) and D2(·).

[0155] As Figure 3 shown, for the encoder D1(·) or D2(·), in this embodiment, it is set as a 4-layer convolutional layer. Specifically:

[0156] 1) The first layer is a 1×1 convolution (output feature number 16), and then it is activated by the Tanh activation function;

[0157] 2) The second layer is a 1×1 convolution (output feature number 8), and then it is activated by the Tanh activation function;

[0158] 3) The third layer is a 1×1 convolution (output feature number 4), and then it is activated by the Tanh activation function;

[0159] 4) The fourth layer is a 1×1 convolution (output feature number 3), and then it is activated by the Tanh activation function to obtain the output of the decoder. The above process is expressed in the form of an expression as:

[0160] F 21 = Tanh(Conv 1×1 (F X or F Y ))

[0161] F 22 = Tanh(Conv 1×1 (F 21 ))

[0162] F 23 = Tanh(Conv 1×1 (F 22 ))

[0163] D 1或2 (F X or F Y ) = Tanh(Conv 1×1 (F 23 ))

[0164] where F 21 、F 22 、F 23 are the outputs of the above first layer, second layer, and third layer respectively.

[0165] 1.3, Bilinear Sampler

[0166] As Figure 1 shown, according to the sampling grid points and use the bilinear sampler to align the images and features:

[0167]

[0168]

[0169] In the formula, S(·) represents the differentiable bilinear sampler, is the image aligned from image X to image Y, is the image aligned from image Y to image X, is the feature aligned from feature F X to feature F Y , is the feature aligned from feature F Y to feature F X .

[0170] 1.4, Objective Function

[0171] The bilateral spatial transformation difference network established above is unsupervised learning. Therefore, the following objective function is designed to learn and train the network to optimize and obtain the final required aligned images X ′ , Y ′ and aligned features F X ′ , F Y ′ .

[0172] (1) Based on the principle that there should be a smaller difference in the unchanged regions between the two-temporal images, the following objective function for minimizing the difference between the unchanged features is set:

[0173]

[0174] In the formula, ‖·‖ F represents the F-norm (Frobenius norm).

[0175] (2) Minimize the background difference between the input and aligned images through cross-reconstruction constraints, and construct the following objective function:

[0176]

[0177] In the formula, represents element-wise multiplication; and To eliminate the impact on training caused by the sampled black edges in the original images X and Y, the specific expressions of M1 and M2 are as follows:

[0178]

[0179] In the formula, is a matrix of all ones,

[0180] (3) Combining the above two objective functions, the total objective function is defined as:

[0181]

[0182] where (i, j) represents the spatial coordinates of the pixel point, F Y F X F F Y F X F

[0183] By solving the above total objective function, X ′ Y ′ F X ′ F Y ′ are optimized.

[0184] 2. Initial feature difference map and change map

[0185] As shown in reference Figure 1 :

[0186] (1) Taking the alignment of feature F X to F Y as an example (similarly for the reverse case), the pixel-level feature difference map D is:

[0187]

[0188] In the formula, D(i, j) and F X ′ F X ′ represent the values of D and F

[0189] (2) According to the feature difference map D, the Automatic Multilevel Thresholding (AMT) algorithm is used for pixel segmentation to obtain the initial binary change map P 0 :

[0190] P 0 = AMT(D)

[0191] In the formula, is the initial binary change map, where the value of P 0 (i, j) is 0 or 1.

[0192] 3. Change map iterative correction network

[0193] As Figure 1 shown, the change map iterative correction network consists of multiple (K) correction modules, which are used to learn the changes in different feature spaces under the guidance of the initial change map for the aligned images, so as to further refine the change detection results of the bilateral spatial transformation difference network, thereby improving the change detection performance.

[0194] As Figure 4 shown, each correction module has a symmetric bilateral network structure, including a feature extraction network and a discriminant network. Taking the alignment transformation of image X to Y as an example, for the k-th correction module:

[0195] 3.1 Feature extraction network

[0196] Feature extraction network

[0197] The feature extraction network outputs the feature difference map according to the aligned image X ′ and the original image Y

[0198] Figure 4 shown, the network structures on the left and right sides of the feature extraction network are symmetric, where:

[0199] (1) The input of the red box part on the left side of the figure is the aligned image X ′

[0200] 1) First, there are two parallel branches in the first layer, which use convolutions of different scales to extract features, namely 5×5 convolution (padding = 2, the number of input image channels is 3, and the number of output features is 20) and 3×3 convolution (padding = 1, the number of input image channels is 3, and the number of output features is 20). After convolution, the output results of the two branches are activated by the Tanh activation function respectively; then, the output results of the two branches are concatenated according to the feature channel dimension and enter the second to third layers;

[0201] 2) The second to third layers are two consecutive 1×1 convolutions (with 40 output feature numbers), and each convolution is activated by the Tanh activation function, and the obtained result enters the fourth layer;

[0202] 3) The fourth layer first undergoes a 1×1 convolution (with 40 output feature numbers), and then is activated by the Sigmoid activation function;

[0203] 4) The output of the fourth layer is successively passed through the Convolutional Block Attention Module (CBAM) and L2 normalization to obtain the feature X ′ 5 of X. ′ . The above process is expressed in the form of an expression as:

[0204] X1 ′ 1 = Tanh(Conv 5×5 (X ′ ))

[0205] X1 ′ 2 = Tanh(Conv 3×3 (X ′ ))

[0206] X3 ′ = Tanh(Conv 1×1 (Tanh(Conv 1×1 ([X1 ′ 1; X1 ′ 2]))))

[0207] X4 ′ = Sig(Conv 1×1 ([X2 ′ ))

[0208] X5 ′ = L2(CBAM([X3 ′ ))

[0209] In the formula, X1 ′ 1, X1 ′ 2 are the features extracted by the two parallel branches of the first layer above, [X1 ′ 1; X1 ′ 2] represents concatenating the two features along the channel dimension, X3 ′ , X4 ′ are the outputs of the third and fourth layers respectively, Sig(·) represents the Sigmoid activation function, and L2(·) represents L2 normalization.

[0210] (2) The input original image Y of the blue frame part on the right side of the figure, and the specific network structure settings are the same as those of the red frame part on the left side, and the corresponding feature Y5 of Y is obtained.

[0211] (3) In the difference perception layer, for the feature X5′ Perform Euclidean distance calculation with Y5 to obtain the feature difference map (of the k-th correction module).

[0212] 3.2, Discriminant network

[0213] (1) The discriminant network outputs a binary classification result based on the feature difference map Output binary classification result:

[0214]

[0215] In the formula, is the probability change map output by the current (k-th) correction module, where represents the discriminant network, and φ is the learnable parameter in the discriminant network.

[0216] Specifically in this embodiment, Figure 4 the shown discriminant network has three convolutional layers, and each convolution is activated by the Tanh activation function. The first layer is a 3×3 convolution (padding = 2, dilation = 2, output feature number 40); the second layer is a 1×1 convolution (output feature number 40); the third layer is a 1×1 convolution (output feature number 3); through the above three convolutional operations, the probability map is obtained

[0217] The above process is expressed in the form of an expression as:

[0218]

[0219] (2) Binary processing

[0220] For the probability change map P k Perform binary processing to obtain the binary change map accordingly

[0221] In this embodiment, the following Dynamic Selection Threshold (DST) strategy is adopted to perform binary processing on the probability change map Perform binary processing as follows:

[0222]

[0223] (1) Set a group of thresholds, and perform a round of binary processing on the change map based on each threshold in turn;

[0224] (2) For the first-round binary processing, calculate the binary change map obtained in the first round and the binary change map P obtained in the previous step k-1The Euclidean distance between them is compared with the given distance judgment parameter Fbefore (in this embodiment, the initial value of Fbefore is set to 1×10 -15 ), and if:

[0225] ① If the Euclidean distance ≥ the distance judgment parameter Fbefore, then the binary change map P of the previous step k-1 is used as the output of the first round;

[0226] ② If the Euclidean distance < the distance judgment parameter Fbefore, then the binary change map obtained in the first round is used as the output, and the value of the distance judgment parameter Fbefore is updated to the Euclidean distance;

[0227] (3) In the binarization of each round from the second round onwards, calculate the Euclidean distance between the binary change map obtained in the current round and the output result of the previous round. If:

[0228] ① If the Euclidean distance ≥ the distance judgment parameter Fbefore, then the output result of the previous round is used as the output of this round;

[0229] ② If the Euclidean distance < the distance judgment parameter Fbefore, then the binary change map obtained in this round is used as the output of this round, and the value of the distance judgment parameter Fbefore is updated to the Euclidean distance.

[0230] Until all the above rounds of iteration end, the binarized change map is output.

[0231] 3.3, objective function

[0232] The above-established correction module is also unsupervised learning. Therefore, the following objective function is designed to learn and train each correction module, so as to optimize and obtain the final required probability change map P k .

[0233] (1) The objective function of each correction module (taking the kth correction module as an example) is expressed as:

[0234]

[0235] In the formula, are the learnable parameters in the feature extraction network, α is the boundary parameter used to control the change map, ω k-1 is the dynamic weight, P k-1 (i, j), ω k-1 (i, j) respectively represent the values of P k-1 , ω k-1 at the coordinate point (i, j).

[0236] The dynamic weight ωk-1 The expression of

[0237]

[0238] is: In the formula, n and γ are weight parameters.

[0239] (2) For the solution problem of the above objective function, this embodiment adopts an alternating iteration method to optimize the learnable parameters and φ in the objective function, that is, when fixing one parameter, optimize the other parameter, and iterate and optimize until convergence. The whole optimization process can be decomposed into the following two minimization problems:

[0240]

[0241] Among them, when optimizing the sub-problem, the backpropagation algorithm is used to optimize the learnable parameters in the feature extraction network so as to obtain the required feature difference map after optimization During the iteration process, when it is within the range of 0 to α, the changing pixels will only contribute to the loss function.

[0242] When optimizing the sub-problem, fix the optimized and iteratively update the learnable parameter φ of the discriminant network to obtain the optimal solution

[0243] In summary, in each correction module (the kth one), first, the alignment image X ′ and the original image Y are processed by the feature extraction network to obtain the feature difference map and then processed by the discriminant network to obtain the probability change map And in the process of generating the above probability change map , based on the binary change map P k-1 obtained in the previous step, the correction module is optimized using the designed objective function to obtain the optimized Finally, the probability change map is subjected to DST binarization processing to obtain the binary change map P k as the final output of this module.

[0244] II. Test and verification

[0245] 1. Test set

[0246] This test uses eight groups of unregistered remote sensing image data sets to verify the effectiveness of the present invention. They are respectively:

[0247] 1) Real datasets, including test sets code-named "Beijing", "UVA", "G1", and "G2".

[0248] 2) Unregistered remotely sensed datasets synthesized by simulation, including test sets code-named "Gloucester", "Tianhe", "Shuguang", and "Sardinia". For the simulated datasets, this test artificially distorted the pre-event images by an angle of 60°.

[0249] Figure 5 Shown is the Beijing dataset. In the figure, (a) and (b) are RGB high-resolution images captured by Google Earth on September 30, 2012, and March 4, 2013, respectively. The size of both images is 500×500×3, and the spatial resolution is 1 meter per pixel. The test objective of this dataset is the change detection ability for building remotely sensed images. In the figure, (c) is the reference image for comparison of the detection results.

[0250] Figure 7 Shown is the UVA dataset. In the figure, (a) and (b) are urban street scenes at the same location captured by an unmanned aerial vehicle (UAV) at different times, belonging to UAV-borne optical remotely sensed images. The size of the images in this dataset is 512×512 pixels. Since the UAV takes pictures during dynamic flight, even when capturing the same scene, the perspective position and angle are different, making it difficult to accurately match many details such as lines, edges, and textures. This poses a huge challenge for change detection of remotely sensed images. In the figure, (c) is the reference image for comparison of the detection results.

[0251] Figure 9 and Figure 11 Shown are the G1 dataset and the G2 dataset, respectively. Both are from the High-Resolution Challenge Dataset (http: / / sw.chreos.org / challenge). In each figure, (a) and (b) are optical images at the same location at different times. The image resolution is 1 meter, and the size is 512×512 pixels. It is obvious that the images at the two moments in both test sets are not aligned, which is used to verify the change detection ability for unaligned images. In each figure, (c) is the reference image for comparison of the detection results.

[0252] Figure 13Shown is the Gloucester dataset, which demonstrates the changes in remote sensing images caused by floods destroying land. In the figure, (a) is the original SPOT image at a certain moment, (b) is the distorted image of (a), and (c) is the ERS image at another moment. The image sizes are all 990×554 pixels. The SPOT image contains three spectral bands, and the spatial resolution is approximately 25 meters. Due to the different acquisition methods of the two remote sensing images, the differences in image properties pose challenges to the change detection task.

[0253] Figure 15 Shown is the Tianhe dataset, which is used to study the changes in the highway-airport area before and after the expansion of Wuhan Tianhe International Airport. In the figure, (a) is the panchromatic image captured by Landsat-7 in July 2002, (b) is the distorted image of (a), and (c) is the optical image captured by Google Earth in June 2013. The sizes are all 666×615 pixels, the coverage area is 7.28×6.73, and the spatial resolution is approximately 11 meters.

[0254] Figure 17 Shown is the Shuguang dataset, which is used to study the situation of the land in the urban and rural areas of Shuguang Village in the suburbs of Dongying City being covered by buildings over time. In the figure, (a) is the SAR image captured by the Radarsat-2 satellite in the C band in June 2008, (b) is the distorted image of (a), and (c) is the optical image captured by Google Earth in September 2012. The sizes are all 921×593 pixels, and the spatial resolution is 8 meters.

[0255] Figure 19 Shown is the Sardinia dataset, which is used to monitor the lake expansion. In the figure, (a) is the near-infrared (NIR) band image captured by the Radarsat-5 satellite in September 1995, (b) is the distorted image of (a), and (c) is the optical image captured by Google Earth in July 1996. The spatial resolution is 30 meters.

[0256] 2. Test environment

[0257] The experimental environment is as follows:

[0258] CPU: Intel i7-9700K, GPU: NVIDIA GeForce RTX 2080Ti, Memory: 32GB, Python 3.6 and PyTorch 1.8.

[0259] In the image change detection network of the present invention: the number of feature channels C1 is set to 20, and the number of feature maps of the bilateral feature extraction network is {C X (or C Y), C1, 2C1, 2C1, 2C1, 2C1}. The discrimination network consists of three convolutional layers. The size of the convolutional kernel of the last layer is 1×1, and the Sigmoid function is used as the activation function. The number of feature maps is {2C1, 2C1, 2C1, C1}. The stride of all layers is 1, and Padding is applied to maintain the spatial dimensions of the feature maps. The network optimizer uses the Adamaxp optimizer. The number of distorted control points is 4, which are distributed at the four corners of the image. The learning rates are set to 0.0001 and 0.001 in the localization network and the bilateral network respectively. The decay coefficient is 0.00001. The number of iterations K of the change map is set to 50, and the tolerable error ∈ is set to 0.0001. λ is set to 0.5, and the weight γ is set to 0.1.

[0260] In the simulation dataset, the parameter α is set to 0.02, 0.5, 0.4, and 0.07 in the Gloucester, Tianhe, Shuguang, and Sardinia datasets respectively. In the real dataset, the parameter α is set to 0.1, 0.3, 0.2, and 0.2 in the Beijing, UVA, G1, and G2 datasets respectively.

[0261] The present invention uses AUC (Area Under the ROC Curve), FP (False Positive), FN (False Negative), OE (Overall Error Number), OA (Total Accuracy), KC (Kappa Coefficient), and F1 (F1 Score) as performance evaluation indicators. The comparison methods include the classical algorithm CVA for high-resolution change detection and the change detection algorithms specifically for multi-source heterogeneous remote sensing images, namely SCASC and GIRMRF.

[0262] 3. Test Results

[0263] The change detection accuracies obtained by the method of the present invention and each existing method for the real dataset and the simulation dataset through the comparative tests are shown in Table 1 and Table 2 below:

[0264] Table 1: Comparative Test of Real Dataset

[0265]

[0266] Table 2: Comparative Test of Simulation Dataset

[0267]

[0268] Figure 6 , Figure 8 , Figure 10 , Figure 12 , Figure 14 , Figure 16 , Figure 18 , Figure 20Image change detection results (change maps) produced by the method of the present invention and various existing methods for real datasets and simulation datasets. Among them Figure 6 is for Figure 5 the Beijing dataset shown, Figure 8 is for Figure 7 the UVA dataset shown, Figure 10 is for Figure 9 the G1 dataset shown, Figure 12 is for Figure 11 the G2 dataset shown, Figure 14 is for Figure 13 the Gloucester dataset shown, Figure 16 is for Figure 15 the Tianhe dataset shown, Figure 18 is for Figure 17 the Shuguang dataset shown, Figure 20 is for Figure 19 the Sardinia dataset shown.

[0269] Combined with the intuitive comparison shown in Table 1 and Table 2 above and the change maps of each group:

[0270] 1) Since Figures 5 to 12 there are unregistered pixels in these real datasets, the CVA method will produce a large number of errors during comparison, making objects that have not changed (such as Figure 7 the buildings between the two images (a) and (b) in the UVA dataset shown) may be mislabeled as change regions. See Figure 6 , 8 , 10, 12, the green pixels in each (a) image.

[0271] 2) Due to the lack of spatial coordinates between the aligned bi-temporal images, the SCASC method results in a large number of false detections. See Figure 6 , 8 , 10, 12, the green pixels in each (b) image; similarly, GIRMRF also has a large number of missed detections. See Figure 6 , 8 , 10, 12, the red pixels in each (c) image.

[0272] 3) The present invention uses a bidirectional spatial transformation network to cancel geometric distortions in remote sensing image change detection and maps the input images to a shared feature transformation domain, enabling the network to effectively distinguish between changed and unchanged objects. At the same time, the present invention effectively suppresses unchanged regions in the difference map, improves the AUC value, and exhibits better performance. Compared with other existing methods, the present invention performs more excellently in reducing FP and FN, and thus obtains a higher F1 score. It shows that the detection network of the present invention achieves a good balance between precision and recall.

[0273] The advantage of the present invention in obtaining the above test results is that it adopts a two-way spatial transformation difference network to adaptively learn the spatial-feature change relationship between dual-temporal images. In this way, not only can image registration be automatically performed, but also shared features can be learned.

[0274] 4) As Figure 14 、 16 As shown in FIGS. 18 and 20, the image change detection for the simulation dataset is presented. Among them, there are a large number of missed detections in the SCASC and GIRMRF methods, so they have a relatively high FN value. In contrast, the present invention can obtain lower FP (marked as green pixels) and FN (marked as red pixels) pixels by aligning the images, and thus has relatively high OA, KC, and F1 values.

[0275] This relatively high performance of the present invention benefits from the unsupervised learning method adopted. Through the feature difference minimization constraint and the cross-reconstruction constraint, dual-temporal heterogeneous images are used in the training stage, enabling the network to not only fully represent the input data but also effectively highlight the changed regions. As a result, the changes between images can be effectively detected, structural distortions and non-essential changes can be suppressed, and the change detection performance of heterogeneous images is significantly improved. Therefore, on the simulation datasets Gloucester, Tianhe, Shuguang, and Sardinia, the present invention achieves the highest AUC scores, verifying that the present invention has good detection performance.

[0276] 5) Figures 13 to 20 As shown, if the registration error between the dual-temporal images is large, combined with the interference of the background of heterogeneous images, it may lead to significant differences and may cause the detection algorithm to misjudge objects that have not changed as changed objects; for example Figure 17 in the field parts of (b) and (c) in Figure 18 are misidentified in (a) and (b) in

[0277] III. Device, Storage Medium, Program Product

[0278] 1. Based on the same inventive concept as the above-mentioned unsupervised image change detection method, the present application also provides an electronic device, which includes a processor and a memory. The memory stores computer-readable code. When the computer-readable code is executed by the processor, the unsupervised image change detection method of the present invention is implemented.

[0279] Among them, the memory includes a non-volatile storage medium and an internal memory; the non-volatile storage medium can store an operating system and computer-readable code. The computer-readable code includes program instructions, and when the program instructions are executed, the processor can be made to execute the unsupervised image change detection method. The processor is used to provide computing and control capabilities to support the operation of the entire electronic device. The memory provides an environment for the operation of the computer-readable code in the non-volatile storage medium. When the computer-readable code is executed by the processor, the processor can be made to execute the unsupervised image change detection method.

[0280] It should be understood that the processor can be a central processing unit, other general-purpose processors, digital signal processors, application-specific integrated circuits, field-programmable gate arrays or other programmable logic devices, transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor.

[0281] 2. The present application also provides a readable storage medium, which can be an internal storage unit of the electronic device described in the foregoing embodiment, such as the hard disk or memory of the computer device. The readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, a smart memory card, a secure digital card, etc. equipped on the electronic device.

[0282] 3. The present application also provides a computer program product, including a computer program or instructions, which implement the unsupervised image change detection method of the present invention when executed by the processor.

[0283] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.

[0284] The present invention is not limited to the above-mentioned embodiments. Without departing from the essence of the present invention, any obvious improvements, substitutions or deformations that those skilled in the art can make all belong to the protection scope of the present invention.

Claims

1. An unsupervised image change detection method, characterized in that: The following image change detection network is adopted, including: A bilateral spatial transformation difference network, which is provided with a bidirectional spatial transformation network, a bilateral feature encoding-decoder, and a bilinear sampler, where: The bidirectional spatial transformation network is used to learn the spatial offset parameters between the two-temporal images X and Y; The bilateral feature encoding-decoder is used to extract the respective depth features of the dual-temporal images X and Y, and generate corresponding reconstructed images based on the respective depth features. and The bilinear sampler aligns the images and features in the spatial domain based on the spatial offset parameters; Generate an initial feature difference map D based on the aligned depth features, and generate an initial binary change map P based on the initial feature difference map D 0 ; A change map iterative correction network iteratively optimizes the binary change map based on the aligned images.

2. The unsupervised image change detection method according to claim 1, wherein: In the two-way spatial transformation network, the spatial coordinate offset B between the two temporal-phase images is obtained through the localization network ctr : B ctr = Lnet([X; Y], S ctr ) where [X; Y] represents concatenating images X and Y along the channel dimension, Lnet(·) represents the localization network, and S ctr is the control point; The positioning network has two layers, where: B1 = Tanh(MaxPool2d 5×5 (Conv 25×25 ([X; Y]))) B2 = Tanh(MaxPool2d 5×5 (Conv 5×5 (B1))) B ctr = Lnet([X; Y], S ctr ) = Linear(AdaptiveAvgPool 24×24 (B2)) where B1 represents the output of the first layer, B2 represents the output of the second layer, Conv 25×25 (·) represents a 25×25 convolution, MaxPool2d 5×5 (·) represents a 5×5 max pooling, Tanh(·) represents the Tanh activation function, Conv 5×5 (·) represents a 5×5 convolution, AdaptiveAvgPool 24×24 (·) represents a 24×24 adaptive average pooling, Linear(·) represents a linear fully connected layer; Based on the control point S ctr and the spatial coordinate offset B ctr , learn the spatial offset parameters between the dual-temporal images X and Y: In the formula, represents the control points for the alignment transformation from image X to image Y, represents the control points for the alignment transformation from image Y to image X, represents the spatial transformation function for the alignment of image X to image Y, and the subscript represents the learnable parameters in the said positioning network, is the grid sampling point after the alignment transformation on image X, and Grid(·) represents the grid generator, where the sampling point coordinates are generated by the TPS interpolation method, is the grid sampling point after the alignment transformation on image Y, represents the spatial transformation function for the alignment of image Y to image X.

3. The unsupervised image change detection method according to claim 2, characterized in that: The bilateral feature encoding-decoder includes a bilateral feature encoder and a bilateral feature decoder, where: Bilateral feature encoder, consisting of two encoders E1(·) and E2(·) with the same structure. Each encoder is provided with L cascaded convolutional layers to extract depth features F from images X and Y respectively X and F Y : where θ1 and θ2 are the learnable parameters in the encoders E1(·) and E2(·), respectively are the convolution kernel weights and biases of the l-th layer in the encoder E1(·), respectively are the convolution kernel weights and biases of the l-th layer in the encoder E2(·), respectively, σ(·) is the activation function, and "*” represents the convolution operation denotes the concatenation of the convolutional layers from layer 1 to layer L Bilateral feature decoder, consisting of two decoders D1(·) and D2(·) with the same structure. Each decoder is composed of multiple cascaded convolutional layers, and decodes through the deep features F X and F Y to obtain the reconstructed images and In the formula, ζ1 and ζ2 are respectively the learnable parameters in the decoders D1(·) and D2(·); The bilinear sampler aligns images and features according to and : where S(·) represents a differentiable bilinear sampler, X ′ is the image after aligning image X to image Y, Y ′ is the image after aligning image Y to image X, F X ′ is the feature after aligning feature F X to feature F Y and F Y ′ is the feature after aligning feature F Y to feature F X and is the aligned feature.

4. The unsupervised image change detection method according to claim 3, wherein: In the encoder, there are 5×5 convolution, 1×1 convolution, 5×5 convolution, and 1×1 convolution in sequence. After each convolution, it is activated by the Tanh activation function, and finally the depth features are output after layer normalization; In the decoder, there are 4 consecutive 1×1 convolutions, and after each convolution, it is activated by the Tanh activation function.

5. The unsupervised image change detection method according to claim 3, wherein: In the bilateral spatial transformation difference network, the following objective function is provided: where (i, j) represents the spatial coordinates of the pixel point, F Y (i, j), F X (i, j), X(i, j), M2(i, j), Y(i, j), M1(i, j) represent respectively F Y 、F X 、X, M2, Y, M1 at the coordinate point (i, j), ‖·‖2 represents the 2-norm, and λ is the weight coefficient.

6. The unsupervised image change detection method according to claim 3, wherein: The initial feature difference map D is obtained through the following formula: where D(i, j) and F X ′ (i, j) respectively represent the values of D and F X ′ at the coordinate point (i, j); The initial binary change map is obtained by performing pixel segmentation on the initial feature difference map D using an automatic multi-threshold segmentation algorithm: P 0 = AMT(D) In the formula, AMT(·) represents the automatic multi-threshold segmentation algorithm; The change iterative correction network is composed of multiple k correction modules, and each correction module respectively includes a feature extraction network and a discriminant network. In the k-th correction module: The feature extraction network has a symmetric bilateral network structure. The network structure on either side is as follows: First, in the first layer, two parallel branches are used to extract image features by 5×5 convolution and 3×3 convolution respectively. After the convolution of the two branches, they are activated by the Tanh activation function, and then the output results of the two branches are concatenated according to the feature channel dimension; The second layer to the third layer are two consecutive 1×1 convolutions, and after each convolution, they are activated by the Tanh activation function; The fourth layer first undergoes 1×1 convolution and then is activated by the Sigmoid activation function; Finally, it passes through the convolutional block attention module and L2 normalization in sequence; In the feature extraction network, the two symmetric sides of the network respectively input the image X ′ and the image Y, and the Euclidean distance is calculated between the obtained outputs to obtain a feature difference map In the discrimination network, according to the feature difference map Output a binary classification result: In the formula, is the probability change graph output by the k-th correction module, represents the discriminant network, φ is the learnable parameter in the discriminant network, Tanh(·) represents the Tanh activation function, Conv 1×1 (·) represents a 1×1 convolution, Conv 3×3 (·) represents a 3×3 convolution; For Perform binarization processing using the following dynamic selection threshold strategy to obtain the binary change diagram P output by the k-th correction module k : The dynamic selection threshold strategy is as follows: Set a group of thresholds, and perform a round of binarization processing on the change graph in turn based on each threshold ; In each round of binarization, calculate the Euclidean distance between the binary change map obtained in the current round and the output result of the previous round. If: ① The Euclidean distance ≥ the distance judgment parameter Fbefore, then use the output result of the previous round as the output of this round; ② The Euclidean distance < the distance judgment parameter Fbefore, then use the binary change map obtained in this round as the output of this round, and update the value of the distance judgment parameter Fbefore to the Euclidean distance; Among which P k-1 As the initial binary change diagram compared with the first round, the initial value of Fbefore is given manually; Until all rounds of iteration end, output the binarized change map.

7. The unsupervised image change detection method according to claim 6, wherein: The objective function of the k-th correction module is: In the formula, are the learnable parameters in the feature extraction network, α is the boundary parameter used to control the change map, and ω k-1 is the dynamic weight, and P k-1 (i, j), ω k-1 (i, j) represent the values of P k-1 and ω k-1 at the coordinate point (i, j); the expression of the dynamic weight ω k-1 is: In the formula, n and γ are weight parameters; The above objective function is optimized through the following alternating iteration method: Among them, during optimization for the sub-problem, the backpropagation algorithm is used to optimize the learnable parameters in the feature extraction network to obtain the required feature difference map after optimization During optimization for the sub-problem, fix the optimized iteratively update the learnable parameters φ of the discriminative network to obtain the optimal solution 8. A computer device, characterized in that: It includes a memory and a processor; The memory is used to store computer programs; The processor is used to execute the computer program and implement the unsupervised image change detection method as described in any one of claims 1 to 7 when executing the computer program.

9. A computer-readable storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a processor, the processor is caused to execute the unsupervised image change detection method according to any one of claims 1 to 7.

10. A computer program product, characterized in that: It includes a computer program, and when the computer program is executed by a processor, the unsupervised image change detection method according to any one of claims 1 to 7 is implemented.