Unsupervised heterogeneous remote sensing image change detection method based on bilateral multi-head image attention self-encoding network

Through the method based on the bilateral multi-head graph attention-encoding network, the problem of inaccurate detection of changing areas caused by data inconsistency in heterogeneous remote sensing images is solved, and the accurate detection of changing areas of heterogeneous remote sensing images is achieved and the detection accuracy is improved.

CN120014479APending Publication Date: 2025-05-16XIAN UNIV OF TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510078096.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Due to data inconsistency, heterogeneous remote sensing images cause inaccurate detection of change areas, which has problems such as false detection and reduced detection accuracy.

Method used

Unsupervised heterogeneous remote sensing image change detection method based on bilateral multi-head graph attention autocoding network is adopted. By constructing a bilateral convolutional autocoding network, a multi-head graph attention mechanism is introduced, and a training is carried out in combination with reconstruction loss, structural consistency loss, domain-related loss and adversarial loss function, and a difference graph is generated for change detection.

Benefits of technology

Accurate detection of the changing areas of heterogeneous remote sensing images is achieved, the interference of pseudo-change is reduced, the accuracy of change detection is improved, spatial dependencies can be adaptively captured, and image change feature learning is carried out under unsupervised conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014479A_ABST
    Figure CN120014479A_ABST
Patent Text Reader

Abstract

The invention discloses an unsupervised heterogeneous remote sensing image change detection method based on a bilateral multi-head image attention self-encoding network. The method specifically comprises the following steps: step 1, constructing a bilateral convolution self-encoder network by using two groups of encoder-decoder pairs; step 2, respectively constructing reconstruction loss and structural consistency loss, introducing a multi-head image attention mechanism into the processing of the auto-encoder network, and constructing domain correlation loss based on the multi-head image attention mechanism; 3, training the network to generate a difference graph; and 4, analyzing the difference image generated in the step 3 through a threshold analysis method, marking a change region as 1, and marking a non-change region as 0, thereby realizing a change detection task of the heterogeneous remote sensing image. According to the method, the problem of inaccurate detection of the change area caused by data inconsistency in the heterogeneous remote sensing image is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and relates to an unsupervised heterogeneous remote sensing image change detection method based on a bilateral multi-head graph attention autoencoder network. Background Art

[0002] Remote sensing image change detection is to analyze dual-phase remote sensing images obtained from the same geographical area at different times, so as to accurately obtain the changes in the observed area (Mercier et al., 2008). At present, this technology has been widely used in remote sensing information processing (Gong Maoguo et al., 2016), computer vision (Radke et al., 2005), urban building monitoring (Li et al., 2019), deforestation (Brunner et al., 2010) and land cover change (Griffiths et al., 2010; Pantze et al., 2010). With the rapid growth of the number of remote sensing satellites and the rapid development of imaging technology, a large number of remote sensing images of different types can be obtained. Heterogeneous remote sensing image change detection technology has become the fastest and most suitable technical means to comprehensively use a variety of remote sensing image data to determine and analyze surface changes, and it is more in line with practical application needs. Therefore, designing a fast, automatic and unsupervised heterogeneous remote sensing image change detection method has very important practical significance and broad application prospects.

[0003] Traditional change detection methods rely on homogeneous data, i.e., a set of images acquired by the same sensor under the same geometry, seasonal conditions, and recording settings. However, these limitations are too stringent for many practical examples. First, the satellite revisit period sets an upper limit on the temporal resolution when monitoring long-term trends and a lower limit on the response time when assessing damage from sudden events. In addition, even if two images are collected with the same configuration, the same feature in the dual-temporal remote sensing images often exhibits significantly different appearances or features due to interference from other objective factors, such as imaging noise, sensor resolution, multi-view imaging, environmental factors (such as illumination, weather, season), illumination conditions for optical data, or humidity and precipitation for Synthetic Aperture Radar (SAR), making the images possibly non-homogeneous. Therefore, the background noise caused by the obvious feature differences in heterogeneous remote sensing images affects the detection of real change areas, resulting in false detections and reduced detection accuracy. Summary of the invention

[0004] The purpose of the present invention is to provide an unsupervised heterogeneous remote sensing image change detection method based on a bilateral multi-head graph attention autoencoder network, which solves the problem of inaccurate detection of changed areas in heterogeneous remote sensing images due to data inconsistency.

[0005] The technical solution adopted by the present invention is an unsupervised heterogeneous remote sensing image change detection method based on a bilateral multi-head graph attention autoencoder network, which specifically includes the following steps:

[0006] Step 1: Use two sets of encoder-decoder pairs to construct a bilateral convolutional autoencoder network.

[0007] Step 2: Based on the network constructed in step 1, reconstruction loss and structural consistency loss are constructed respectively, and a multi-head graph attention mechanism is introduced in the processing of the autoencoder network, based on which domain-related loss is constructed;

[0008] Step 3: Consider the bilateral autoencoder network constructed in step 1 as a generator, construct the adversarial loss function of the corresponding discriminator, train the bilateral autoencoder network with the loss function constructed in step 2, and use the trained network to generate a difference map;

[0009] Step 4: Analyze the difference map generated in step 3 by threshold analysis method, mark the changed area as "1" and the non-changed area as "0", so as to achieve the change detection task of heterogeneous remote sensing images.

[0010] The present invention is also characterized in that:

[0011] In step 1, the bilateral adversarial autoencoder network model consists of two sets of encoder-decoder sub-networks including:

[0012] Encoder E1(X):X→Z1 and the corresponding decoder D l (Z): Right now:

[0013] Encoder E2(Y): Y→Z2 and corresponding decoder D2(Z): Right now:

[0014] Among them, X and Y represent the two registered dual-time remote sensing images of the input, Z1 and Z2 represent the encoding layer or latent space of encoders E1 and E2 respectively, and Z is the encoding layer representation of the tensor in Z1 and Z2. Indicates that X passes through E1 and D l The reconstructed image of Represents the reconstructed image of Y after E2 and D2. The two sets of encoder-decoder pairs are constructed by a deep fully convolutional network.

[0015] The specific process of step 2 is:

[0016] Step 2.1, the first regularization term of the objective function of the bilateral convolutional autoencoder network is used as the reconstruction loss term, and the reconstruction loss term L recon It is expressed as:

[0017]

[0018] Among them, I1 is a randomly extracted image block from image X, and I2 is a randomly extracted image block from image Y. It is I1 through E l and D l The reconstructed image generated is is the reconstructed image of I2 generated by E2 and D2, and M is the mean square error;

[0019] Step 2.2, randomly extract image block I1 from image X, use encoder E1 to convert image block I1 to latent space, then use decoder D2 corresponding to image Y to decode the features in the latent space and generate an image block with the same style as Y As shown in the following formula (4); randomly extract image block I2 from image Y, use encoder E2 to convert it, and then use decoder D1 corresponding to image X to decode it to generate an image block with the same style as X As shown in the following formula (5); To ensure that the pixels not affected by the change before and after the transformation remain consistent, the structural consistency loss function L is used struct To constrain this conversion process:

[0020]

[0021] Among them, δ(·,·|π) is a weighted loss function, and the elements of the prior weight π is the probability that pixel i∈{1,…,n} does not change; π i Stored in the matrix Π∈[0,1] H×W In the figure, H is the height of the image and W is the width of the image;

[0022] The prior weight matrix Π stores the unchanged probabilities of all pixels. Initially, the probabilities of all pixels are set to 0. After each round of training, Π is updated by calculating the difference image Δ:

[0023] Π=1-Δ(7)

[0024]

[0025] Among them, Δ is an indicator to measure the degree of pixel change, ranging from [0,1], c x represents the number of channels in image X, c y represents the number of channels in the image Y, and d(·) represents a pixel-level distance metric;

[0026] Step 2.3, calculate the domain-dependent loss L dcl .

[0027] The specific process of step 2.3 is:

[0028] Given input images X and Y, use encoders E1 and E2 to map the input images X and Y to the latent space domain Z, respectively, and obtain two feature representations Z1 and Z2. Consider each 1×1×c grid in Z1 and Z2 as a node, and generate a graph K, where c is the number of feature channels. Let V1 represent a set of nodes containing all nodes of Z1, and V2 represent a set of nodes containing all nodes of Z2. Construct a complete bipartite graph K = (V, E) to model the node-level relationship between X and Y, where V = V1∪V2 and For each (m,n)∈E, e mn represents the correlation coefficient between m∈V1 and n∈V2, where and Represents the feature vector of node m and node n in the latent space domain, and the calculation formula is as follows:

[0029]

[0030] in, and are the linear transformations applied to each node in Z1 and Z2, respectively,

[0031] Based on the correlation coefficient e mn , perform self-attention mechanism on each node in V2 and define a mn To represent the influence of node n on node m, the correlation coefficient e is calculated by using the SoftMax function mn After normalization, we get the following formula (10):

[0032]

[0033] Where k represents a node in the node set V2, which means traversing all nodes. p represents the source node, e pk is defined as being proportional to the similarity of the feature vectors called nodes p and k; in order to propagate attention information from all nodes in V2 to the mth node in V1, a linear transformation matrix is ​​applied to each node feature in V2 The mathematical formula for the aggregation representation of node m becomes As shown in the following formula (11):

[0034]

[0035] By expressing the aggregation With node features Fusion, get feature representation As shown in the following formula (12):

[0036]

[0037] Among them, ∥ represents the vector concatenation operator, which calculates all Get a new feature map Through the following formulas (13)-(16), a new feature map of image Y is calculated based on X

[0038]

[0039] Among them, g represents a node in the node set V1, which means traversing all nodes, j represents the source node, and e nm is defined as being proportional to the similarity between the feature vectors of node n and node m; e jg is defined as being proportional to the similarity between the feature vectors of node j and node g; a linear transformation matrix is ​​applied to each node feature in V1 The mathematical formula for the aggregate representation of node n becomes:

[0040]

[0041] By expressing the aggregation With node features Fusion, get feature representation As shown in the following formula (16):

[0042]

[0043] Compute all Get a new feature map

[0044] Add O1 attention heads o1 to V2, repeat formulas (9)-(12), and finally calculate the feature map obtained by each attention head o1 Put together and get

[0045]

[0046] Add O2 attention heads o2 to V1, repeat formulas (13)-(16), and finally calculate the feature map obtained by each attention head o2 Put together and get

[0047]

[0048] Domain-dependent loss L dcl The calculation is as follows:

[0049]

[0050] Among them, d2(·) is a pixel-level distance metric function.

[0051] The specific process of step 3 is: input the image blocks I1 and I2 of the dual-time images X and Y into the bilateral convolutional autoencoder network constructed in step 1, and train the constraint network through the loss function to generate images with similar styles. and After training is completed, the dual-time images X and Y are input into the network, and finally the difference map Δ is calculated based on the generated images.

[0052] In step 3, the loss function used in training is the reconstruction loss L recon , consistency loss L struct , domain-dependent loss L dcl And the adversarial loss function L adver .

[0053] In step 3, the adversarial loss function L adver The construction process is: input data pair and I1, and I2 are sent to the discriminant network, and the converted image is distinguished by the discriminator P1 Pixel distribution of the real image I1, and the difference between the converted image and the discriminator P2 The pixel distribution of the real image I2 is used to construct the adversarial loss function L between the discriminator P1 and the discriminator P2 adver , the formula is as follows:

[0054]

[0055] Among them, L GAN,1→2 is the adversarial loss for converting images from style 1 to style 2, L GAN,2→1 is the adversarial loss for converting an image from style 2 to style 1; the total loss L total It is expressed as follows:

[0056] L total =L recon +αL struct +βL dcl +δL adver (twenty three)

[0057] Where, the weights α, β, and δ are hyperparameters used to balance the four loss terms.

[0058] The specific process of step 4 is as follows: after the network training is completed, the image and are put into the current network to obtain the final difference map Δ. The difference image Δ is obtained using formula (8), and the classic Otsu method is used to obtain the final change detection result map CM. After obtaining the final threshold threshold, all pixel values ​​are compared with the threshold threshold. If it is greater than the threshold, the pixel is considered to belong to the change class; if it is less than the threshold, the pixel belongs to the unchanged class, as shown in the following formula (24):

[0059]

[0060] The beneficial effects of the present invention are as follows:

[0061] 1) Accurately detect changed areas: Through the design of adversarial loss, the interference of pseudo-changes on detection results is reduced, thus improving the accuracy of change detection.

[0062] 2) Adaptively capture spatial dependencies: The multi-head graph attention mechanism can capture detailed features between images, especially showing better robustness when processing heterogeneous remote sensing images.

[0063] 3) Completely unsupervised: Without the need for a large amount of labeled data, the network can learn image change features in an unsupervised manner, which has better versatility. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 It is a framework diagram of the unsupervised heterogeneous remote sensing image change detection method based on the bilateral multi-head graph attention mechanism autoencoder network of the present invention;

[0065] Figure 2 It is the process of processing data by a single-side convolutional autoencoder;

[0066] Figure 3 It is the image before the change of the Italy dataset;

[0067] Figure 4 This is the changed image of the Italy dataset;

[0068] Figure 5 This is the image before the change of the California dataset;

[0069] Figure 6 This is the changed image of the California dataset;

[0070] Figure 7 This is a comparison of the change graphs obtained on the Italy dataset;

[0071] Figure 8 This is a comparison of the change graphs obtained on the California dataset. DETAILED DESCRIPTION

[0072] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0073] Example 1

[0074] The present invention is based on the unsupervised heterogeneous remote sensing image change detection method of the bilateral multi-head graph attention mechanism autoencoder network. The autoencoder network (Bipartite Multi-head GraphAttention Autoencoder Network, BMGAAN) combined with the bilateral multi-head graph attention mechanism is used to realize the unsupervised heterogeneous remote sensing image change image, which is used to deal with the remote sensing image change detection problem. The framework diagram is as follows Figure 1 The specific steps include:

[0075] Step 1: Use two sets of encoder-decoder pairs to construct a bilateral convolutional autoencoder network (BMGAAN).

[0076] Step 2: Based on the network constructed in step 1, reconstruction loss and structural consistency loss are constructed respectively, and a multi-head graph attention mechanism is introduced in the processing of the autoencoder network, based on which domain-related loss is constructed.

[0077] In step 3, the bilateral autoencoder network constructed in step 1 is regarded as a generator, and the adversarial loss function of the corresponding discriminator is constructed. The bilateral autoencoder network is trained by combining the loss function constructed in step 2, and the trained network is used to generate a difference map.

[0078] Step 4: Analyze the difference map generated in step 3 by threshold analysis method, mark the changed area as "1" and the non-changed area as "0", so as to achieve the change detection task of heterogeneous remote sensing images.

[0079] Example 2

[0080] In step 1, the bilateral adversarial autoencoder network model consists of two sets of encoder-decoder sub-networks including:

[0081] Encoder E1(X):X→Z1 and the corresponding decoder D l (Z): Right now:

[0082] Encoder E2(Y): Y→Z2 and corresponding decoder D2(Z): Right now:

[0083] Among them, X and Y represent the two registered dual-time remote sensing images of the input, Z1 and Z2 represent the encoding layer or latent space of encoders E1 and E2 respectively, and Z is the encoding layer representation of the tensor in Z1 and Z2. Indicates that X passes through E1 and D l The reconstructed image of Represents the reconstructed image of Y after E2 and D2. The two sets of encoder-decoder pairs are constructed by a deep fully convolutional network, where the encoder consists of two 3×3×50 convolutional layers plus a LeakyReLU activation function, followed by a 3×3×3 convolutional layer plus a Tanh activation function, and finally obtains a hidden layer feature Z of size 100×100×3; the last layer of the decoder is a 3×3×C convolutional layer plus a Tanh activation function, where C is the number of channels of the image, which restores the features to the image data size of the corresponding domain space.

[0084] Example 3

[0085] The specific process of step 2 is:

[0086] The first regularization term of the BMGAAN network objective function is the reconstruction loss term, which aims to force the reconstructed output image obtained after the network encoding and decoding to be as close to the input image as possible. Therefore, the reconstruction loss term L recon It can be expressed as:

[0087]

[0088] Among them, I1 is a randomly extracted image block from image X, and I2 is a randomly extracted image block from image Y. It is I1 through E l and D l The reconstructed image generated is is the reconstructed image of I2 generated by E2 and D2, and M is the mean square error (MSE).

[0089] For image X, randomly extract image block I1 from it, then use encoder E1 to convert it to the latent space, and then use decoder D2 corresponding to image Y to decode the features in the latent space to generate an image block with the same style as Y. As shown in the following formula (4). For image Y, randomly extract image block I2 from it, use encoder E2 to convert it, and then use decoder D1 corresponding to image X to decode it to generate an image block with the same style as X As shown in the following formula (5); Figure 2 As shown in the figure, it describes the process of processing data by the one-sided convolutional self-encoder, where w, h are the width and height of the image block I1, C1, C2 are I1 and To ensure that the pixels not affected by the change should remain consistent before and after the transformation, and the pixels affected by the change should remain relatively different, we use the structural consistency loss function L struct To constrain this conversion process:

[0090]

[0091] Among them, δ(,·|π) is a weighted loss function, whose weight is determined by the prior weight π. The elements of the prior weight π can be interpreted as the probability that a pixel i∈{1,…,n} remains unchanged. i Stored in the matrix Π∈[0,1] H×W In the figure, H is the height of the image and W is the width of the image.

[0092] The prior weight matrix Π stores the unchanged probabilities of all pixels. Initially, the probabilities of all pixels are set to 0. After each round of training, Π is updated by calculating the difference image Δ:

[0093] Π=1-Δ (7)

[0094]

[0095] Among them, Δ is an indicator to measure the degree of pixel change, ranging from [0,1], c x represents the number of channels in image X, c y represents the number of channels in the image Y, and d(·) represents a pixel-level distance metric; due to the computational cost of large-scale data, the Euclidean distance is used. Pixels with large Δ values ​​are more likely to change, and their assigned weight π is smaller. On the contrary, pixels with small Δ values ​​are more likely to remain unchanged, and their assigned weight π is larger. This update method ensures that the network pays attention to unchanged pixels, while giving smaller weights to changed pixels to reduce their impact on network training. This enables the model to map unchanged areas more accurately without being disturbed by changed areas, thereby improving the robustness of the model.

[0096] Example 4

[0097] Based on Example 3, in order to align the representation of the image, a structure based on the Multi-head Graph Attention Mechanism (MHGAM) is proposed. The multi-head attention mechanism is used to mine the potential spatial feature relationship in the spatial domain. This method does not rely on manual annotation and does not require specific prior knowledge. It is assumed that the unchanged area can represent more spatial correlations, thereby improving the feature expression effect. The implementation method of each module will be described in detail below.

[0098] Given input images X and Y, we first map the input images X and Y to the latent space domain Z using encoders E1 and E2, obtaining two feature representations Z1 and Z2, respectively. Then, we consider each 1×1×c grid in Z1 and Z2 as a node, generating a graph K, where c is the number of feature channels. Let V1 denote a set of nodes containing all nodes of Z1, and V2 denote a set of nodes containing all nodes of Z2. We construct a complete bipartite graph K = (V, E) to model the node-level relationship between X and Y, where V = V1∪V2 and pass and We further define two subgraphs of K. For each (m,n)∈E, e mn represents the correlation coefficient between m∈V1 and n∈V2. and represents the feature vector of node m and node n in the latent space domain. The more similar the features are between the images before and after the change, the more likely it is an unchanged region and more spatial correction information should be transferred to it. Therefore, e mn is defined as being proportional to the similarity of the feature vectors called node m and node n. Here, the inner product of the feature vectors is used to calculate the similarity. In order to adaptively capture a better representation between the features of two nodes, the node features are linearly transformed and then the correlation score is calculated by the inner product:

[0099]

[0100] in, and are the linear transformations applied to each node in Z1 and Z2, respectively. They can also be thought of as two weight matrices, initialized randomly and learned in each round of training.

[0101] Based on the correlation coefficient e mn , perform self-attention mechanism on each node in V2. Define a mn To represent the influence of node n on node m, this is achieved by using the SoftMax function to calculate the correlation coefficient e mn Normalized to get:

[0102]

[0103] Where k represents a node in the node set V2, and p represents the source node. pk is defined as being proportional to the similarity of the feature vectors of node p and node k. That is, a mnThe attention paid to node m is measured according to the view of node n. To propagate the attention information from all nodes in V2 to the mth node in V1, a linear transformation matrix is ​​applied to each node feature in V2 Then, the mathematical formula for the aggregate representation of node m becomes:

[0104]

[0105] Therefore, by expressing the aggregation With node features Fusion, a more powerful feature representation As shown in the following formula (12):

[0106]

[0107] Here, ∥ represents the vector concatenation operator. This attention mechanism is a single feed-forward neural layer. All It is called a new feature map Similarly, a new feature map of image Y can be calculated based on X by using the following formulas (13) to (16):

[0108]

[0109] Where g represents a node in the node set V1, indicating traversal of all nodes. j represents the source node. jg is defined to be proportional to the similarity between the feature vectors of node j and node g.

[0110]

[0111] Among them, a linear transformation matrix is ​​applied to each node feature in V1

[0112]

[0113] Among them, all the It is called a new feature map At this point, the feature map is completed and However, in heterogeneous remote sensing image change detection, the image features before and after the change may be quite different, and only updating the feature map once will miss the feature information. Therefore, a multi-head graph attention mechanism (MHGAM) is designed. The approach is to add O1 attention heads o1 to V2, repeat formulas (9)-(12), and finally update the feature map obtained by each attention head o1. Put together and get As the final output of MHGAM:

[0114]

[0115] Repeating formulas (13)-(16) adds O2 attention heads o2 to V1, and finally the feature map obtained by each attention head o2 is Putting them together, we can get

[0116]

[0117] MHGAM allows the model to learn features in different spaces. It processes input features in parallel through multiple independent attention heads. Each attention head can learn different feature representations and relationships, thereby enhancing the expressiveness of the model. Therefore, the domain-dependent loss L dcl The calculation is as follows:

[0118]

[0119] Among them, d2(·) is a pixel-level distance metric function. In order to meet the limitation of computational cost, Euclidean distance is preferred. Using domain correlation loss L dcl , establishes the node matching relationship between feature images, thereby reducing the impact of spectral feature differences of image pairs on target area analysis. By mapping spatial information from one feature domain to another through MHGAM, a more effective image embedding representation is obtained, so that the heterogeneity between cross-domain image pairs is better handled.

[0120] Example 5

[0121] The specific process of step 3 is as follows: the bilateral autoencoder network constructed in step 1 can be regarded as a pair of generators F and G, which are used to convert between heterogeneous data to generate images with more consistent styles. However, the images generated by the convolutional autoencoder network alone may not fully achieve the expected style transfer effect, so a discriminator needs to be introduced to judge the style consistency between the generated image and the original image. By jointly training the generator and the discriminator, the ideal style transfer goal can be gradually achieved. A simple solution is to convert the input data pair ( and I1, and I2) are sent to the discriminant network, and the converted image is distinguished by the discriminator P1 Pixel distribution of the real image I1, and the difference between the converted image and the discriminator P2 The pixel distribution of the real image I2 is used to construct the adversarial loss function of the discriminator P1 and the discriminator P2, forcing the generated image style to be more consistent. Mathematically, it can be expressed as:

[0122]

[0123] Among them, L GAN,1→2 is the adversarial loss for converting images from style 1 to style 2, L GAN,2→1 is the adversarial loss for converting an image from style 2 to style 1. The model adjusts the generators F and G to minimize these losses, while adjusting the discriminators P1 and P2 to maximize these losses.

[0124] Total loss L total It is expressed as follows:

[0125] L total =L recon +αL stuct +βL dcl +δL adver (twenty three)

[0126] Where, the weights α, β, and δ are hyperparameters used to balance the four loss terms.

[0127] The image blocks I1 and I2 of the dual-time images X and Y are input into the network and passed through four loss functions L recon , L struct , L dcl and L adver Train the constraint network to generate images with similar styles and I1, After training, the dual-time images X and Y are input into the network, and finally the generated and X, and Y calculate the difference map Δ, correspond X corresponds to I1, correspond Y corresponds to I2.

[0128] Example 6

[0129] In step 4, after the network training is completed, the images and are input into the current network to obtain the final difference map Δ.

[0130] The difference image Δ is obtained using formula (8). The classic Otsu method is used to obtain the final change detection result map CM. The key to this threshold selection algorithm is to obtain the optimal threshold. After obtaining the final threshold threshold, all pixel values ​​are compared with the threshold threshold. If it is greater than the threshold, the pixel is considered to belong to the change class, and vice versa.

[0131]

[0132] Example 7

[0133] The effect of the present invention can be specifically illustrated by simulation experiments.

[0134] 1. Experimental Setup

[0135] The entire network model is built using the TensorFlow framework. All experiments are run on a personal computer with a GeForce RTX 2080Ti graphics card, 11GB memory, and 16G RAM. The software environment is Pycharm under Windows 10.

[0136] 2. Experimental content

[0137] First, in the training phase, we used the Adam optimizer, which has the characteristics of adaptive learning rate and minimizes the error through the stochastic gradient descent method. The entire training process lasted 200 cycles. In each cycle, the network receives 10 pairs of randomly generated 100×100 image blocks for learning. In the adversarial training phase, we built a fully connected network as the discriminator. The number of neurons in this network increases with the increase of network depth. The specific configuration is 25, 100, 200, 50, and 1. Then, using the trained network, we generated difference maps on two datasets. The Italian dataset is as follows: Figure 3 , Figure 4 As shown, Figure 3 is the image before the change, Figure 4 is the changed image, the California dataset is as follows Figure 5 , Figure 6 As shown, Figure 5 is the image before the change, Figure 6 is the image after the change. Figure 3 , Figure 4 as well as Figure 5 , Figure 6 is the input image of the network. Finally, the final change map is obtained by the threshold method.

[0138] 3. Experimental results

[0139] Figure 7 (a) is the reference graph of the changes in the Italy dataset. Figure 7 (b) is the change graph obtained by CAA on the Italy dataset. Figure 7 (c) is the change map obtained by BMGAAN on the Italy dataset; TP is the true positive, that is, the number of pixels correctly predicted in the changed area; TN is the true negative, the number of pixels correctly predicted in the unchanged area; FP is the false positive, the number of pixels incorrectly predicted in the changed area; FN is the false negative, the number of pixels incorrectly predicted in the unchanged area. The challenge of change detection in the Italy dataset lies in the different acquisition bands of the images and the rugged mountainous terrain. The drastic changes in the altitude of the mountains lead to color differences in the images, causing pixels in the areas that should remain unchanged to be mistakenly detected as changed areas, resulting in higher false positives, as shown by the green marks in the change map. From Figure 7 (c) It can be seen that the number of FNs of this method is smaller, indicating that it is better than other methods in suppressing background noise.

[0140] Figure 8 (a) is the reference diagram of the California dataset. Figure 8 (b) is the change graph obtained by CAA on the California dataset. Figure 8 (c) is a change map obtained by BMGAAN on the California dataset; change detection in the California dataset faces a significant challenge, namely the flat terrain where floods occur. Although the expansion of the main river channel caused by floods is relatively easy to identify, accurately detecting the small lakes and puddles scattered around the river channel is more challenging. Figure 8 As can be seen in (c), the proposed method demonstrates a strong ability in accurately identifying the river channel expansion area caused by floods.

Claims

1. An unsupervised heterogeneous remote sensing image change detection method based on a bilateral multi-head graph attention autoencoder network, characterized by: The specific steps include: Step 1, construct a bilateral convolutional autoencoder network using two sets of encoder-decoder pairs; Step 2: Based on the network constructed in step 1, reconstruction loss and structural consistency loss are constructed respectively, and a multi-head graph attention mechanism is introduced in the processing of the autoencoder network, based on which domain-related loss is constructed; Step 3: Consider the bilateral autoencoder network constructed in step 1 as a generator, construct the adversarial loss function of the corresponding discriminator, train the bilateral autoencoder network with the loss function constructed in step 2, and use the trained network to generate a difference map; Step 4: Analyze the difference map generated in step 3 by threshold analysis method, mark the changed area as "1" and the non-changed area as "0", so as to achieve the change detection task of heterogeneous remote sensing images.

2. The unsupervised heterogeneous remote sensing image change detection method based on bilateral multi-head graph attention autoencoder network according to claim 1 is characterized in that: In step 1, the bilateral adversarial autoencoder network model consists of two sets of encoder-decoder sub-networks including: Encoder E1(X):X→Z1 and the corresponding decoder Right now: Encoder E2(Y):Y→Z2 and the corresponding decoder Right now: Among them, X and Y represent the two registered dual-time remote sensing images of the input, Z1 and Z2 represent the encoding layer or latent space of encoders E1 and E2 respectively, and Z is the encoding layer representation of the tensor in Z1 and Z2. Indicates that X passes through E1 and D l The reconstructed image of Represents the reconstructed image of Y after E2 and D2. The two sets of encoder-decoder pairs are constructed by a deep fully convolutional network.

3. The unsupervised heterogeneous remote sensing image change detection method based on bilateral multi-head graph attention autoencoder network according to claim 2 is characterized by: The specific process of step 2 is: Step 2.1, the first regularization term of the objective function of the bilateral convolutional autoencoder network is used as the reconstruction loss term, and the reconstruction loss term L recon It is expressed as: Among them, I1 is a randomly extracted image block from image X, and I2 is a randomly extracted image block from image Y. It is I1 through E l and D l The reconstructed image generated is is the reconstructed image of I2 generated by E2 and D2, and M is the mean square error; Step 2.2, randomly extract image block I1 from image X, use encoder E1 to convert image block I1 to latent space, then use decoder D2 corresponding to image Y to decode the features in the latent space and generate an image block with the same style as Y As shown in the following formula (4); randomly extract image block I2 from image Y, use encoder E2 to convert it, and then use decoder D1 corresponding to image X to decode it to generate an image block with the same style as X As shown in the following formula (5); To ensure that the pixels not affected by the change before and after the transformation remain consistent, the structural consistency loss function L is used struct To constrain this conversion process: Among them, δ(·,·|π) is a weighted loss function, and the elements of the prior weight π is the probability that pixel i∈{1,…,n} does not change; π i Stored in the matrix Π∈[0,1] H×W In the figure, H is the height of the image and W is the width of the image; The prior weight matrix Π stores the unchanged probabilities of all pixels. Initially, the probabilities of all pixels are set to 0. After each round of training, Π is updated by calculating the difference image Δ: Π=1-Δ (7) Among them, Δ is an indicator to measure the degree of pixel change, ranging from [0,1], c x represents the number of channels in image X, c y represents the number of channels in the image Y, and d(·) represents a pixel-level distance metric; Step 2.3, calculate the domain-dependent loss L dcl .

4. The unsupervised heterogeneous remote sensing image change detection method based on bilateral multi-head graph attention autoencoder network according to claim 2 is characterized by: The specific process of step 2.3 is as follows: Given input images X and Y, use encoders E1 and E2 to map the input images X and Y to the latent space domain Z, respectively, and obtain two feature representations Z1 and Z2. Consider each 1×1×c grid in Z1 and Z2 as a node, and generate a graph K, where c is the number of feature channels. Let V1 represent a set of nodes containing all nodes of Z1, and V2 represent a set of nodes containing all nodes of Z2. Construct a complete bipartite graph K = (V, E) to model the node-level relationship between X and Y, where V = V1∪V2 and For each (m,n)∈E, e mn represents the correlation coefficient between m∈V1 and n∈V2, where and Represents the feature vector of node m and node n in the latent space domain, and the calculation formula is as follows: in, and are the linear transformations applied to each node in Z1 and Z2, respectively, Based on the correlation coefficient e mn , perform self-attention mechanism on each node in V2 and define a mn To represent the influence of node n on node m, the correlation coefficient e is calculated by using the SoftMax function mn After normalization, we get the following formula (10): Where k represents a node in the node set V2, which means traversing all nodes. p represents the source node, e pk is defined as being proportional to the similarity of the feature vectors called nodes p and k; in order to propagate attention information from all nodes in V2 to the mth node in V1, a linear transformation matrix is ​​applied to each node feature in V2 The mathematical formula for the aggregation representation of node m becomes As shown in the following formula (11): By expressing the aggregation With node features Fusion, get feature representation As shown in the following formula (12): Among them, ∥ represents the vector concatenation operator, which calculates all Get a new feature map Through the following formulas (13) to (15), a new feature map of image Y is calculated based on X Among them, g represents a node in the node set V1, which means traversing all nodes, j represents the source node, and e nm is defined as being proportional to the similarity between the feature vectors of node n and node m; e jg is defined as being proportional to the similarity between the feature vectors of node j and node g; a linear transformation matrix is ​​applied to each node feature in V1 The mathematical formula for the aggregate representation of node n becomes: By expressing the aggregation With node features Fusion, get feature representation As shown in the following formula (15): Compute all Get a new feature map Add O1 attention heads o1 to V2, repeat formulas (10)-(12), and finally calculate the feature map obtained by each attention head o1 Put together and get Add O2 attention heads o2 to V1, repeat formulas (13)-(15), and finally calculate the feature map obtained by each attention head o2 Put together and get Domain-dependent loss L dcl The calculation is as follows: Among them, d2(·) is a pixel-level distance metric function.

5. The unsupervised heterogeneous remote sensing image change detection method based on bilateral multi-head graph attention autoencoder network according to claim 2 is characterized by: The specific process of step 3 is: input the image blocks I1 and I2 of the dual-time images X and Y into the bilateral convolutional autoencoder network constructed in step 1, and train the constraint network through the loss function to generate images with similar styles and After training is completed, the dual-time images X and Y are input into the network, and finally the difference map Δ is calculated based on the generated images.

6. The unsupervised heterogeneous remote sensing image change detection method based on bilateral multi-head graph attention autoencoder network according to claim 5 is characterized by: In step 3, the loss function used in training is the reconstruction loss L recon , consistency loss L struct , domain-dependent loss L dcl And the adversarial loss function L adver .

7. The unsupervised heterogeneous remote sensing image change detection method based on bilateral multi-head graph attention autoencoder network according to claim 6 is characterized by: In step 3, the adversarial loss function L adver The construction process is: input data pair and I1, and I2 are sent to the discriminant network, and the converted image is distinguished by the discriminator P1 Pixel distribution of the real image I1, and the difference between the converted image and the discriminator P2 The pixel distribution of the real image I2 is used to construct the adversarial loss function L between the discriminator P1 and the discriminator P2 adver , the formula is as follows: Among them, L GAN,1→2 is the adversarial loss for converting images from style 1 to style 2, L GAN,2→1 is the adversarial loss for converting an image from style 2 to style 1; the total loss L total It is expressed as follows: L total =L recon +αL struct +βL dcl +δL adver (23) Where, the weights α, β, and δ are hyperparameters used to balance the four loss terms.

8. The unsupervised heterogeneous remote sensing image change detection method based on bilateral multi-head graph attention autoencoder network according to claim 7 is characterized by: The specific process of step 4 is as follows: after the network training is completed, the image and are put into the current network to obtain the final difference map Δ, the difference image Δ is obtained using formula (8), and the final change detection result map CM is obtained using the classic Otsu method. After the final threshold threshold is obtained, all pixel values ​​are compared with the threshold threshold. If the pixel is greater than the threshold, it is considered that the pixel belongs to the change class; if the pixel is less than the threshold, the pixel belongs to the unchanged class, as shown in the following formula (24):

Citation Information

Cited By

  • Multi-task unsupervised change detection method fusing image domain alignment and segmentation network

    CN121725373A

  • Multi-task unsupervised change detection method fusing graph domain alignment and segmentation network

    CN121725373B

  • Bridge health intelligent monitoring operation and maintenance method, system and equipment based on generative reasoning and medium

    CN121745921A

  • Bridge health intelligent monitoring operation and maintenance method, system and device based on generative inference and medium

    CN121745921B