Remote sensing image change detection method based on state-driven discrete diffusion model

Through the state-driven discrete diffusion model, combined with multiple network modules to extract and fuse remote sensing image features, the problem of information balancing in remote sensing image change detection is solved, and the detection accuracy and efficiency are improved.

CN120689767APending Publication Date: 2025-09-23DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510818086.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing remote sensing image change detection methods find it difficult to take into account local information, global information, and edge change information, resulting in low detection efficiency and accuracy.

Method used

Based on the state-driven discrete diffusion model, by constructing the forward diffusion model, deep network model and backward diffusion model, combined with the change detection encoder module, Fourier transform module, feature pyramid network, bidirectional cross attention mechanism and UNet backbone network, multi-scale image features are extracted and feature fusion and backward diffusion processing are performed to obtain change detection results of high-resolution remote sensing images.

Benefits of technology

It effectively takes into account local information, global information and edge change information, improves the accuracy and efficiency of remote sensing image change detection, and realizes the full extraction of high-resolution remote sensing image features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689767A_ABST
    Figure CN120689767A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image change detection method based on a state-driven discrete diffusion model. The method comprises the following steps: acquiring a remote sensing sample image; performing model training on the constructed remote sensing image change detection model based on the state-driven discrete diffusion model through the training set and the verification set to obtain an optimal remote sensing image change detection model; the remote sensing image change detection model comprises a forward diffusion model based on state transition, a deep network model, a back diffusion model and a classifier module; and according to the optimal remote sensing image change detection model, realizing change detection of the dual-temporal high-resolution remote sensing image in the test set. The invention provides a fusion strategy for the collaborative challenge of an existing remote sensing change detection model in the process of integrating local features, global context and edge details, and provides a new solution thought for improving the efficiency and precision of change detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing image change detection, and in particular to a remote sensing image change detection method based on a state-driven discrete diffusion model. Background Art

[0002] Remote sensing image change detection aims to systematically analyze temporal changes by comparing multi-temporal remote sensing images acquired from the same geographic coordinates. Change detection has widespread applications in environmental monitoring, urban planning, and disaster assessment. However, a major challenge in remote sensing image change detection stems from global and local feature differences in multi-temporal remote sensing images caused by factors such as varying illumination and climatic conditions. These differences generate artifacts that mask true ground changes. The rapid development of remote sensing image acquisition technology has facilitated the generation of high-resolution spatiotemporal data. With the increasing availability of large amounts of high-resolution, multi-temporal remote sensing imagery, research interest in remote sensing image change detection has increased significantly. Change detection technology has continued to advance in recent decades. Initially, change detection relied primarily on traditional image processing and statistical methods, which primarily focused on pixel-based and feature-based approaches. While these methods are effective for simple changes, they often underperform in complex scenes.

[0003] Currently, deep neural network algorithms dominate this field, outperforming traditional techniques by effectively learning discriminative features from large-scale, high-quality samples. However, common deep learning-based change detection methods currently use CNNs and Transformers to obtain distinguishable feature representations for classification. CNN-based change detection methods struggle to simultaneously leverage local spatial information and capture long-range dependencies, making it difficult to accurately describe change boundaries. Transformer-based change detection methods often have high computational overhead. Meanwhile, some studies have used continuous diffusion models based on Gaussian noise to address change detection tasks, but the diffusion process of adding and denoising noise results in low model efficiency. In summary, existing network models and methods struggle to balance local, global, and edge change information, resulting in low efficiency and accuracy in remote sensing image change detection. Summary of the Invention

[0004] The present invention provides a remote sensing image change detection method based on a state-driven discrete diffusion model to overcome the above technical problems.

[0005] In order to achieve the above object, the technical solution of the present invention is:

[0006] A remote sensing image change detection method based on a state-driven discrete diffusion model comprises the following steps:

[0007] S1: Acquire dual-temporal high-resolution remote sensing images for change detection;

[0008] The dual-temporal high-resolution remote sensing image dataset is randomly divided into training set, validation set and test set;

[0009] And obtain the ground truth map of the dual-temporal high-resolution remote sensing images in the training set and the validation set, and dimensionally splice the dual-temporal high-resolution remote sensing images with the ground truth map to obtain the remote sensing sample images for model input;

[0010] S2: A remote sensing image change detection model based on a state-driven discrete diffusion model is constructed using a training set and a validation set to train the model and obtain the optimal remote sensing image change detection model; the remote sensing image change detection model includes a forward diffusion model based on state transition, a deep network model, a backward diffusion model, and a classifier module;

[0011] S21: Model training: performing forward diffusion processing on the ground truth image in the remote sensing sample image using a pixel-level state transition strategy constructed by the forward diffusion model to obtain a remote sensing transition image;

[0012] The deep network model includes a change detection encoder module, a Fourier transform module, a feature pyramid network FPN, a bidirectional cross attention mechanism module, and a UNet-based backbone network module;

[0013] The change detection encoder module is used to extract the multi-scale image pixel feature map of the dual-phase high-resolution remote sensing image in the remote sensing sample image, and the multi-scale image pixel feature map is subjected to feature fusion to obtain the multi-scale encoding feature map;

[0014] The multi-scale coded feature map is Fourier transformed by the Fourier transform module to obtain a multi-scale mixed difference feature map containing the frequency domain and spatial domain of the dual-temporal remote sensing image;

[0015] The feature pyramid network FPN is used to fuse the multi-scale mixed difference feature map to obtain a high-dimensional difference feature map;

[0016] The high-dimensional difference feature map is fused with the feature map output by the fourth block module in the UNet-based backbone network module through the bidirectional cross-attention mechanism module, and the feature is enhanced through spatial-channel attention to obtain the cross-attention feature map that integrates the spatial-channel attention features.

[0017] The pixel features of the cross-attention feature map are extracted through the UNet-based backbone network module to obtain the pixel change probability feature map;

[0018] The pixel features in the pixel change probability feature map obtained by the deep network model are back-diffused through the back-diffusion model to iteratively obtain the predicted change probability map of the dual-temporal remote sensing image. The predicted change probability image is binarized through the classifier module to obtain a black and white binary change feature map to obtain the change detection results of the dual-temporal remote sensing image, and then the trained remote sensing image change detection model is obtained.

[0019] S22: Model verification: Based on the constructed total model loss function, the trained remote sensing image change detection model is verified using the validation set to determine whether the output of the trained remote sensing image change detection model converges; if so, the trained remote sensing image change detection model is confirmed to be the optimal remote sensing image change detection model; otherwise, the weight parameters of the trained remote sensing image change detection model are adaptively adjusted based on the back propagation method, and step S21 is repeated;

[0020] S3: Implement change detection on the dual-temporal high-resolution remote sensing images in the test set according to the optimal remote sensing image change detection model.

[0021] Furthermore, the expression of the pixel-level state transition strategy constructed by the forward diffusion model in S21 is:

[0022]

[0023] Where: Indicates that image pixel i, j changes from the initial state After t steps of forward diffusion, the state The conditional transition probability of It represents the image pixel state of the dual-temporal high-resolution remote sensing image at the t-th iteration, that is, the remote sensing transition image at any time is determined based on the image pixel state; Represents the image pixel status of the dual-temporal high-resolution remote sensing image at the initial iteration; represents the set of state transition matrices and Q t represents the state transition matrix at the tth iteration and represents the pixel state transition probability used for pixel state transition and β t represents the pixel state transition probability value at the tth iteration; x t represents the remote sensing transition image corresponding to the tth iteration; i, j represent the image pixel positions of the dual-temporal high-resolution remote sensing image.

[0024] Furthermore, the change detection encoder module includes a first convolution block, a first residual convolution block, a second residual convolution block, a third residual convolution block and a coding fusion module connected in sequence;

[0025] The first convolution block is used to perform a convolution operation on the dual-temporal high-resolution remote sensing image in the remote sensing sample image to obtain a first convolution feature map;

[0026] The first residual convolution block, the second residual convolution block, and the third residual convolution block each include a first two-dimensional convolution layer, a second two-dimensional convolution layer, a first normalization layer, and a first nonlinear activation function layer with different convolution kernels and connected in sequence;

[0027] The first two-dimensional convolutional layer is used to perform a two-dimensional convolution operation on its input;

[0028] The second two-dimensional convolutional layer is used to perform a two-dimensional convolution operation on the output of the first two-dimensional convolutional layer;

[0029] The first normalization layer is used to normalize the output of the second two-dimensional convolutional layer;

[0030] The first nonlinear activation function layer is used to perform nonlinear activation operation on the output of the normalization layer;

[0031] The first residual convolution block is used to perform residual convolution processing on the first convolution feature map to obtain a second convolution feature map; the second residual convolution block is used to perform residual convolution processing on the second convolution feature map to obtain a third convolution feature map; the third residual convolution block is used to perform residual convolution processing on the third convolution feature map to obtain a fourth convolution feature map;

[0032] The coding fusion module is used to splice and fuse the first convolution feature map, the second convolution feature map, the third convolution feature map and the fourth convolution feature map to obtain a multi-scale coding feature map.

[0033] Furthermore, the UNet-based backbone network module includes a first Block module, a second Block module, a third Block module, and a fourth Block module connected in sequence;

[0034] The first Block module, the second Block module, and the third Block module all include a first residual inner convolution block and a convolution downsampling layer with different convolution kernels that are connected in sequence;

[0035] The first residual inner convolution block includes a third two-dimensional convolution, a second normalization layer, a second nonlinear activation function layer, and a fully connected layer connected in sequence;

[0036] The third two-dimensional convolution is used to perform a two-dimensional convolution operation on its input;

[0037] The second normalization layer is used to normalize the output of the third two-dimensional convolution;

[0038] The second nonlinear activation function layer is used to perform a nonlinear activation operation on the output of the second normalization layer;

[0039] The fully connected layer is used to perform a fully connected operation on the output of the second nonlinear activation function layer;

[0040] The convolutional downsampling layer is used to downsample the output of the fully connected layer;

[0041] The first Block module is used to extract pixel features of the multi-scale coding feature map to obtain the fifth convolution feature map; the second Block module is used to extract pixel features from the fifth convolution feature map to obtain the sixth convolution feature map; the third Block module is used to extract pixel features from the sixth convolution feature map to obtain the seventh convolution feature map;

[0042] The fourth Block module includes a second residual inner convolution block and a multi-head attention block that are different from the convolution kernels of other Block modules;

[0043] The second residual inner convolution block is used to perform a convolution operation on the seventh convolution feature map to obtain an eighth convolution feature map; the multi-head attention block is used to perform a multi-head attention operation on the eighth convolution feature map to extract image pixel features to obtain a pixel change probability feature map.

[0044] Furthermore, the method for obtaining a multi-scale mixed difference feature map including frequency domain and spatial domain of a dual-temporal remote sensing image in S21 specifically includes:

[0045] S211: Based on the fast Fourier transform method, a frequency domain transform operation is performed on the multi-scale coding feature maps corresponding to each of the dual-phase high-resolution remote sensing images to obtain two multi-scale frequency domain pixel feature maps corresponding to different phases;

[0046] The pixel feature difference between the two multi-scale frequency domain pixel feature maps is used as the multi-scale frequency domain pixel difference feature, thereby obtaining a multi-scale frequency domain pixel difference feature map;

[0047] S212: Using the pixel feature difference between the two multi-scale image pixel feature maps corresponding to the dual-temporal high-resolution remote sensing image as a multi-scale spatial domain difference feature, thereby obtaining a multi-scale spatial domain difference feature map;

[0048] S213: performing feature map splicing on the multi-scale frequency domain pixel difference feature map and the multi-scale spatial domain difference feature map through a preset splicing layer to obtain a spliced ​​feature map;

[0049] And through the preset two-dimensional convolution module, a two-dimensional convolution operation is performed on the spliced ​​feature map to obtain a multi-scale mixed difference feature map.

[0050] Furthermore, in S21, the pixel features in the pixel change probability feature map are back-diffused by the back-diffusion model to obtain the expression of the predicted change probability map of the dual-temporal remote sensing image:

[0051] The expression for obtaining the predicted change probability map of the dual-temporal remote sensing image is:

[0052]

[0053] Where: Represents the predicted pixel change feature map; Represents the reverse state transition probability matrix at the tth iteration; represents the confidence probability that the image pixel i, j belongs to a certain pixel state at time step t, as predicted by the deep network model with weight parameter θ; represents the conditional transition probability of image pixel i, j from the state at time step t to the state at time step t-1 under the deep network model with weight parameter θ; Represents the remote sensing transition image x corresponding to the tth iteration t The pixel status of each pixel in .

[0054] Furthermore, the construction formula of the total model loss function L in S22 is

[0055] L=L bce +L edge

[0056]

[0057]

[0058] Where: L bce represents the cross entropy loss function; x i The predicted change probability map is represented by The determined probability map; y i represents the ground truth map corresponding to the dual-temporal high-resolution remote sensing image; N represents the total number of samples; L edge represents the edge structure consistency loss function; Respectively represent the Sobel edge gradient amplitude of the predicted change map and its corresponding ground truth map at a certain pixel point; ΔI i They respectively represent the Laplacian response value of the predicted change map and its corresponding ground truth map at a certain pixel point.

[0059] Furthermore, the bidirectional cross attention mechanism module includes a first cross attention mechanism module CrossAttention, a second cross attention mechanism module CrossAttention and a space-channel attention module;

[0060] The first cross-attention mechanism module CrossAttention is used to take the high-dimensional difference feature map as the input of the query vector Q1, and the feature map output by the fourth Block module in the UNet-based backbone network as the input of the key vector K1 and the value vector V1 to perform spatial cross-attention fusion and obtain the first attention feature map;

[0061] And the expression of the first cross attention mechanism module CrossAttention is

[0062]

[0063] Where: Q1 represents the query vector, and its input is the high-dimensional difference feature map; K1, V1 represent the key vector and value vector respectively, and their input is the feature map output by the fourth Block module in the UNet-based backbone network; d k1 represents the dimension of the key vector K1; T is the transpose;

[0064] The second cross-attention mechanism module CrossAttention is used to take the feature map output by the fourth Block module in the UNet-based backbone network as the input of the query vector Q2, and the high-dimensional difference feature map as the input of the key vector K2 and the value vector V2 to perform spatial cross-attention fusion and obtain the second attention feature map;

[0065] And the expression of the second cross attention mechanism module CrossAttention is

[0066]

[0067] Where: Q2 represents the query vector, whose input is the feature map output by the fourth Block module in the UNet-based backbone network; K2, V2 represent the key vector and value vector respectively, whose input is the high-dimensional difference feature map; d k2 represents the dimension of the key vector K2;

[0068] The spatial-channel attention module is used to perform channel splicing on the first attention feature map and the second attention feature map, and perform feature enhancement on the feature map after channel splicing to obtain a cross-attention feature map that integrates the spatial-channel attention features.

[0069] Beneficial effect: The present invention provides a remote sensing image change detection method based on a state-driven discrete diffusion model, which realizes the discrete diffusion process for the dual-phase high-resolution remote sensing image change detection task through the constructed pixel-level state transfer strategy, and extracts the multi-scale image pixel feature map of the dual-phase high-resolution remote sensing image in the remote sensing sample image through the change detection encoder module to obtain the multi-scale coding feature map; Fourier transform is performed on the multi-scale coding feature map to obtain a multi-scale mixed difference feature map, and feature fusion is performed on the multi-scale mixed difference feature map through the feature pyramid network FPN to obtain a high-dimensional difference feature map, which is used to obtain the fused feature map through the bidirectional cross attention mechanism module. The cross-attention feature map of the combined spatial-channel attention feature. The present invention combines a bidirectional cross-attention mechanism module to fuse key features, combines channel attention to emphasize key feature channels, and spatial attention to highlight spatially significant areas, thereby realizing the conditional guidance of the bidirectional cross-attention mechanism on the diffusion model, effectively taking into account local information, global information, and edge change information. Compared with the traditional change detection network, the method described in the present invention realizes change detection by means of a discrete diffusion process, which more fully extracts features from multi-temporal remote sensing images, effectively solves the problem of poor feature extraction effect of high-resolution remote sensing images, and improves the accuracy of remote sensing image change detection. In addition, the use of pixel-level state control for change detection in high-resolution remote sensing images has important theoretical significance. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0071] Figure 1 This is a flow chart of a remote sensing image change detection method based on a state-driven discrete diffusion model according to the present invention;

[0072] Figure 2 Schematic diagram of the structure of the backbone network of the deep network model in this embodiment;

[0073] Figure 3 This is a schematic diagram of an example result of the test data set in this embodiment. DETAILED DESCRIPTION

[0074] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0075] This embodiment provides a remote sensing image change detection method based on a state-driven discrete diffusion model. Figure 1 As shown, the steps include:

[0076] S1: Acquire dual-temporal high-resolution remote sensing images for change detection;

[0077] The dual-temporal high-resolution remote sensing image dataset is randomly divided into training set, validation set and test set;

[0078] And obtain the ground truth map of the dual-temporal high-resolution remote sensing images in the training set and the validation set, and dimensionally splice the dual-temporal high-resolution remote sensing images with the ground truth map to obtain the remote sensing sample images for model input;

[0079] Specifically, the dual-temporal high-resolution remote sensing images in this embodiment are derived from the CLCD dataset, which contains 600 pairs of 512*512 images. Considering the size of video memory and sample size, the CLCD data is cropped to 256*256 images without overlap. Based on the default partitioning method of the dataset, the CLCD dataset is partitioned to obtain 1440 / 480 / 480 pairs of dual-temporal remote sensing images for model training, validation, and testing. Dual-temporal remote sensing images are two sets of high-resolution remote sensing images corresponding to different times within the same shooting area. In the fields of remote sensing image processing and computer vision, a ground truth map is an image used to mark or annotate the true category or state of each pixel in an image. Ground truth maps are usually obtained by manual annotation and provide accurate reference information for network model algorithms to evaluate their performance or serve as labels in supervised learning. In deep learning, especially in semantic segmentation tasks, ground truth maps are used to train networks to identify different objects or regions in images. For example, in the urban landscape segmentation task, the ground truth map will label each pixel in the image as a category such as "building", "road", "vegetation", etc.

[0080] S2: A remote sensing image change detection model based on a state-driven discrete diffusion model is constructed using a training set and a validation set to train the model and obtain the optimal remote sensing image change detection model; the remote sensing image change detection model includes a forward diffusion model based on state transition, a deep network model, a backward diffusion model, and a classifier module;

[0081] S21: Model training: performing forward diffusion processing on the ground truth image in the remote sensing sample using a pixel-level state transition strategy constructed by the forward diffusion model to obtain a remote sensing transition image;

[0082] Specifically, the process of the forward diffusion model is:

[0083]

[0084] Where: β t ∈[0,1],1-β t Corresponding to the pixel state transition probability used in the pixel transition sample image of this embodiment; Indicates the image pixel status, set and Represents the determined pixel state and the uncertain pixel state respectively; Q t Represents a state transfer matrix, Q t The form of explicitly shows the unidirectionality, that is, all pixels in the uncertain pixel state will never be transformed back to the determined pixel state;

[0085] Based on formula (1) and formula (2), the expression of pixel-level state transition strategy is obtained as follows:

[0086]

[0087] Where: Represents the forward diffusion process, that is, the image pixel (i, j) is transferred from the initial state After t steps of forward diffusion, the state The conditional transition probability of Represents the image pixel state of the dual-temporal high-resolution remote sensing image at the t-th iteration, that is, the pixel state of the remote sensing transition image at any time t is determined based on the image pixel state; Represents the image pixel status of the dual-temporal high-resolution remote sensing image at the initial iteration; represents the set of state transition matrices and Q t represents the state transition matrix at the tth iteration and represents the pixel state transition probability used for pixel state transition and β t represents the pixel state transition probability value at the tth iteration; xt represents the remote sensing transition image corresponding to the tth iteration; i, j represent the image pixel positions of the dual-temporal high-resolution remote sensing image.

[0088] This embodiment uses a pixel-level state transition strategy to obtain a remote sensing transition image at an intermediate moment in the diffusion process at any intermediate time step t without the need for stepwise sampling, thereby enabling faster model training. The ground truth image corresponding to the dual-temporal remote sensing image is subjected to a state transition operation according to the diffusion process, thereby destroying the pixel states in the dual-temporal remote sensing image, and the destroyed ground truth image and its corresponding dual-temporal remote sensing image are input into a deep network model together. In this embodiment, the ground truth image and the full-change sample image are regarded as two extreme states of state transition, and the pixel states of the transition image at the intermediate moment in the diffusion process are randomly selected from these two extreme states using a state transition formula to obtain pixel-destroyed images of the ground truth image to varying degrees.

[0089] The deep network model includes a change detection encoder module, a Fourier transform module, a feature pyramid network FPN, a bidirectional cross attention mechanism module, and a UNet-based backbone network module;

[0090] Through the change detection encoder module containing the twin convolutional neural network, the multi-scale image pixel feature map of the dual-phase high-resolution remote sensing image in the remote sensing sample image is extracted, and the multi-scale image pixel feature map is feature fused to obtain the multi-scale encoding feature map;

[0091] In a specific embodiment, the change detection encoder module includes a first convolution block, a first residual convolution block, a second residual convolution block, a third residual convolution block, and a coding fusion module connected in sequence;

[0092] The first convolution block is used to perform a convolution operation on the remote sensing transition image to obtain a first convolution feature map. Specifically, the first convolution block, Block 1, includes a 7×7 two-dimensional convolution layer, a normalization layer, and a nonlinear activation function layer. The operation methods of each network layer are all well-known technologies and will not be described in detail.

[0093] The first residual convolution block, the second residual convolution block, and the third residual convolution block all include a first two-dimensional convolution layer, a second two-dimensional convolution layer, a first normalization layer, and a first nonlinear activation function layer with different convolution kernels and connected in sequence; the first residual convolution block, Block 2, the second residual convolution block, Block 3, and the third residual convolution block, Block 4, are all composed of residual blocks, and each residual block includes a 3×3 two-dimensional convolution layer, two 1×1 two-dimensional convolution layers, three normalization layers, and three nonlinear activation function layers. The specific encoder network structure of its change detection encoder module is shown in Table 1:

[0094] Table 1. Encoder network structure

[0095]

[0096] The first two-dimensional convolutional layer is used to perform a two-dimensional convolution operation on its input;

[0097] The second two-dimensional convolutional layer is used to perform a two-dimensional convolution operation on the output of the first two-dimensional convolutional layer;

[0098] The first normalization layer is used to normalize the output of the second two-dimensional convolutional layer;

[0099] The first nonlinear activation function layer is used to perform nonlinear activation operation on the output of the normalization layer;

[0100] The first residual convolution block is used to perform residual convolution processing on the first convolution feature map to obtain a second convolution feature map; the second residual convolution block is used to perform residual convolution processing on the second convolution feature map to obtain a third convolution feature map; the third residual convolution block is used to perform residual convolution processing on the third convolution feature map to obtain a fourth convolution feature map; the coding fusion module is used to splice and fuse the first convolution feature map, the second convolution feature map, the third convolution feature map and the fourth convolution feature map to obtain a multi-scale coding feature map;

[0101] The multi-scale coded feature map is Fourier transformed by the Fourier transform module to obtain a multi-scale mixed difference feature map containing the frequency domain and spatial domain of the dual-temporal remote sensing image;

[0102] The specific steps include:

[0103] S211: Based on the fast Fourier transform method, a frequency domain transform operation is performed on the multi-scale coding feature maps corresponding to each of the dual-phase high-resolution remote sensing images to obtain two multi-scale frequency domain pixel feature maps corresponding to different phases; the pixel feature difference in the two multi-scale frequency domain pixel feature maps is used as the multi-scale frequency domain pixel difference feature, thereby obtaining a multi-scale frequency domain pixel difference feature map;

[0104] S212: Using the pixel feature difference between the two multi-scale image pixel feature maps corresponding to the dual-temporal high-resolution remote sensing image as a multi-scale spatial domain difference feature, thereby obtaining a multi-scale spatial domain difference feature map;

[0105] S213: performing feature map splicing on the multi-scale frequency domain pixel difference feature map and the multi-scale spatial domain difference feature map through a preset splicing layer to obtain a spliced ​​feature map; and performing a two-dimensional convolution operation on the spliced ​​feature map through a preset 1×1 two-dimensional convolution module to obtain a multi-scale mixed difference feature map;

[0106] The multi-scale mixed difference feature map is subjected to layer-by-layer feature fusion through a feature pyramid network (FPN) to obtain a high-dimensional difference feature map. The network structure of the feature pyramid network (FPN) is well known in the art and is only used here to process the multi-scale mixed difference feature map in this embodiment through the existing network structure.

[0107] The high-dimensional difference feature map and the feature map output by the fourth block module in the UNet-based backbone network are fused through the bidirectional cross-attention mechanism module, and the feature is enhanced through spatial-channel attention to obtain the cross-attention feature map that integrates the spatial-channel attention features.

[0108] Specifically, the bidirectional cross attention mechanism module includes a first cross attention mechanism module CrossAttention, a second cross attention mechanism module CrossAttention and a space-channel attention module;

[0109] The first cross-attention mechanism module CrossAttention is used to take the high-dimensional difference feature map as the input of the query vector Q1, and the feature map output by the fourth Block module in the UNet-based backbone network as the input of the key vector K1 and the value vector V1 to perform spatial cross-attention fusion and obtain the first attention feature map;

[0110] And the expression of the first cross attention mechanism module CrossAttention is

[0111]

[0112] Where: Q1 represents the query vector, and its input is the high-dimensional difference feature map; K1, V1 represent the key vector and value vector respectively, and their input is the feature map output by the fourth Block module in the UNet-based backbone network; d k1 represents the dimension of the key vector K1; T is the transpose;

[0113] The second cross-attention mechanism module CrossAttention is used to take the feature map output by the fourth Block module in the UNet-based backbone network as the input of the query vector Q2, and the high-dimensional difference feature map as the input of the key vector K2 and the value vector V2 to perform spatial cross-attention fusion and obtain the second attention feature map;

[0114] And the expression of the second cross attention mechanism module CrossAttention is

[0115]

[0116] Where: Q2 represents the query vector, whose input is the feature map output by the fourth Block module in the UNet-based backbone network; K2, V2 represent the key vector and value vector respectively, whose input is the high-dimensional difference feature map; d k2 represents the dimension of the key vector K2;

[0117] The spatial-channel attention module is used to perform channel splicing on the first attention feature map and the second attention feature map, and combine channel attention to emphasize the key feature channel and spatial attention to highlight the spatially significant area, i.e., feature enhancement, to obtain a cross-attention feature map that integrates the spatial-channel attention features; the specific implementation method of feature enhancement is a known technical means; specifically, this embodiment fuses the high-dimensional difference feature map with the high-dimensional feature information in the diffusion model, i.e., the feature map output by the fourth Block module in the backbone network based on UNet, through the cross-attention mechanism shared by the position encoding and normalization layer, and uses channel attention to automatically focus on important feature dimensions and spatial attention to enhance the response of key areas;

[0118] The pixel features of the cross-attention feature map are extracted through the backbone network module based on UNet to obtain the pixel change probability feature map; specifically, the backbone network module based on UNet includes a first Block module, a second Block module, a third Block module and a fourth Block module connected in sequence; each Block module has two residual blocks, and each residual block has two 3×3 two-dimensional convolution layers, two normalization layers, two nonlinear activation function layers and a fully connected layer; Block1, Block 2 and Block 3 all include a 3×3 two-dimensional convolution downsampling layer; Block4 includes a multi-head attention block. Table 2 shows the specific network structure, and the network structure model is as follows Figure 2 As shown;

[0119] Table 2. Network structure of the backbone network based on UNet

[0120]

[0121] The first Block module, the second Block module, and the third Block module all include a first residual inner convolution block and a convolution downsampling layer with different convolution kernels that are connected in sequence;

[0122] The first residual inner convolution block includes a third two-dimensional convolution, a second normalization layer, a second nonlinear activation function layer, and a fully connected layer connected in sequence;

[0123] The third two-dimensional convolution is used to perform a two-dimensional convolution operation on its input;

[0124] The second normalization layer is used to normalize the output of the third two-dimensional convolution;

[0125] The second nonlinear activation function layer is used to perform a nonlinear activation operation on the output of the second normalization layer;

[0126] The fully connected layer is used to perform a fully connected operation on the output of the second nonlinear activation function layer;

[0127] The convolutional downsampling layer is used to downsample the output of the fully connected layer;

[0128] The first Block module is used to extract pixel features of the multi-scale coding feature map to obtain the fifth convolution feature map; the second Block module is used to extract pixel features from the fifth convolution feature map to obtain the sixth convolution feature map; the third Block module is used to extract pixel features from the sixth convolution feature map to obtain the seventh convolution feature map;

[0129] The fourth Block module includes a second residual inner convolution block and a multi-head attention block that are different from the convolution kernels of other Block modules;

[0130] The second residual inner convolution block is used to perform a convolution operation on the seventh convolution feature map to obtain an eighth convolution feature map; the multi-head attention block is used to perform a multi-head attention operation on the eighth convolution feature map to extract image pixel features to obtain a pixel change probability feature map;

[0131] The pixel features in the pixel change probability feature map are back-diffused by the back-diffusion model to obtain the predicted change probability map of the dual-temporal remote sensing image. Specifically, the back-diffusion process is established, and its expression is:

[0132]

[0133] Where: I a ,I b represents the corresponding dual-time remote sensing image; t represents the iterative time step of the diffusion model; x t Represents the transition image at the middle moment of the diffusion process; The predicted binary fine change mask image, i.e., the predicted pixel change probability feature map, is obtained through the classifier module; Has representation The confidence score and the predicted probability of the change map have dual meanings; f θ represents a deep network model with parameters θ;

[0134] Then, according to the reverse diffusion process, the expression for the predicted change probability map of the dual-temporal remote sensing image is obtained as follows:

[0135]

[0136] Where: Represents the reverse state transition probability matrix at the tth iteration; A map showing the predicted changes of dual-temporal remote sensing images; represents the confidence probability that the image pixel (i, j) predicted by the deep network model with weight parameter θ belongs to a certain pixel state at time step t; Indicates the conditional transition probability of the image pixel (i, j) from the state at time step t to the state at time step t-1 under the deep network model with weight parameter θ; through the reverse state transition probability matrix, x t and As the input of the back-diffusion model, the conversion sampling can convert a part of the pixels to a precise and certain pixel state at each time step, thereby obtaining a predicted change probability map;

[0137] The predicted change probability image is binarized by the classifier module to obtain a black and white binary change feature map, thereby obtaining a change detection result of the dual-temporal remote sensing image and then obtaining a trained remote sensing image change detection model; wherein the method of binarizing the predicted change probability image is a well-known technical means and will not be elaborated on here;

[0138] S22: Model validation: constructing the total model loss function

[0139] And the construction formula of the total model loss function L is

[0140] L=L bce +L edge

[0141]

[0142] Where: L bce represents the cross entropy loss function; x i The predicted change probability map is represented by The determined probability map; y i represents the ground truth map corresponding to the dual-temporal high-resolution remote sensing image; N represents the total number of samples; L edge represents the edge structure consistency loss function; Respectively represent the Sobel edge gradient amplitude of the predicted change map and its corresponding ground truth map at a certain pixel point; ΔI i They represent the Laplacian response value of the predicted change map and its corresponding ground truth map at a certain pixel point;

[0143] Based on the constructed total model loss function, the trained remote sensing image change detection model is validated using the validation set to determine whether the output of the trained remote sensing image change detection model converges; if so, the trained remote sensing image change detection model is confirmed to be the optimal remote sensing image change detection model; otherwise, the weight parameters of the trained remote sensing image change detection model are adaptively adjusted based on the back-propagation method, and step S21 is repeated;

[0144] S3: Implement change detection on the dual-temporal high-resolution remote sensing images in the test set according to the optimal remote sensing image change detection model.

[0145] Example experiment:

[0146] The method described in this embodiment performs a change detection experiment on a high-resolution remote sensing image using the CLCD dataset. The results of the experiment are shown in Table 3:

[0147] Table 3. CLCD detection accuracy (%)

[0148] Classification accuracy Precision 71.51 Recall 59.85 F1 65.16 IoU 48.33 OA 95.24

[0149] Among them: Precision represents the actual correct proportion of pixels predicted by the model as "changed", Recall represents the proportion of pixels successfully detected by the model among the real "changed" pixels, F1 represents the weighted average of Precision and Recall, reflecting the overall accuracy of the model in detecting the "change" category, IoU (Intersection over Union) represents the proportion of the intersection of the predicted change area and the real change area to their union, OA (Overall Accuracy) represents the proportion of pixels predicted correctly by the model (including changed and unchanged), such as Figure 3 The following shows a specific example of the test dataset;

[0150] In order to more objectively evaluate the role of each step in the remote sensing image change detection method based on a state-driven discrete diffusion model proposed in this embodiment, an existing ablation experiment is added for illustration. On the basis of the common prototype network, a single module or a combination of different modules is added to compare the experimental results. The specific experimental results are shown in Table 4:

[0151] Table 4. Classification accuracy of different modules (%)

[0152]

[0153] The following conclusions can be drawn from the above experiments:

[0154] (1) The experimental results in Table 3 show that the method described in this embodiment has a good detection effect in remote sensing image change detection, which proves the superior performance of the method described in this embodiment in high-resolution remote sensing change detection.

[0155] (2) The ablation experiment data in Table 4 show that the detection result of adding interactive operations on feature information in the frequency domain and spatial domain is better than the result of only operating in the spatial domain, which proves that the use of frequency domain information in the method described in this embodiment is conducive to improving the detection effect.

[0156] (3) The ablation experiment data in Table 4 show that the detection result of adding the hierarchical fusion operation of multi-scale feature information is better than the result of a single high-dimensional feature, which proves that the full use of multi-scale feature information in the method described in this embodiment is conducive to improving the detection effect.

[0157] (4) The ablation experiment data in Table 4 show that the detection results of using the bidirectional cross-attention mechanism to embed difference feature information into the diffusion model are better than the results of direct fusion, which proves that the use of the bidirectional attention mechanism in the method described in this embodiment is conducive to improving the detection effect.

[0158] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A remote sensing image change detection method based on a state-driven discrete diffusion model, characterized in that: The specific steps include: S1: Acquire dual-temporal high-resolution remote sensing images for change detection; The dual-temporal high-resolution remote sensing image dataset is randomly divided into training set, validation set and test set; The method for obtaining the remote sensing sample images input to the model in the training set and the validation set is as follows: obtaining the ground truth map of the dual-temporal high-resolution remote sensing image, and dimensionally splicing the dual-temporal high-resolution remote sensing image with the ground truth map to obtain the remote sensing sample images input to the model; S2: A remote sensing image change detection model based on a state-driven discrete diffusion model is constructed using a training set and a validation set to train the model and obtain the optimal remote sensing image change detection model; the remote sensing image change detection model includes a forward diffusion model based on state transition, a deep network model, a backward diffusion model, and a classifier module; S21: Model training: performing forward diffusion processing on the ground truth image in the remote sensing sample image using a pixel-level state transition strategy constructed by the forward diffusion model to obtain a remote sensing transition image; The deep network model includes a change detection encoder module, a Fourier transform module, a feature pyramid network FPN, a bidirectional cross attention mechanism module, and a UNet-based backbone network module; The change detection encoder module is used to extract the multi-scale image pixel feature map of the dual-phase high-resolution remote sensing image in the remote sensing sample image, and the multi-scale image pixel feature map is subjected to feature fusion to obtain the multi-scale encoding feature map; The multi-scale coded feature map is Fourier transformed by the Fourier transform module to obtain a multi-scale mixed difference feature map containing the frequency domain and spatial domain of the dual-temporal remote sensing image; The feature pyramid network FPN is used to fuse the multi-scale mixed difference feature map to obtain a high-dimensional difference feature map; The high-dimensional difference feature map is fused with the feature map output by the fourth block module in the UNet-based backbone network module through the bidirectional cross-attention mechanism module, and the feature is enhanced through spatial-channel attention to obtain the cross-attention feature map that integrates the spatial-channel attention features. The pixel features of the cross-attention feature map are extracted through the UNet-based backbone network module to obtain the pixel change probability feature map; The pixel features in the pixel change probability feature map obtained by the deep network model are back-diffused through the back-diffusion model to iteratively obtain the predicted change probability map of the dual-temporal remote sensing image. The predicted change probability image is binarized through the classifier module to obtain a black and white binary change feature map to obtain the change detection results of the dual-temporal remote sensing image, and then the trained remote sensing image change detection model is obtained. S22: Model verification: Based on the constructed total model loss function, the trained remote sensing image change detection model is verified through the validation set to determine whether the output of the trained remote sensing image change detection model converges; If yes, then confirm that the trained remote sensing image change detection model is the optimal remote sensing image change detection model; otherwise, adaptively adjust the weight parameters of the trained remote sensing image change detection model based on the back propagation method, and repeat step S21; S3: Implement change detection on the dual-temporal high-resolution remote sensing images in the test set according to the optimal remote sensing image change detection model.

2. The remote sensing image change detection method based on a state-driven discrete diffusion model according to claim 1, characterized in that: The expression of the pixel-level state transition strategy constructed by the forward diffusion model in S21 is: Where: Indicates that image pixel i, j changes from the initial state After t steps of forward diffusion, the state The conditional transition probability of It represents the image pixel state of the dual-temporal high-resolution remote sensing image at the t-th iteration, that is, the remote sensing transition image at any time is determined based on the image pixel state; Represents the image pixel status of the dual-temporal high-resolution remote sensing image at the initial iteration; represents the set of state transition matrices and Q t represents the state transition matrix at the tth iteration and represents the pixel state transition probability used for pixel state transition and β t represents the pixel state transition probability value at the tth iteration; x t represents the remote sensing transition image corresponding to the tth iteration; i, j represent the image pixel positions of the dual-temporal high-resolution remote sensing image.

3. The remote sensing image change detection method based on a state-driven discrete diffusion model according to claim 2, characterized in that: The change detection encoder module includes a first convolution block, a first residual convolution block, a second residual convolution block, a third residual convolution block and a coding fusion module connected in sequence; The first convolution block is used to perform a convolution operation on the dual-temporal high-resolution remote sensing image in the remote sensing sample image to obtain a first convolution feature map; The first residual convolution block, the second residual convolution block, and the third residual convolution block each include a first two-dimensional convolution layer, a second two-dimensional convolution layer, a first normalization layer, and a first nonlinear activation function layer with different convolution kernels and connected in sequence; The first two-dimensional convolutional layer is used to perform a two-dimensional convolution operation on its input; The second two-dimensional convolutional layer is used to perform a two-dimensional convolution operation on the output of the first two-dimensional convolutional layer; The first normalization layer is used to normalize the output of the second two-dimensional convolutional layer; The first nonlinear activation function layer is used to perform nonlinear activation operation on the output of the normalization layer; The first residual convolution block is used to perform residual convolution processing on the first convolution feature map to obtain a second convolution feature map; the second residual convolution block is used to perform residual convolution processing on the second convolution feature map to obtain a third convolution feature map; the third residual convolution block is used to perform residual convolution processing on the third convolution feature map to obtain a fourth convolution feature map; The coding fusion module is used to splice and fuse the first convolution feature map, the second convolution feature map, the third convolution feature map and the fourth convolution feature map to obtain a multi-scale coding feature map.

4. The remote sensing image change detection method based on a state-driven discrete diffusion model according to claim 3, characterized in that: The UNet-based backbone network module includes a first Block module, a second Block module, a third Block module, and a fourth Block module connected in sequence; The first Block module, the second Block module, and the third Block module all include a first residual inner convolution block and a convolution downsampling layer with different convolution kernels that are connected in sequence; The first residual inner convolution block includes a third two-dimensional convolution, a second normalization layer, a second nonlinear activation function layer, and a fully connected layer connected in sequence; The third two-dimensional convolution is used to perform a two-dimensional convolution operation on its input; The second normalization layer is used to normalize the output of the third two-dimensional convolution; The second nonlinear activation function layer is used to perform a nonlinear activation operation on the output of the second normalization layer; The fully connected layer is used to perform a fully connected operation on the output of the second nonlinear activation function layer; The convolutional downsampling layer is used to downsample the output of the fully connected layer; The first Block module is used to extract pixel features of the multi-scale coding feature map to obtain the fifth convolution feature map; the second Block module is used to extract pixel features from the fifth convolution feature map to obtain the sixth convolution feature map; the third Block module is used to extract pixel features from the sixth convolution feature map to obtain the seventh convolution feature map; The fourth Block module includes a second residual inner convolution block and a multi-head attention block that are different from the convolution kernels of other Block modules; The second residual inner convolution block is used to perform a convolution operation on the seventh convolution feature map to obtain an eighth convolution feature map; the multi-head attention block is used to perform a multi-head attention operation on the eighth convolution feature map to extract image pixel features to obtain a pixel change probability feature map.

5. The remote sensing image change detection method based on a state-driven discrete diffusion model according to claim 4, characterized in that: The method for obtaining a multi-scale mixed difference feature map containing frequency domain and spatial domain of a dual-temporal remote sensing image in S21 specifically includes: S211: Based on the fast Fourier transform method, a frequency domain transform operation is performed on the multi-scale coding feature maps corresponding to each of the dual-phase high-resolution remote sensing images to obtain two multi-scale frequency domain pixel feature maps corresponding to different phases; The pixel feature difference between the two multi-scale frequency domain pixel feature maps is used as the multi-scale frequency domain pixel difference feature, thereby obtaining a multi-scale frequency domain pixel difference feature map; S212: Using the pixel feature difference between the two multi-scale image pixel feature maps corresponding to the dual-temporal high-resolution remote sensing image as a multi-scale spatial domain difference feature, thereby obtaining a multi-scale spatial domain difference feature map; S213: performing feature map splicing on the multi-scale frequency domain pixel difference feature map and the multi-scale spatial domain difference feature map through a preset splicing layer to obtain a spliced ​​feature map; And through the preset two-dimensional convolution module, a two-dimensional convolution operation is performed on the spliced ​​feature map to obtain a multi-scale mixed difference feature map.

6. The remote sensing image change detection method based on a state-driven discrete diffusion model according to claim 5, characterized in that: In S21, the pixel features in the pixel change probability feature map are back-diffused by the back-diffusion model to obtain the expression of the predicted change probability map of the dual-temporal remote sensing image: The expression for obtaining the predicted change probability map of the dual-temporal remote sensing image is: Where: Represents the reverse state transition probability matrix at the tth iteration; A map showing the predicted changes of dual-temporal remote sensing images; represents the confidence probability that the image pixel i, j belongs to a certain pixel state at time step t, as predicted by the deep network model with weight parameter θ; represents the conditional transition probability of image pixel i, j from the state at time step t to the state at time step t-1 under the deep network model with weight parameter θ; Represents the remote sensing transition image x corresponding to the tth iteration t The pixel status of each pixel in .

7. The remote sensing image change detection method based on a state-driven discrete diffusion model according to claim 6, characterized in that: The formula for constructing the total model loss function L in S22 is: L=L bce +L edge Where: L bce represents the cross entropy loss function; x i The predicted change probability map is represented by The determined probability map; y i represents the ground truth map corresponding to the dual-temporal high-resolution remote sensing image; N represents the total number of samples; L edge represents the edge structure consistency loss function; Respectively represent the Sobel edge gradient amplitude of the predicted change map and its corresponding ground truth map at a certain pixel point; ΔI i They respectively represent the Laplacian response value of the predicted change map and its corresponding ground truth map at a certain pixel point.

8. The remote sensing image change detection method based on a state-driven discrete diffusion model according to claim 1, characterized in that: The bidirectional cross attention mechanism module includes a first cross attention mechanism module CrossAttention, a second cross attention mechanism module CrossAttention and a space-channel attention module; The first cross-attention mechanism module CrossAttention is used to take the high-dimensional difference feature map as the input of the query vector Q1, and the feature map output by the fourth Block module in the UNet-based backbone network as the input of the key vector K1 and the value vector V1 to perform cross-attention fusion and obtain the first attention feature map; And the expression of the first cross attention mechanism module CrossAttention is Where: Q1 represents the query vector, and its input is the high-dimensional difference feature map; K1, V1 represent the key vector and value vector respectively, and their input is the feature map output by the fourth Block module in the UNet-based backbone network; d k1 represents the dimension of the key vector K1; is transposed; The second cross-attention mechanism module CrossAttention is used to take the feature map output by the fourth Block module in the UNet-based backbone network as the input of the query vector Q2, and the high-dimensional difference feature map as the input of the key vector K2 and the value vector V2 to perform spatial cross-attention fusion and obtain the second attention feature map; And the expression of the second cross attention mechanism module CrossAttention is Where: Q2 represents the query vector, and its input is the feature map output by the fourth Block module in the UNet-based backbone network; K2, V2 represent the key vector and value vector respectively, and their input is the high-dimensional difference feature map; d k2 represents the dimension of the key vector K2; The spatial-channel attention module is used to perform channel splicing on the first attention feature map and the second attention feature map, and perform feature enhancement on the feature map after channel splicing to obtain a cross-attention feature map that integrates the spatial-channel attention features.