A remote sensing image change detection method based on sparse change self-attention mechanism

By using a probabilistic change model and self-attention module based on a sparse change self-attention mechanism, the problems of high computational complexity and insufficient applicability in change detection of high-resolution remote sensing images are solved, achieving efficient and unified change detection results.

CN115761476BActive Publication Date: 2026-01-06WUHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211173201.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-26
Publication Date
2026-01-06
Estimated Expiration
2042-09-26

AI Technical Summary

Technical Problem

Existing deep learning methods suffer from high computational complexity and energy consumption in high-resolution remote sensing image change detection. Furthermore, the manually designed network structures lack applicability and are deficient in theoretical support and the integration of domain knowledge.

Method used

Based on the sparse change self-attention mechanism, a sparse change self-attention module is designed by constructing a probabilistic change model for conditional decomposition and parameterizing the deep change detection network architecture, thereby achieving computationally efficient and task-adaptive change detection.

Benefits of technology

It achieves improved computational efficiency and detection performance for unified change detection in high-resolution remote sensing imagery, reduces computational resource requirements, and is applicable to a variety of change detection tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115761476B_ABST
    Figure CN115761476B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of remote sensing image change detection method based on sparse change self-attention mechanism, for computing, memory efficient high-precision unified change detection.The present application combines the principle of deep learning, probability graph theory, proposes a unified probability change modeling theory framework, the joint distribution of random variable in change process is conditionally decomposed, according to different prior assumptions, different factors can be decomposed, these factors are the theoretical representation of the architecture of deep change detection model, further parameterize these decomposition factors using the proposed sparse change self-attention module, so as to obtain specific task adaptive, computationally efficient deep change detection model architecture.The present application can solve the problem that existing architecture design lacks theoretical basis and has high computational complexity, and can realize unified processing of various change detection tasks and rapid change detection of large-scale remote sensing image pairs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of high-resolution remote sensing image recognition, and specifically relates to a method for detecting changes in remote sensing images based on a sparse change self-attention mechanism. Background Technology

[0002] With the launch of high-resolution remote sensing satellites both domestically and internationally, and the continuous development of sensor technology, people are now able to acquire a large amount of high-resolution remote sensing imagery. Compared to medium- and low-resolution remote sensing imagery, high-resolution remote sensing imagery possesses rich spatial details, continuous spectrum, and dynamic change information. The geometric structure of ground features is more obvious, their location layout is clearer, and information such as texture and size is more refined. This allows for a clear representation of the characteristic distribution and spatiotemporal relationships of ground features, making the dynamic and accurate interpretation of ground features possible.

[0003] Remote sensing change detection, as a fundamental Earth vision technology, is widely used in practical tasks such as urban development monitoring, disaster damage assessment, and flood monitoring, greatly promoting the progress of sustainable development goals. Remote sensing change detection aims to find precise areas of change using dual-temporal remote sensing image pairs. Based on the change of interest, it can be divided into two types of change detection and semantic change detection (including one-to-many and many-to-many change modes). Two-type change detection focuses only on the occurrence of change, while semantic change detection also considers the direction of change. Traditional change detection paradigms driven by models based on handcrafted features combined with classifiers are highly dependent on the adaptability of handcrafted feature operators to the current scene. For example, scale-invariant feature transform operators have good descriptive ability for multi-scale changes but are difficult to fully describe areas with radiometric differences. In complex Earth observation scenarios, local geometric offsets between image pairs, global-local radiometric differences, and multi-scale problems of changed areas often coexist, making traditional handcrafted feature operators difficult to handle.

[0004] Deep learning, as an automatic hierarchical feature representation learning framework, has revolutionized the field of remote sensing change detection with its data-driven change detection paradigm. It learns robust spatiotemporal feature representations through massive and diverse change training samples, overcoming the three challenges mentioned above. Convolutional neural networks and their variants combined with attention mechanisms, such as Siamese fully convolutional neural networks and Siamese self-attention transform networks, compared to traditional methods, can adaptively model long-distance change context information to alleviate the local geometric offset problem, extract radiation-independent discriminative features to overcome radiation difference problems, and use pyramid hierarchical features to overcome scale differences in change regions. Therefore, they can achieve robust change detection with low false alarms and high recall, significantly improving the intelligence of change detection systems and reducing the cost of manually delineating change patches.

[0005] Although deep learning methods based on self-attention mechanisms have greatly improved the performance of change detection in high-resolution images, their high computational and memory costs (complexity quadratic with image size) make large-scale city-level and national-level surface monitoring applications extremely difficult. Furthermore, manually designed change detection network structures are typically only suitable for one type of change detection task, namely binary or semantic change detection. This requires a separate, targeted network design, placing extremely high demands on designers' domain knowledge, analytical skills, and model design experience.

[0006] The fundamental problem lies in the lack of theoretical support for the deep change detection architecture design and the lack of integration of domain knowledge into the general self-attention operator.

[0007] To address the aforementioned issues, a unified theoretical framework for modeling various change detection tasks is urgently needed. This framework would allow for mathematical derivation to provide design principles for a deep change detection architecture. Furthermore, by incorporating domain knowledge of remote sensing change, namely change sparsity and local correlation, a domain-specific computationally and memory-efficient self-attention operator would be designed to parameterize the aforementioned theoretical deep change detection architecture. This would overcome the current challenges of "lack of theory" and "high energy consumption" in the field of change detection. Summary of the Invention

[0008] The purpose of this invention is to propose a unified change detection method for high-resolution remote sensing images based on a sparse change self-attention mechanism.

[0009] The proposed unified change detection method for high-resolution remote sensing images based on a sparse change self-attention mechanism first models the probabilistic change process and performs conditional decomposition on the joint distribution of random variables during the change process. According to different prior assumptions, different factors can be obtained, which are the theoretical representation of the depth change detection model architecture. Furthermore, the proposed sparse change self-attention module is used to parameterize these decomposition factors, thereby obtaining a specific task-adaptive and computationally efficient depth change detection model architecture, thus overcoming the problems of lack of theoretical basis and high computational complexity in architecture design.

[0010] This invention provides a unified change detection method for high-resolution remote sensing images based on a sparse change self-attention mechanism, the implementation steps of which are as follows:

[0011] Step 1: Construct a unified probabilistic change model framework. Based on the assumptions of different change detection tasks, perform conditional decomposition of the joint distribution of variables within the framework to obtain the sub-task factors.

[0012] Step 2: Construct a sparse variation self-attention module and use it to parameterize the subtask factors obtained in Step 1 to obtain the specific deep variation detection network architecture.

[0013] Step 3: Construct a high-resolution remote sensing image change detection dataset, split it into a training set and a test set, and perform geometric, radiometric and other data augmentation operations on the training set and normalize it.

[0014] Step 4: Construct a unified optimization objective function for change detection. Use the training set of the change detection dataset obtained in Step 3 to train the deep change detection network architecture obtained in Step 2 until the convergence condition is met.

[0015] Step 5: Using the depth change detection network architecture and corresponding parameters obtained in Step 4, predict the test set, and solve the change probability map according to the corresponding probability change model to obtain the change region result.

[0016] Furthermore, step 1 specifically includes the following sub-steps:

[0017] Step 1.1: Establish a probabilistic change model, considering the change as state S at time t. t State S at time t+1 t+1 The transition process, and the value of each state and its corresponding image I. t Directly related, therefore the joint distribution of the process from time t to t+1 is P(S t ,S t+1 ,I t ,I t+1 Based on the causal relationship between the state and the image, the joint distribution can be decomposed into several factors:

[0018] P(S t ,S t+1 ,I t ,I t+1 )=P(I t )P(I t+1 |I t )P(S t |I t )P(S t+1 |S t ,I t+1 )

[0019] Where P(I) t ),P(I t+1 |I t ) respectively represent image I t with I t+1 The data distribution, which is based on image I t Given a conditional data distribution P(I) t+1 |I t ) indicates image I t+1 In the context of image I t The data was obtained from sampling the covered geographic locations; P(S) t |It ),P(S t+1 |S t ,I t+1 ) represent the dependency relationship between the state distribution at two time points and their corresponding images, such as P(S) t+1 |S t ,I t+1 ) represents state S t+1 Simultaneously dependent on state S t With Image I t+1 .

[0020] Step 1.2: Based on the probabilistic change model, derive the theoretical framework for binary change detection. Since binary change detection is state-independent change detection, focusing only on the occurrence of change, it corresponds to state transitions in the probabilistic change model. Based on this objective assumption, the state-related factor P(S) in the probabilistic change model can be analyzed. t |I t ),P(S t+1 |S t ,I t+1 ) are merged to obtain P(S) t+1 ,S t |I t+1 ,I t Since the second type of change is independent of the specific value of the state, the factor P(S) t+1 ,S t |I t+1 ,I t ) using P(C|I t+1 ,I t Let C be a random variable representing a type II change; the factor P(C|I) is re-expressed. t+1 ,I t This is the theoretical framework of the two-type change detection model.

[0021] Step 1.3: Based on the probabilistic change model, derive the theoretical framework for semantic change detection. For semantic change detection, it is necessary to consider the specific values ​​of the state, utilizing the state-change condition independence assumption, i.e., S... t ⊥S t+1 |C, can be applied to factor P(S) t |I t )P(S t+1 |S t ,I t+1 Merge and then decompose:

[0022] P(S t |I t )P(S t+1 |S t ,I t+1 )=P(S t+1 ,S t |It+1 ,I t )=P(S t+1 |C,I t+1 )P(S t |C,I t )P(C|I t+1 ,I t )

[0023] Among them, the subtask factor P(S) t+1 |C,I t+1 ) and P(S t |C,I t Both ) represent semantic segmentation tasks conditioned by two types of changes, and the subtask factor P(C|I) represents a semantic segmentation task. t+1 ,I t ) represents a binary change detection task; mathematically, this decomposition breaks down the semantic change detection task into two semantic segmentation tasks conditioned on change and one binary change detection task.

[0024] Furthermore, step 2 specifically includes the following sub-steps:

[0025] Step 2.1: Construct the sparse variation self-attention module; this module consists of a layer normalization layer, a sparse variation multi-head self-attention layer, a residual layer normalization layer, and a convolutional multi-layer perceptron layer stacked in this order. Given an input feature map x and a sparse variation prior probability map p... c (Applying a fully convolutional network to the features output by a regular self-attention module for sparse variation prior probability estimation), firstly, three 1x1 convolutional layers are used to transform the input feature map x into feature maps Q, K, and V respectively. Layer normalization is then performed on these three feature maps. Next, the probability map p... c Three feature maps, Q, K, and V, are input into a sparsely varying multi-head self-attention layer to obtain context-enhanced feature maps of the varying regions. These context-enhanced feature maps are then normalized and added to the input feature maps, forming a residual normalization layer. The residual normalized features are further input into a convolutional multilayer perceptron layer for local position encoding and nonlinear transformation, and then added to the input features to form a residual connection, yielding the output feature map. In the sparsely varying multi-head self-attention layer, a probability map p is utilized. c Threshold segmentation is performed to obtain a coarse segmentation map of the changed regions. Based on this segmentation map, the features Q, K, and V are indexed spatially to obtain the feature Q of the changed regions. c ,K c V c Then perform the standard self-attention calculation.

[0026] Step 2.2: Parameterize subtask factors using sparse variation self-attention modules. An existing visual backbone network (such as ResNet, Swin-Transformer) is used as the image encoder to obtain a temporally independent hierarchical dense feature set from dual-temporal images. Further, one regular self-attention module and three sparse variation self-attention modules are constructed at the top layer of the image encoder for variation feature representation. All parameters of the image encoder and self-attention modules are used to parameterize the subtask factors. Specifically, the parameterized image encoder is used to extract the hierarchical dense feature set {x}. (i)}, i = 2, 3, 4, 5, then the feature x (5) The input is fed into a regular self-attention module to enhance the long-range spatial context information of this feature, resulting in the context feature x. c This is used for subsequent prior estimation of sparse changes, yielding the prior probability p of the sparse changes. c ; the prior probability p of sparse variation c With hierarchical dense feature set {x (4) x (3) x (2) The input is fed into three sparse variation self-attention modules to further extract spatiotemporal context information, and the output variables to be solved in the subtask factors are then completed, thus achieving parameterization. The calculation of the sparse variation self-attention part adopts a progressive aggregation method, that is, first, the low-resolution feature x is aggregated... (4) With contextual features x c Summation, along with the prior probability p of sparse variation. c The input is fed into the first sparse variation self-attention module to obtain the first spatiotemporal context feature; then the first spatiotemporal context feature is combined with the medium-resolution feature x. (3) Summation, along with the prior probability p of sparse variation. c The input is fed into the second sparse variation self-attention module to obtain the second spatiotemporal context feature; finally, the second spatiotemporal context feature is combined with the high-resolution feature x. (2) Summation, along with the prior probability p of sparse variation. c The input is fed into the third sparse variation self-attention module to obtain the third spatiotemporal context feature.

[0027] Furthermore, step 3 specifically includes the following sub-steps:

[0028] Step 3.1: Use drones or high-resolution satellites to acquire a large number of paired high spatial resolution remote sensing images;

[0029] Step 3.2: Pair-label N types of ground interest target samples in remote sensing images, compare the target sample masks of two time phases to obtain change masks, and make a deep learning change detection dataset by combining the sample mask, change mask and the corresponding regional and temporal images.

[0030] Step 3.3: Divide the deep learning change detection dataset into two parts in a 4:1 ratio: training set A for model training and test set B for evaluating model performance.

[0031] Step 3.4: Normalize A and B, and apply geometric and radiometric data augmentation methods such as horizontal flipping, vertical flipping, random rotation, and random color dithering only to A.

[0032] Furthermore, step 4 specifically includes the following sub-steps:

[0033] Step 4.1, construct a unified optimization objective function for change detection, specifically in the following form:

[0034]

[0035] Where L region The loss function is estimated based on the sparse prior information (which is the sum of the binary cross-entropy loss function and the Dice loss function). Let be the semantic segmentation loss function for phase t1 (and be the cross-entropy loss function). Let L be the semantic segmentation loss function for phase t2 (where L is the cross-entropy loss function). change λ1 and λ2 are the semantic modulation coefficients of the corresponding loss functions, respectively, with a default value of 1.

[0036] Step 4.2: Train the depth change detection network obtained in Step 2 using training set A. First, initialize the network parameters, input data in batches, perform forward computation in the network, obtain the loss value by uniformly optimizing the objective function through change detection, use the loss value to perform backpropagation gradient calculation, use the neural network optimizer to update the parameters, and iterate continuously until the model converges.

[0037] The unified change detection method for high-resolution remote sensing images based on a sparse change self-attention mechanism proposed in this invention has the following two significant features. First, it invents a design principle for a depth change detection network architecture based on probabilistic change processes. It utilizes a probabilistic graphical model to uniformly model arbitrarily complex change processes, and further decomposes the probabilistic graph conditionally to obtain sub-task factors, i.e., the theoretical architecture of the depth change detection network. Second, it invents a sparse change self-attention mechanism that combines prior knowledge of remote sensing change, such as change sparsity and local correlation. Based on this mechanism, a sparse change self-attention module is constructed, parameterizing the sub-task factors of the aforementioned theoretical architecture to form a specific, task-adaptive, and computationally efficient depth change detection model architecture. Attached Figure Description

[0038] Figure 1This is a schematic diagram of the probability change model designed in step 1 of embodiment 1 of the present invention.

[0039] Figure 2 This is a schematic diagram of the sparse variation self-attention module designed in step 2.1 of Embodiment 1 of the present invention.

[0040] Figure 3 This is a schematic diagram of the network structure used for parameterization in step 2.2 of Embodiment 1 of the present invention.

[0041] Figure 4 This is a sample image from the change detection dataset in step 3.2 of Embodiment 1 of the present invention.

[0042] Figure 5 The image shows the change detection result obtained in step 5 of embodiment 1 of the present invention. The left image is the semantic segmentation image of the change region at the first moment, and the right image is the semantic segmentation image of the change region at the second moment. Detailed Implementation

[0043] The following examples further illustrate the outstanding features and significant advancements of the present invention, which are intended to illustrate the invention but are in no way limiting it.

[0044] Example 1

[0045] (a) Constructing a unified probabilistic change modeling framework, such as Figure 1 Based on the semantic-change conditional independence assumption of the semantic change detection task, the joint distribution of variables is conditionally decomposed within the framework to obtain the subtask factors.

[0046] 1.1 Establish a probabilistic change model, considering the change as state S at time t. t State S at time t+1 t+1 The transition process, and the value of each state and its corresponding image I. t Directly related, therefore the joint distribution of the process from time t to t+1 is P(S t ,S t+1 ,I t ,I t+1 Based on the causal relationship between the state and the image, the joint distribution can be decomposed into several factors:

[0047] P(S t ,S t+1 ,I t ,I t+1 )=P(I t )P(I t+1 |I t )P(S t |I t )P(S t+1 |S t ,It+1 )

[0048] Where P(I) t ),P(I t+1 |I t ) respectively represent image I t with I t+1 The data distribution, which is based on image I t Given a conditional data distribution P(I) t+1 |I t ) indicates image I t+1 In the context of image I t The data was obtained from sampling the covered geographic locations; P(S) t |I t ),P(S t+1 |S t ,I t+1 ) represent the dependence of the state distribution at two time points on their corresponding images, where P(S) t+1 |S t ,I t+1 ) represents state S t+1 Simultaneously dependent on state S t With Image I t+1 ;

[0049] 1.2. Based on the probabilistic change model, a theoretical framework for binary change detection is derived. Since binary change detection is state-independent, focusing only on the occurrence of change, it corresponds to state transitions in the probabilistic change model. Based on this objective assumption, the state-related factor P(S) in the probabilistic change model can be analyzed. t |I t ),P(S t+1 |S t ,I t+1 ) are merged to obtain P(S) t+1 ,S t |I t+1 ,I t Since the second type of change is independent of the specific value of the state, the factor P(S) t+1 ,S t |I t+1 ,I t ) using P(C|I t+1 ,I t Let C be a random variable representing a type II change; the factor P(C|I) is re-expressed. t+1 ,I t This constitutes the theoretical framework of the two-type change detection model;

[0050] Since this embodiment is for a semantic change detection task, this step can be omitted;

[0051] 1.3. Based on the probabilistic change model, a theoretical framework for semantic change detection is derived. For semantic change detection, the specific values ​​of the state need to be considered, utilizing the state-change condition independence assumption, i.e., S... t ⊥S t+1 |C, can be applied to factor P(S) t |I t )P(S t+1 |S t ,I t+1 Merge and then decompose:

[0052] P(S t |I t )P(S t+1 |S t ,I t+1 )=P(S t+1 ,S t |I t+1 ,I t )=P(S t+1 |C,I t+1 )P(S t |C,I t )P(C|I t+1 ,I t )

[0053] Among them, the subtask factor P(S) t+1 |C,I t+1 ) and P(S t |C,I t Both ) represent semantic segmentation tasks conditioned by two types of changes, and the subtask factor P(C|I) represents a semantic segmentation task. t+1 ,I t ) represents a binary change detection task; mathematically, this decomposition breaks down the semantic change detection task into two semantic segmentation tasks conditioned on change and one binary change detection task.

[0054] (ii) Constructing sparse variable self-attention modules, such as Figure 2 And by using it to parameterize the above sub-task factors, the specific depth change detection network architecture can be obtained;

[0055] 2.1. A sparse variation self-attention module is constructed using the PyTorch programming framework. This module consists of a layer normalization layer, a sparse variation multi-head self-attention layer, a residual normalization layer, and a convolutional multilayer perceptron layer stacked in this order.

[0056] Given an input feature map x and a sparse variation prior probability map p c First, three 1x1 convolutional layers are used to transform it into feature maps Q, K, and V respectively. Layer normalization is then performed on these three feature maps. Next, the probability map p... cThe three feature maps Q, K, and V are input into a sparse variation multi-head self-attention layer to obtain context-enhanced feature maps of the variation regions. These context-enhanced feature maps are then normalized and added to the input feature maps (i.e., a residual normalization layer), which is then input to the remaining layers of the current module. The residual normalized features are further input into a convolutional multilayer perceptron layer for local position encoding and nonlinear transformation, and then added to the input features to form a residual connection, yielding the output feature map. The sparse variation prior probability map p... c The prior probability of sparse variation is estimated by applying a fully convolutional network to the features output by a regular self-attention module.

[0057] In the sparsely variable multi-head self-attention layer, a probabilistic graph p is utilized. c Threshold segmentation is performed to obtain a coarse segmentation map of the changed regions. Based on this segmentation map, the features Q, K, and V are indexed spatially to obtain the feature Q of the changed regions. c ,K c V c Then perform the standard self-attention calculation.

[0058] 2.2 Parameterizing Subtask Factors Using Sparse Variation Self-Attention Modules. An existing visual backbone network is used as the image encoder to obtain a temporally independent hierarchical dense feature set from dual-temporal images. Furthermore, one conventional self-attention module and three sparse variation self-attention modules are constructed at the top layer of the image encoder for variation feature representation; specifically, the parametric image encoder is used to extract the hierarchical dense feature set {x}. (i)}, i = 2, 3, 4, 5, then the feature x (5) The input is fed into a regular self-attention module to enhance the long-range spatial context information of this feature, resulting in the context feature x. c This is used for subsequent prior estimation of sparse changes, yielding the prior probability p of the sparse changes. c ; the prior probability p of sparse variation c With hierarchical dense feature set {x (i) The input is fed into three sparse variation self-attention modules to further extract spatiotemporal context information, and the variables to be solved in the subtask factors are output, thus completing the parameterization; the sparse variation prior probability p is then used to further extract the spatiotemporal context information. c With hierarchical dense feature set {x (4) x (3) x (2) The input is fed into three sparse variation self-attention modules to further extract spatiotemporal context information, and the output variables to be solved in the subtask factors are then completed, thus achieving parameterization. The calculation of the sparse variation self-attention part adopts a progressive aggregation method, that is, first, the low-resolution features x are aggregated... (4) With contextual features x c Summation, along with the prior probability p of sparse variation.c The input is fed into the first sparse variation self-attention module to obtain the first spatiotemporal context feature; then the first spatiotemporal context feature is combined with the medium-resolution feature x. (3) Summation, along with the prior probability p of sparse variation. c The input is fed into the second sparse variation self-attention module to obtain the second spatiotemporal context feature; finally, the second spatiotemporal context feature is combined with the high-resolution feature x. (2) Summation, along with the prior probability p of sparse variation. c The input is fed into the third sparse variation self-attention module to obtain the third spatiotemporal context feature. The overall network construction is as follows: Figure 3 .

[0059] (III) Using the publicly available high-resolution remote sensing image change detection dataset SECOND, which contains 2968 pairs of multi-source optical remote sensing images, each image is 512×512 pixels in size, we split it into a training set (2374 pairs) and a test set (594 pairs), and performed geometric, radiometric and other data augmentation operations on the training set and normalized it.

[0060] 3.1 Acquire a large number of paired high spatial resolution remote sensing images using drones or high-resolution satellites. Since the publicly available dataset already contains images, this step can be skipped.

[0061] 3.2 Since the paired-annotated remote sensing images in this publicly available dataset already contain six types of ground interest (ROI) target samples, a change mask is obtained by comparing the target sample masks from two different time periods. The sample masks, change masks, and corresponding regional and temporal images are then used to construct a deep learning change detection dataset. An example of this dataset can be found in [link to dataset]. Figure 4 ;

[0062] 3.3 The dataset is divided into two parts in an approximate 4:1 ratio: a training set A (2374 pairs) for model training and a test set B (594 pairs) for evaluating model performance.

[0063] 3.4. Use Python to write an algorithm to normalize A and B, and only apply geometric and radiometric data augmentation methods such as horizontal flipping, vertical flipping, random rotation, and random color dithering to A.

[0064] (iv) Construct a unified optimization objective function for change detection, and train the deep change detection network architecture using the training set until the convergence condition is met.

[0065] Step 4.1, construct a unified optimization objective function for change detection, specifically in the following form:

[0066]

[0067] Where L regionTo estimate the loss function for sparse priors, Let t1 be the semantic segmentation loss function. Let L be the semantic segmentation loss function for phase t2. change Let λ1 and λ2 be the semantic modulation coefficients of the corresponding loss functions, both of which are 1.

[0068] Step 4.2: Train the depth change detection network using training set A. First, initialize the network parameters using the Kaiming initialization method. The model training parameters are set as follows: 15,000 training steps, 16 samples per batch, batch image size of 256×256, and an SGD optimizer with momentum of 0.9. The initial learning rate is set to 0.03, and a poly learning rate decay rule with a decay coefficient of 0.9 is used to dynamically adjust the learning rate. Iterate the forward computation and backpropagation of the model until the required number of iterations is reached.

[0069] (v) Input a brand new image to be detected and use it to predict the trained model. Then, use the probability change model in 1.1 to perform probabilistic inference to obtain the prediction. The result is as follows: Figure 5 .

[0070] The specific embodiments described herein are merely illustrative of the spirit of the invention. Those skilled in the art to which this invention pertains may make various modifications or additions to the described specific embodiments or use similar methods to substitute them, without departing from the spirit of the invention or exceeding the scope defined by the appended claims.

Claims

1. A remote sensing image change detection method based on sparse change self-attention mechanism, characterized in that, Comprise the following steps: Step 1, construct a unified probability change model framework, according to the assumption conditions of different change detection tasks, conditionally decompose the joint distribution of variables in the framework, and obtain subtask factors; Step 2, construct a sparse change self-attention module, and use its parameterization step 1 to obtain subtask factors to obtain a specific deep change detection network architecture; Step 3, construct a high-resolution remote sensing image change detection dataset, split it into a training set and a test set, and perform data enhancement operations on the training set and normalize it; The specific implementation of step 3 includes the following substeps, Step 3.1, use a UAV or high-resolution satellite to obtain a large number of paired high-spatial-resolution remote sensing images; Step 3.2, compare the two-phase target sample masks to obtain the change mask, and use the sample mask, change mask, and corresponding region, phase image to create a deep learning change detection dataset; Step 3.3, divide the deep learning change detection dataset into two parts according to a 4:1 ratio, a training set A for model training and a test set B for evaluating model performance; Step 3.4, normalize A and B, and only use A to perform geometric and radiometric data enhancement methods such as horizontal flipping, vertical flipping, random rotation, and random color jittering; Step 4, construct a unified optimization objective function for change detection, and use the training set obtained in step 3 to train the deep change detection network architecture obtained in step 2 until the convergence condition is met; The specific implementation of step 4 includes the following substeps, Step 4.1, construct a unified optimization objective function for change detection, which has the following specific form: wherein L region is a varying sparsity prior estimation loss function, is a semantic segmentation loss function for the t1 time phase, is a semantic segmentation loss function for the t2 time phase, L change is a binary varying loss function, and λ1 and λ2 are semantic modulation coefficients for the corresponding loss functions; Step 4.2, train the deep change detection network obtained in step 2 using the training set; first, initialize the network parameters, input data in batches, perform forward calculation in the network, obtain the loss value through the change detection unified optimization objective function, perform backward gradient calculation using the loss value, update the parameters using the neural network optimizer, and iterate until the model converges; Step 5, use the deep change detection network architecture and corresponding parameters obtained in step 4 to predict the test set, and solve the change probability map according to the corresponding probability change model to obtain the change region result.

2. The remote sensing image change detection method based on sparse change self-attention mechanism according to claim 1, characterized in that: The specific implementation of step 1 includes the following substeps, Step 1.1: Establish a probabilistic change model, considering the change as state S at time t. t State S at time t+1 t+1 The transition process, and the value of each state and its corresponding image I. t Directly related, therefore the joint distribution of the process from time t to t+1 is P(S t ,S t+1 ,I t ,I t+1 Based on the causal relationship between the state and the image, the joint distribution can be decomposed into several factors: P(S t ,S t+1 ,I t ,I t+1 )=P(I t )P(I t+1 |I t )P(S t |I t )P(S t+1 |S t ,I t+1 ) Where P(I) t ),P(I t+1 |I t ) respectively represent image I t with I t+1 The data distribution, which is based on image I t Given a conditional data distribution P(I) t+1 |I t ) indicates image I t+1 In the context of image I t The data was obtained from sampling the covered geographic locations; P(S) t |I t ),P(S t+1 |S t ,I t+1 ) represent the dependence of the state distribution at two time points on their corresponding images, where P(S) t+1 |S t ,I t+1 ) represents state S t+1 Simultaneously dependent on state S t With Image I t+1 ; Step 1.2, based on the probability change model, derive the two-class change detection theory framework; since the two-class change detection is state-independent change detection, that is, only focus on the occurrence of change, corresponding to the state transition in the probability change model; based on this objective assumption, the state-related factors P(S t |I t ), P(S t+1 |S t ,I t+1 ) can be merged to obtain P(S t+1 ,S t |I t+1 ,I t ), since the two-class change is irrelevant to the specific value of the state, the factor P(S t+1 ,S t |I t+1 ,I t ) is represented by P(C|I t+1 ,I t ), C is a random variable representing two-class change; the factor P(C|I t+1 ,I t ) is the theoretical framework of the two-class change detection model; Step 1.3, based on the probability change model, derive the semantic change detection theory framework; for semantic change detection, the specific value of the state needs to be considered, using the state-change condition independent assumption, that is, S t ⊥S t+1 |C, the factor P(S t |I t ) P(S t+1 |S t ,I t+1 ) can be merged and re-decomposed: P(S t |I t )P(S t+1 |S t ,I t+1 )=P(S t+1 ,S t |I t+1 ,I t )=P(S t+1 |C,I t+1 )P(S t |C,I t )P(C|I t+1 ,I t ) Among them, the subtask factor P(S) t+1 |C,I t+1 ) and P(S t |C,I t Both ) represent semantic segmentation tasks conditioned by two types of changes, and the subtask factor P(C|I) represents a semantic segmentation task. t+1 ,I t ) represents a binary change detection task; mathematically, this decomposition breaks down the semantic change detection task into two semantic segmentation tasks conditioned on change and one binary change detection task.

3. The remote sensing image change detection method based on sparse change self-attention mechanism according to claim 1, characterized in that: The specific implementation of step 2 includes the following substeps, Step 2.1, constructing a sparse change self-attention module; the sparse change self-attention module is stacked by a layer normalization layer, a sparse change multi-head self-attention layer, a residual layer normalization layer, and a convolutional multi-layer perception layer in this order; given an input feature map x and a sparse change prior probability map p c First, three 1x1 convolutional layers are used to transform them into feature maps Q, K, and V, respectively. The three feature maps are subjected to layer normalization processing. Then, the probability map p c is input into the sparse change multi-head self-attention layer along with the three feature maps Q, K, and V to obtain a feature map enhanced by the context of the change region. The context-enhanced feature map is further subjected to layer normalization and added to the input feature map, i.e., the residual layer normalization layer, and input into the remaining layers of the current module. The residual normalized feature is further input into the convolutional multi-layer perception layer for local position coding and nonlinear transformation, and added to the input feature to form a residual connection to obtain an output feature map. The sparse change prior probability map p c is obtained by applying a fully convolutional network to the features output by the conventional self-attention module for sparse change prior probability estimation. Step 2.2, parameterize the subtask factors using the sparse change self-attention module; use an existing visual backbone network as an image encoder to obtain a hierarchical dense feature set that is independent of the time phase of the image, and further construct one regular self-attention module and three sparse change self-attention modules at the top of the image encoder for change feature representation; Specifically, a parametric image encoder is used to extract a hierarchical dense feature set {x (i)}, i = 2, 3, 4, 5, and then the feature x (5) is input to a conventional self-attention module to enhance the long-distance spatial context information of the feature, to obtain a context feature x c , which is used for subsequent sparse change prior estimation, to obtain a sparse change prior probability p c ; Sparse change prior probability p c and hierarchical dense feature set {x (i)} are input into 3 sparse change self-attention modules to further extract spatio-temporal context information, and output the to-be-solved variables in the sub-task factors, i.e. parameterization is completed; sparse change prior probability p c and hierarchical dense feature set {x (4) , x (3) , x (2)} are input into 3 sparse change self-attention modules to further extract spatio-temporal context information, and output the to-be-solved variables in the sub-task factors, i.e. parameterization is completed; The calculation of the sparse variation self-attention part adopts progressive aggregation, that is, first aggregates the low-resolution features x (4) With contextual features x c Summation, along with the prior probability p of sparse variation. c The input is fed into the first sparse variation self-attention module to obtain the first spatiotemporal context feature; then the first spatiotemporal context feature is combined with the medium-resolution feature x. (3) Summation, along with the prior probability p of sparse variation. c The input is fed into the second sparse variation self-attention module to obtain the second spatiotemporal context feature; finally, the second spatiotemporal context feature is combined with the high-resolution feature x. (2) Summation, along with the prior probability p of sparse variation. c The input is fed into the third sparse variation self-attention module to obtain the third spatiotemporal context feature.

4. The remote sensing image change detection method based on sparse change self-attention mechanism according to claim 3, characterized in that: The existing visual backbone network includes ResNet, Swin-Transformer.

5. The remote sensing image change detection method based on sparse change self-attention mechanism according to claim 1, characterized in that: The change sparse prior estimation loss function and the two-class change loss function use the sum of the two-class cross-entropy loss function and the dice loss function, and the semantic segmentation loss function is the cross-entropy loss function.

Citation Information

Patent Citations

  • Remote sensing image building target efficient extraction method based on attention mechanism

    CN113780149A

  • System for deciding driving situation of vehicle

    KR102388806B1