Remote sensing image change detection method based on tensor decomposition and self-attention mechanism

By combining tensor decomposition and self-attention mechanism, the high computational complexity and information loss problems during high-resolution remote sensing image processing in the prior art are solved, and a more efficient and accurate change detection effect is achieved.

CN119964006AActive Publication Date: 2025-05-09XIDIAN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510077778.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-09
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

When processing high-resolution remote sensing images, the prior art has high computational complexity, serious information loss, and it is difficult to effectively combine local and global information, resulting in poor change detection effect.

Method used

The remote sensing image change detection method based on tensor decomposition and self-attention mechanism is adopted, and the spatial structure information of the image is preserved through the tensor neural network, and the global features are comprehensively captured in combination with the Transformer's self-attention mechanism.

Benefits of technology

It significantly improves the accuracy and efficiency of change detection, reduces the complexity of calculation steps and feature extraction, enhances the classification accuracy of complex land objects, and achieves rapid and accurate change detection under limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964006A_ABST
    Figure CN119964006A_ABST
Patent Text Reader

Abstract

The invention provides a remote sensing image change detection method based on tensor decomposition and a self-attention mechanism. The method mainly solves the problem that an existing detection technology is poor in complex ground feature change classification effect. According to the scheme, the method comprises the following steps: 1) preprocessing an input image to obtain three-channel image data; 2) constructing a detection model in which a tensor neural network and Transform are combined, and taking the three-channel image data as model input; 3) space structure information of the image is reserved by using a tensor neural network, the calculation amount is reduced, global features are comprehensively captured by a self-attention mechanism, and the accuracy and robustness of change detection are enhanced; 4) performing iterative training on the model to realize optimization; and 5) performing pixel-level classification on the global feature tensor obtained by the optimized model through a classifier to generate a change detection image. The method can improve the precision and efficiency of change detection under the condition of not increasing significant calculation overhead, and is suitable for change detection tasks of high-resolution remote sensing images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image detection technology, and further relates to remote sensing image change detection technology, specifically a remote sensing image change detection method based on tensor decomposition and self-attention mechanism, which can be used in the fields of urban expansion, ecological environment monitoring, and natural disaster response. Technical Background

[0002] Remote sensing image change detection is a key technology in remote sensing data processing, which aims to detect surface changes by comparing remote sensing images at different time points. Change detection technology has been widely used in urban expansion, ecological environment monitoring, natural disaster response and other fields. With the acceleration of global urbanization and the increasing severity of environmental problems, how to accurately and efficiently detect surface changes has become an urgent need in urban planning, environmental management and disaster assessment. However, with the improvement of remote sensing image resolution and the increase of data dimensions (such as multispectral and hyperspectral images), traditional change detection methods have gradually exposed many shortcomings when dealing with high-dimensional data. For example, early pixel-level change detection methods, such as change vector analysis (CVA) and differential image method, although they have good results in low-resolution data, they perform poorly in high-resolution data. The main reason is that they are too sensitive to noise and cannot make full use of complex spatial structure information.

[0003] In recent years, the development of deep learning technology has provided new ideas for change detection in remote sensing images. In particular, models such as convolutional neural networks (CNNs) and Transformers have performed well in feature extraction and classification of large-scale image data. CNN extracts local features of images through convolution kernels and has achieved remarkable results in image classification and change detection tasks. For example, Daudt et al. proposed a change detection method based on deep convolutional networks, which extracts image features through multi-layer convolutions and uses differential images to judge surface changes. Nevertheless, the local receptive field of CNN limits its ability to capture global features, and it is difficult to effectively combine local and global information when faced with complex surface changes. The Transformer model has shown great advantages in extracting long-distance dependent features by introducing a self-attention mechanism, and has gradually been introduced into the change detection task of remote sensing images. The Vision Transformer (ViT) model proposed by Dosovitskiy et al. achieves image classification and feature extraction by dividing the image into blocks and performing self-attention analysis on the relationship between these blocks. However, when processing images, Transformer needs to flatten the two-dimensional image into a one-dimensional vector. This flattening process may cause the loss of spatial information of the image, especially for high-resolution and multispectral images, the impact is more obvious, which limits the application of Transformer in complex remote sensing tasks.

[0004] Current change detection algorithms often involve a lot of calculations and highly complex network structures when processing high-resolution remote sensing images, resulting in high computational costs and resource consumption of the algorithms, especially in scenarios with limited hardware resources, which make it difficult to achieve efficient operation. Summary of the invention

[0005] The purpose of the present invention is to overcome the deficiencies of the above-mentioned prior art, and propose a remote sensing image change detection method based on tensor decomposition and self-attention mechanism, which mainly solves the problem that the prior art has poor classification effect on complex land object changes. Traditional models usually have high computational complexity, information loss and limited classification effect. The present invention introduces tensor decomposition and combines it with the self-attention mechanism to construct a new remote sensing image change detection network, uses tensor neural network to retain the spatial structure information of the image and reduce the amount of calculation, and the self-attention mechanism comprehensively captures global features, enhances the accuracy and robustness of change detection, and improves the accuracy and efficiency of change detection without increasing significant computational overhead. It is suitable for change detection tasks of high-resolution remote sensing images.

[0006] To achieve the above object, the technical solution adopted by the present invention includes the following steps:

[0007] (1) Preprocessing the input multi-temporal remote sensing images to reduce the illumination and geometric differences in the input images and obtain preprocessed three-channel image data;

[0008] (2) Construct a tensor transformer model that combines a tensor neural network with a transformer. Use the preprocessed three-channel image data as the model input data and obtain global features through model processing. The implementation steps are as follows:

[0009] (2.1) Decomposing and optimizing the preprocessed three-channel image data through the tensor neural network, extracting data features, and obtaining the core tensor G;

[0010] (2.2) Use the self-attention mechanism of Transformer to globally model the features extracted by the tensor decomposition network and mine the long-distance dependencies and change features in the image; that is, use the core tensor G of the tensor neural network directly as the input unit of Transformer to preserve the spatial and spectral structure of the image, and then calculate based on Transformer to obtain the global feature tensor F;

[0011] (3) Iteratively train the tensor transformer model and optimize it by continuously adjusting the model parameters to obtain the final model for the change detection task;

[0012] (4) The global feature tensor obtained by the final model is classified at the pixel level through the classifier to generate a binary change detection map with the same size as the input image.

[0013] Compared with the prior art, the present invention has the following beneficial effects:

[0014] First, the present invention effectively extracts key features from remote sensing images by introducing tensor neural networks, retains spatial structural information, and avoids the loss of a large amount of structural information in traditional flattening operations. Compared with traditional high-dimensional data processing methods, the tensor decomposition method significantly improves the accuracy of change detection.

[0015] Second, the present invention innovatively proposes a new network structure that combines tensor decomposition with the Transformer model. It reduces the number of model parameters through tensor decomposition, and uses the Transformer's self-attention mechanism to capture the global features in the image, thus overcoming the limitation of the limited receptive field of the convolutional neural network. In the change detection task, the network can extract features more comprehensively and effectively improve the classification accuracy of complex land object changes.

[0016] Third, since the present invention adopts a method that combines efficient tensor decomposition with Transformer, the calculation steps are reduced and the complexity of the feature extraction process is optimized; through the combination of effective feature dimensionality reduction and self-attention mechanism, the model shows higher efficiency and stronger robustness in the change detection task of high-resolution and multispectral images, thereby achieving fast and accurate change detection under limited resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0018] Figure 1 Flow chart for realizing the method of the present invention;

[0019] Figure 2 A schematic diagram of the structure of the tensor Transformer model constructed in the present invention;

[0020] Figure 3 An example of the LEVIR-CD dataset used in the present invention;

[0021] Figure 4 Schematic diagram of the simulation result of scenario 1 in the embodiment of the present invention;

[0022] Figure 5 This is a simulation result diagram of scenario 2 in an embodiment of the present invention. DETAILED DESCRIPTION

[0023] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.

[0024] Example 1: Reference Figure 1-2 This example provides a remote sensing image change detection method based on tensor decomposition and self-attention mechanism. The specific implementation steps are as follows:

[0025] Step 1. Preprocess the input multi-temporal remote sensing image, including brightness normalization and image registration operations, to reduce the illumination and geometric differences in the input image and obtain the preprocessed three-channel image data.

[0026] The above brightness normalization operation is as follows:

[0027] Let X1(t) and X2(t) be two multi-temporal remote sensing data, and use the global normalization formula to standardize each image:

[0028]

[0029] Among them, μ i and σ i Respectively represent data X i The mean and standard deviation of X i ' is the normalized data.

[0030] The above-mentioned image registration operation specifically aligns corresponding pixel points in multi-temporal images to the same geographic coordinate system. The registration method includes the following steps:

[0031] (1.1) Use the scale-invariant feature transform algorithm SIFT or the orientation-rotation algorithm ORB to extract key point features in the image;

[0032] (1.2) Match the key point features of multi-temporal images using the nearest neighbor algorithm;

[0033] (1.3) Aligning the two images using a geometric transformation method, where the geometric transformation method includes at least an affine transformation and a projection transformation.

[0034] Step 2. Build a tensor transformer model that combines tensor neural network and transformer, such as Figure 2 As shown, the tensor neural network uses the tensor feature extraction layer to realize feature extraction. The feature extraction layer uses the inverse process of Tucker decomposition, that is, a matrix group containing N matrices is used to extract features from the original data x to obtain tensor y. The preprocessed three-channel image data is used as the model input data, and the global features are obtained after model processing. The implementation steps are as follows:

[0035] (2.1) The preprocessed three-channel image data is feature decomposed and optimized through the tensor neural network, the data features are extracted, and the core tensor G is obtained; the implementation steps are as follows:

[0036] (2.1.1) Design a tensor neural network that combines multi-layer tensor decomposition and nonlinear mapping, which includes an input layer, a multi-layer Tucker decomposition module, a ReLU activation function layer, and a normalization layer;

[0037] (2.1.2) The input layer receives the preprocessed three-channel image and stacks it in time to construct the input tensor X:

[0038] X∈R H×W×2C ,

[0039] Where H and W represent the height and width of the image respectively; C is the number of channels of each image;

[0040] (2.1.3) The multi-layer Tucker decomposition module converts the input tensor X into a feature representation through a tensor decomposition operation, that is, the tensor X is decomposed into a core tensor G and a factor matrix using Tucker decomposition:

[0041] X≈G×1U1×2U2×3U3,

[0042] Among them, 1U1 is the factor matrix of H and the local spatial features of the retained image; 2U2 is the factor matrix of W and the local spatial features of the retained image; 3U3 is the factor matrix of the channel dimension;

[0043] (2.1.4) After each decomposition layer, the ReLU activation function and batch normalization operation are applied to enhance the feature representation capability, accelerate training convergence, and output the core tensor G.

[0044] (2.2) The self-attention mechanism of Transformer is used to globally model the features extracted by the tensor decomposition network, and to mine the long-distance dependencies and change features in the image; that is, the core tensor G of the tensor neural network is directly used as the input unit of Transformer to preserve the spatial and spectral structure of the image, and then the global feature tensor F is obtained by calculation based on Transformer. The implementation steps are as follows:

[0045] (2.2.1) The core tensor G generated by the tensor decomposition network:

[0046]

[0047] Among them, r1 represents the spatial dimension after H compression; r2 represents the spatial dimension after W compression; r3 is the channel dimension;

[0048] (2.2.2) Directly regard r3 as the number of tensor blocks, take each tensor block as the input of a self-attention unit, and use the Transformer's multi-head self-attention mechanism MHSA to operate on the tensor block to obtain the global feature tensor F:

[0049]

[0050] Where d is the number of output channels of Transformer.

[0051] The above uses Transformer's multi-head self-attention mechanism MHSA to operate on tensor blocks, where the calculation formula for each attention head is as follows:

[0052]

[0053] Where Q, K, and V represent the matrices of query, key, and value, respectively; d kRepresents the normalization factor of the attention weight.

[0054] Step 3. Iteratively train the tensor transformer model, optimize by continuously adjusting the model parameters, and obtain the final model for the change detection task. In this embodiment, the predicted change detection map Y and the actual labeled change detection map Y are minimized by supervised learning. true The error between is optimized; the steps are as follows:

[0055] (3.1) Load the training dataset, including multi-temporal remote sensing image pairs [X1, X2] and corresponding change detection labels Y true ;

[0056] (3.2) The input image pair [X1, X2] is passed through the tensor neural network and Transformer in turn, and finally a change detection map P is generated through the classifier;

[0057] (3.3) Through the cross entropy loss function L BCE Calculate the predicted graph P and the labeled graph Y true The error between; the cross entropy loss function L BCE The binary cross entropy loss BCE is used as the objective function, which is defined as follows:

[0058]

[0059] Where N is the number of pixels; Y true (i) is the annotated change detection label, which takes the value of 0 or 1; P(i) is the output probability of the model, which takes the value between [0,1];

[0060] (3.4) Back-propagate the error to the weights of the tensor neural network and Transformer, and use the back-propagation algorithm to adjust the parameters of each layer;

[0061] (3.5) Use the Adam optimizer or SGD optimizer to update the trainable parameters of the network and minimize the loss function.

[0062] Step 4. Use the classifier to perform pixel-level classification on the global feature tensor obtained by the final model to generate a binary change detection map with the same size as the input image. The implementation is as follows:

[0063] (4.1) Use the convolution classifier to directly convert the feature tensor F into a change probability map P of the same size as the input image:

[0064] P = σ(Conv 1×1 (F))∈R H×W ,

[0065] Among them, Conv1×1 is a 1×1 convolution operation, which is used to compress the number of channels d to 1; σ is the Sigmoid activation function, which is used to map the output of the convolution to the interval [0,1];

[0066] (4.2) The probability map P is converted into a binary change detection map Y using a binarization operation:

[0067]

[0068] Among them, P(x,y) represents the change probability of pixel (x,y) in the probability map, τ is a threshold; 1 represents the changed area, and 0 represents the unchanged area.

[0069] Embodiment 2: The overall implementation steps of the detection method provided in this embodiment are the same as those in Embodiment 1. Now, an example is given to further describe the implementation process of the present invention in detail:

[0070] Step 1. Preprocess the input multi-temporal remote sensing images. In order to ensure the consistency of the input image data and eliminate unnecessary noise interference, the two multi-temporal remote sensing data X1(t) and X2(t) are first preprocessed, including:

[0071] Brightness normalization: Remote sensing images usually have brightness differences due to sensor conditions or lighting changes. In order to reduce the impact of such differences on change detection results, each image is normalized using the global normalization formula:

[0072]

[0073] Among them, μ i and σ i Respectively represent data X i The mean and standard deviation of X i ' is the normalized data.

[0074] Image registration: Multi-temporal remote sensing images may be spatially misaligned due to factors such as sensor viewing angle, shooting time or terrain undulation. In order to ensure the accuracy of image change detection, multi-temporal images need to be registered. Image registration aims to align corresponding pixels in multi-temporal images to the same geographic coordinate system. The registration method usually includes the following steps: 1a) Feature point extraction: Use algorithms such as SIFT (Scale-Invariant Feature Transform) or ORB (OrientedFAST and Rotated BRIEF) to extract key point features in the image. 1b) Feature point matching: Match the key points of multi-temporal images through the nearest neighbor algorithm or other matching strategies. 1c) Registration transformation: Use geometric transformation methods such as affine transformation and projection transformation to align the two images.

[0075] Through the preprocessing steps of brightness normalization and image registration, the illumination and geometric differences in the input images can be effectively reduced, ensuring that the subsequent tensor decomposition and self-attention mechanism can accurately extract the change features.

[0076] Step 2. The tensor neural network extracts data features and reduces the amount of computation. In the remote sensing image change detection task, three-channel data usually contains rich spatial and spectral information (such as RGB or near-infrared combination), and direct dimensionality reduction may lead to information loss. Therefore, in this step, the core goal of the tensor decomposition network is not simply dimensionality reduction, but to extract key features and reduce computational complexity by performing efficient feature decomposition and optimized representation of the three-channel data. The specific steps are as follows:

[0077] 1) Stack the preprocessed three-channel images X1(t) and X2(t) by time to construct the input tensor X:

[0078] X∈R H×W×2C

[0079] Where: H and W represent the height and width of the image, respectively; C is the number of channels of each image (for example, C = 3 for RGB channels); 2C represents the number of channels of a dual-temporal image. The tensor X is used as the input of the tensor decomposition network, which avoids the spatial correlation lost by the traditional flattening operation by maintaining the structure of multi-channel and spatial information.

[0080] 2) Tensor decomposition operation. Through the tensor decomposition operation, the input tensor X is converted into a more compact feature representation, rather than directly reducing the dimension. The main goal is to extract key features between channels and spatial dimensions. Tucker decomposition is used to decompose the tensor X into a core tensor and a factor matrix:

[0081] X≈G×1U1×2U2×3U3

[0082] Among them, G is the core tensor, which represents the efficient and compact features of the image; 1U1 and 2U2 are the factor matrices of the spatial dimension, which retain the local spatial features of the image; 3U3 is the factor matrix of the channel dimension, which extracts the relevant information between channels. Through this decomposition method, it is possible to reduce redundant calculations and storage requirements while maintaining the image structure information.

[0083] 3) The tensor neural network further optimizes the image features by combining multi-layer tensor decomposition and nonlinear mapping. The specific design includes: a) Input layer: accepts the tensor representation X of the three-channel stacked image. b) Multi-layer tensor decomposition: gradually optimizes the features by stacking multiple layers of Tucker decomposition modules. The core tensor G is compressed in each layer to extract more abstract spatial and channel features. c) Activation and normalization: After each decomposition layer, the ReLU activation function and batch normalization operation are applied to enhance the feature representation capability and accelerate training convergence. The key advantage of the tensor neural network is that it not only retains the spatial-channel structure of the image, but also effectively reduces redundant information, while providing high-quality feature representation for subsequent models.

[0084] 4) Although three-channel data does not require direct dimensionality reduction, the tensor decomposition network significantly reduces the computational complexity by optimizing data representation. The main reasons include: 1) Compact feature representation: The dimension of the core tensor G is significantly smaller than the original tensor X, which reduces the amount of subsequent operations. 2) Avoid flattening operations: Traditional methods need to flatten the image into a one-dimensional sequence, resulting in a computational complexity of O(n 2 )(n is the number of pixels). The length of the flattened sequence is n, which corresponds to the total number of pixels in the image. If a self-attention mechanism such as Transformer is used, its computational complexity is O(n 2 ), because it is necessary to calculate the correlation between each unit in the sequence. This is not suitable for high-resolution images (such as 1024*1024 size data, n=10 6 ) will cause extremely high computational requirements. The flattening operation loses the spatial structural information of the image (such as the relationship between adjacent pixels), causing the model to relearn these relationships through additional calculations, further increasing the computational burden. Tensor decomposition operates directly on the tensor form, retaining structural information and reducing redundant calculations. 3) Efficient feature extraction: By jointly modeling spatial and channel information, tensor decomposition can significantly reduce unnecessary calculations while extracting key change features.

[0085] 5) Output. The output of the tensor decomposition network is the core tensor G, which retains the key spatial and spectral features of the image and has a much smaller dimension than the original input tensor X. The compact representation of the core tensor G not only effectively reduces the computational complexity, but also provides high-quality input for the subsequent Transformer module. Through tensor decomposition, the multi-dimensional features of the image are optimized into an efficient feature representation. Compared with traditional methods, the core tensor G reduces redundant information while avoiding the loss of structural information caused by the flattening operation, laying the foundation for subsequent global feature extraction and change detection.

[0086] Step 3. Transformer processing features based on tensor units. The core goal of this step is to use the self-attention mechanism of the Transformer to further globally model the features extracted by the tensor decomposition network, and to mine long-distance dependencies and change features in the image. In order to solve the high computational complexity problem caused by the traditional Transformer flattening the image into a one-dimensional sequence, the present invention directly uses the core tensor G of the tensor decomposition network as the input unit, retaining the spatial and spectral structure of the image, thereby greatly reducing the computational cost.

[0087] s1. Input data. In step 2, the input remote sensing image X is processed by the tensor decomposition network to generate the core tensor G:

[0088]

[0089] Among them, r1 and r2 represent the compressed spatial dimensions, r3 is the channel dimension (or the number of tensor blocks), and the tensor G is used as the input of the Transformer model in this step. Different from the traditional Transformer operation that needs to flatten the image into a one-dimensional sequence, the present invention directly performs block-level operations on G, retaining the spatial and channel information of the image, avoiding the high computational complexity caused by serialization.

[0090] s2. Transformer self-attention mechanism. In traditional Transformer, the computational complexity of the self-attention mechanism is O(n 2 ), where n is the number of pixels after flattening. This high complexity will lead to serious computational overhead in high-resolution images. In order to solve this problem, the improvement measures of the present invention are as follows: 1) Input design based on tensor blocks. Different from traditional pixel serialization, the tensor G maintains the tensor structure of the image. For dimension G∈R r1×r2×r3 Instead of flattening r1 and r2, r3 is directly regarded as the number of "tensor blocks", each of which is the input of a self-attention unit. In this way, the attention mechanism no longer operates on O(n 2 ), but operates on tensor blocks, which greatly reduces the computational complexity. 2) Multi-Head Self-Attention (MHSA). The present invention adopts a multi-head self-attention mechanism, using multiple independent attention heads to focus on different features in the image. The calculation formula for each attention head is as follows:

[0091]

[0092] Where Q, K, and V represent the matrices of query, key, and value, respectively; d krepresents the normalization factor of the attention weight; these operations are performed on the channel dimension r3 of the tensor block, and the computational complexity becomes O(r3 2 ), which is much lower than the traditional O(n 2 ).

[0093] S3. Output. Through the operation of the multi-head self-attention mechanism, the global feature tensor F is obtained:

[0094]

[0095] Among them, r1 and r2 represent the spatial dimensions of the image, and d is the number of output channels of the Transformer (usually greater than r3).

[0096] Step 4. In this step, the global feature tensor F is classified at the pixel level by a classifier to generate a binary change detection map that is consistent with the size of the input image.

[0097] In order to improve the accuracy of detection, the present invention also involves the optimization and training of the model. By continuously adjusting the model parameters, the performance of the model in the change detection task can be improved.

[0098] 1. Input. The global feature tensor F obtained from step 3 is used as the input of this step and is expressed as:

[0099]

[0100] Among them, r1 and r2 represent the spatial dimensions of the image, and d represents the channel dimension of the feature, which is usually larger than the number of channels of the original input.

[0101] 2. The process of generating change detection graph:

[0102] Convolution classifier: The purpose is to directly convert the feature tensor F into a change detection map P of the same size as the input image. The specific operation is: use a 1×1 convolution operation to perform channel compression on the feature tensor F:

[0103] P = σ(Conv 1×1 (F))∈R H×W

[0104] Among them, Conv 1×1 is a 1×1 convolution operation, the purpose of which is to compress the number of channels d to 1; σ is the Sigmoid activation function, which maps the output of the convolution to the interval [0,1], indicating the probability of change.

[0105] Thresholding to generate change detection map: The purpose is to convert the probability map P into a binary change detection map Y in order to identify the changed area in the image. The operation is a simple binarization operation, defined as follows:

[0106]

[0107] Among them, P(x,y) represents the change probability of pixel (x,y) in the probability map, τ is a threshold, usually set to 0.5; 1 represents the changed area, and 0 represents the unchanged area.

[0108] 3. Model training and optimization. In the process of generating detection results, model training and optimization are essential parts. The goal of model training is to minimize the difference between the predicted change detection map Y and the actual labeled change detection map Y through supervised learning. true 1) Definition of loss function. The present invention adopts binary cross-entropy loss (BCE) as the main objective function to calculate the error between the detection image P and the annotation image Y predicted by the model. true The error between . The loss function is defined as follows:

[0109]

[0110] Where N is the number of pixels; Y true (i) is the annotated change detection label, which takes a value of 0 or 1; P(i) is the output probability of the model, which takes a value between [0,1]. 2) Steps of the training process. The training process of the model is usually divided into the following steps: Load the training dataset, including multi-temporal remote sensing image pairs [X1, X2] and the corresponding change detection label Y true ; Forward propagation: the input image pair [X1, X2] passes through the tensor decomposition network and Transformer in turn, and finally generates a change detection map P through the classifier; Loss calculation: through the cross entropy loss function L defined above BCE Calculate the predicted graph P and the labeled graph Y true ; Back propagation: propagate the error back to the weights of the tensor network and Transformer, and use the back propagation algorithm to adjust the parameters of each layer; Parameter update: Use the Adam optimizer or SGD optimizer to update the trainable parameters of the network and minimize the loss function.

[0111] 4. Output. The generated detection result image Y is a binary image with the same size as the input image:

[0112] Y∈{0,1} H×W

[0113] Among them, 1 represents the area that has changed; 0 represents the area that has not changed. In addition, the output detection map Y can be saved as an image file (such as PNG, TIFF format) or used for visualization. In the actual change detection task, the change detection map Y is usually used as a comparative visualization result, and is placed together with the original images X1 and X2 to intuitively show which areas have changed.

[0114] The effect of the present invention is further described below in conjunction with simulation experiments.

[0115] 1. Simulation conditions:

[0116] The simulation experiment of the present invention is carried out in the hardware environment of GPU3070ti and the software environment of Python3.8.

[0117] 2. Simulation content:

[0118] In the simulation process, the present invention adopts a remote sensing image change detection method based on Tensor Transformer to simulate the remote sensing image change detection task under the same scene at different times. The specific steps include:

[0119] 1) Data preprocessing: Standardize remote sensing images to meet the input requirements of the model.

[0120] 2) Model training: The proposed Tensor Transformer model is trained using the training dataset.

[0121] 3) Model validation: Evaluate model performance using the test dataset.

[0122] 3. Simulation results: The simulation results are as follows: Figure 4 and Figure 5 .

[0123] See attached Figure 2 , commonly used data include A before the change, B after the change and the change label, respectively, as Figure 2 As shown in (a)-(c) in the figure. (a): The image before the change, which is one of the input images in the change detection task of the present invention, that is, the image A before the change, shows the surface coverage information of a specific area. (b): The image after the change, which is the second input image in the change detection task of the present invention, that is, the image B after the change, shows the surface information of the same area in different time periods. (c): The change annotation map, which is the change annotation map (Ground Truth) used to train the model, in which the white area represents the changed pixels and the black area represents the unchanged pixels.

[0124] See attached Figure 2, which shows the structure of the model proposed in the present invention, specifically including the following modules: 2a) Input module. The input image pair includes an image before the change (Input A) and an image after the change (Input B). These two image pairs represent the surface information of a certain area at different times, and are input into the model for change detection analysis. 2b) TBDM module performs feature extraction and differential analysis on the input image pair through a tensor neural network (TBDM). This module uses tensor operations to directly extract the difference features of the image in the high-dimensional space, avoiding the flattening operation in the traditional method, effectively reducing the computational complexity, and retaining the spatial and spectral information of the image. 2c) Encoder module (Encoder). The encoder further extracts the features output by the TBDM module and compresses the high-dimensional features into low-dimensional representations. The encoder captures the local and global features of the image through convolution operations and multi-layer network structures, providing support for the subsequent decoder to generate change detection maps. 2d) Decoder module (Decoder). The decoder is responsible for reconstructing the low-dimensional features output by the encoder into a change detection map that is consistent with the size of the input image. Through step-by-step upsampling operations, the decoder can restore the spatial resolution of the image and generate a change detection map. 2e) Output module. The final output includes the predicted change detection map (Predict) and the labeled change detection map (Label). The prediction results are used to compare and evaluate the model performance, where the white area represents the surface pixels that have changed and the black area represents the surface pixels that have not changed. Figure 2 The model structure of the present invention and the synergistic effect of each module are clearly demonstrated, and the innovation and practical application value of the present invention in remote sensing image change detection are intuitively reflected.

[0125] Reference Figure 4 and Figure 5 , these two figures show the prediction results of the present invention in the remote sensing image change detection task, clearly demonstrating the model's ability to detect changed areas. Specifically, they include:

[0126] 1. The image before the change, such as Figure 4 (a) and Figure 5 As shown in (a):

[0127] Figure 4 and Figure 5 (a) in the figure shows the images before the change in the two scenes, showing the surface information of the target area before the change.

[0128] 2. The changed image, such as Figure 4 (b) and Figure 5 As shown in (b):

[0129] Figure 4 and Figure 5(b) in the figure shows the changed images in the two scenes, reflecting the surface information of the same area after the change.

[0130] 3. Change detection results predicted by the model, such as Figure 4 (c) and Figure 5 As shown in (c):

[0131] Figure 4 and Figure 5 (c) in the figure are the change detection results predicted by the present invention for two different scenes, wherein: the white area represents the detected change area; the black area represents the area without change.

[0132] 4.Real change annotation, such as Figure 4 (d) and Figure 5 As shown in (d):

[0133] Figure 4 and Figure 5 (d) in the figure are respectively the change detection images with real annotations for two different scenes of the present invention, which are used to evaluate the model performance.

[0134] The network model proposed in this invention is shown in the following performance indicators:

[0135] Performance indicators:

[0136] 1. Precision: 0.8850. The model can effectively avoid false positives and accurately identify pixels in the changed area.

[0137] 2. Recall: 0.9524. The model performs well in covering the changed area and can capture all the change information to a large extent.

[0138] 3. F1 score (F1): 0.9175. As a comprehensive indicator of precision and recall, the F1 value reflects the excellent performance of the model in balancing false positives and false negatives.

[0139] 4. Mean Intersection over Union (MIoU): 0.8475. The MIoU indicator reflects the degree of overlap between the predicted results and the actual change area, indicating that the model has a high adaptability in different scenarios.

[0140] Training metrics:

[0141] 1. Single epoch training time: 9 minutes and 32 seconds. This result demonstrates the efficient training performance of the model on a large-scale remote sensing dataset.

[0142] 2. Total model parameters: 0.9M. The model is compact in design and has a small number of parameters, making it suitable for resource-constrained embedded environments.

[0143] 3. Test images: 2048 pairs, total prediction time: 66 seconds. In practical applications, the model has the ability of fast reasoning and can efficiently complete the change detection task of large-scale images.

[0144] The performance of the model has reached a high level in key indicators such as precision, recall, F1 score and MIoU, proving its superiority in change detection tasks. In particular, its low computational complexity and efficient reasoning ability make the model have broad application prospects in practical scenarios.

[0145] The present invention can achieve the following main purposes:

[0146] (I) Effectively reduce spatial information loss and improve detection accuracy

[0147] The traditional Transformer model needs to flatten the two-dimensional image into a one-dimensional vector, which leads to the loss of a large amount of spatial structure information, especially in high-resolution and multi-spectral images. This information loss significantly affects the detection accuracy. The present invention adopts a tensor neural network to retain the spatial and spectral information of the image, and represents the image features in a tensor manner, effectively avoiding the information loss caused by the flattening operation. Therefore, compared with the existing Transformer method, the present invention significantly improves the accuracy of change detection, especially when dealing with complex surface change scenes.

[0148] (II) Reduce computational complexity and improve computational efficiency

[0149] Although Transformer is good at extracting local features when extracting high-resolution image features, it often requires a complex calculation process, resulting in high computational complexity and long calculation time. The present invention optimizes the feature extraction process by combining tensor decomposition with the Transformer self-attention mechanism. Table 2 shows the comparison of different models in terms of parameter quantity and computational complexity under the same computing resources. Compared with other Transformer models, the present invention has significantly fewer parameters and lower computational complexity.

[0150] The comparison of the computational complexity of the traditional Transformer model and the model of the present invention is shown in Table 1, and the comparison of the model parameters and the training complexity is shown in Table 2:

[0151] Table 1. Comparison of computational complexity

[0152]

[0153] Table 2. Comparison of model parameters and training complexity

[0154]

[0155] Tensor decomposition is used to reduce the dimensionality of input data, thereby reducing the complexity of network calculations and ensuring lower computational costs. Compared with traditional deep learning methods, the present invention can complete change detection tasks more efficiently and is suitable for resource-constrained environments, such as embedded remote sensing data processing systems.

[0156] (III) Improve the ability to capture global features and enhance model robustness

[0157] In the prior art, convolutional neural networks are difficult to effectively capture global information in images due to the limitations of their receptive fields. In particular, in change detection tasks, the lack of global information often leads to inaccurate recognition of changes in objects. The present invention uses a self-attention mechanism that can model the global dependencies of images during feature extraction, and combines tensor neural networks to ensure the comprehensive use of multi-dimensional information. This combination enables the model to exhibit higher robustness in capturing the global features of complex objects, effectively reducing false detections and missed detections caused by incomplete local features.

[0158] The remote sensing image change detection method based on tensor decomposition and self-attention mechanism of the present invention shows good prospects in multiple practical application fields. First, in urban expansion and environmental monitoring, the present invention can be used to monitor the development and changes of urban areas, dynamic changes of ecological environment such as forest coverage in real time, and can provide timely data support for scientific planning and environmental protection decision-making. Secondly, the present invention also has important application value in natural disaster assessment and emergency response, and can quickly and accurately assess the surface changes in the disaster area, providing a reliable basis for formulating rescue measures and post-disaster reconstruction. In the field of agricultural management, the present invention helps farmers and agricultural managers to take effective measures in time to improve agricultural production efficiency and management accuracy by monitoring the growth status of crops and pests and diseases. In addition, due to the low computational complexity and small number of parameters, the present invention is very suitable for deployment in embedded devices and edge computing scenarios to achieve real-time change detection in resource-constrained environments. In addition, the present invention also has a wide range of application potential in other remote sensing image processing tasks, such as image classification and target detection, and can provide efficient and economical solutions for fields such as geographic information systems GIS, military reconnaissance and resource census, and promote the further development and application of remote sensing technology.

[0159] The above comparative analysis proves the correctness and effectiveness of the method proposed in the present invention.

[0160] The parts not described in detail in the present invention belong to the common knowledge of those skilled in the art. The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Obviously, for professionals in this field, after understanding the content and principle of the present invention, it is possible to make various modifications and changes in form and details without departing from the principle and structure of the present invention. However, these solutions based on the idea of ​​the present invention and using the multi-agent reinforcement learning algorithm and the time slot allocation algorithm for joint design are all within the scope of protection of the present invention.

Claims

1. A remote sensing image change detection method based on tensor decomposition and self-attention mechanism, characterized in that: The steps include: (1) Preprocessing the input multi-temporal remote sensing images to reduce the illumination and geometric differences in the input images and obtain preprocessed three-channel image data; (2) Construct a tensor transformer model that combines a tensor neural network with a transformer. Use the preprocessed three-channel image data as the model input data and obtain global features through model processing. The implementation steps are as follows: (2.1) Decomposing and optimizing the preprocessed three-channel image data through the tensor neural network, extracting data features, and obtaining the core tensor G; (2.2) Use the self-attention mechanism of Transformer to globally model the features extracted by the tensor decomposition network and mine the long-distance dependencies and change features in the image; that is, use the core tensor G of the tensor neural network directly as the input unit of Transformer to preserve the spatial and spectral structure of the image, and then calculate based on Transformer to obtain the global feature tensor F; (3) Iteratively train the tensor transformer model and optimize it by continuously adjusting the model parameters to obtain the final model for the change detection task; (4) The global feature tensor obtained by the final model is classified at the pixel level through the classifier to generate a binary change detection map with the same size as the input image.

2. The method according to claim 1, characterized in that: The preprocessing in step (1) includes brightness normalization and image registration operations.

3. The method according to claim 2, characterized in that: the brightness normalization operation is specifically as follows: Let X1(t) and X2(t) be two multi-temporal remote sensing data, and use the global normalization formula to standardize each image: in, μ i and σ i Respectively represent data X i The mean and standard deviation of X i ' is the normalized data.

4. The method according to claim 2, characterized in that: The image registration operation specifically aligns corresponding pixel points in multi-temporal images to the same geographic coordinate system. The registration method includes the following steps: (1.1) Use the scale-invariant feature transform algorithm SIFT or the orientation-rotation algorithm ORB to extract key point features in the image; (1.2) Match the key point features of multi-temporal images using the nearest neighbor algorithm; (1.3) Aligning the two images using a geometric transformation method, where the geometric transformation method includes at least an affine transformation and a projection transformation.

5. The method according to claim 1, characterized in that: The tensor neural network described in step (2) realizes feature extraction by using a tensor feature extraction layer. The feature extraction layer uses the inverse process of Tucker decomposition, that is, a matrix group containing N matrices is used to extract features from the original data x to obtain a tensor y.

6. The method according to claim 1, characterized in that: Step (2.1) performs feature decomposition and optimization on the preprocessed three-channel image data through a tensor neural network, extracts data features, and obtains a core tensor G; the implementation steps are as follows: (2.1.1) Design a tensor neural network that combines multi-layer tensor decomposition and nonlinear mapping, which includes an input layer, a multi-layer Tucker decomposition module, a ReLU activation function layer, and a normalization layer; (2.1.2) The input layer receives the preprocessed three-channel image and stacks it in time to construct the input tensor X: X∈R H×W×2C , Where H and W represent the height and width of the image respectively; C is the number of channels of each image; (2.1.3) The multi-layer Tucker decomposition module converts the input tensor X into a feature representation through a tensor decomposition operation, that is, the tensor X is decomposed into a core tensor G and a factor matrix using Tucker decomposition: X≈G×1U1×2U2×3U3, Among them, 1U1 is the factor matrix of H and the local spatial features of the retained image; 2U2 is the factor matrix of W and the local spatial features of the retained image; 3U3 is the factor matrix of the channel dimension; (2.1.4) After each decomposition layer, the ReLU activation function and batch normalization operation are applied to enhance the feature representation capability, accelerate training convergence, and output the core tensor G.

7. The method according to claim 6, characterized in that: Step (2.2) is calculated based on Transformer to obtain the global feature tensor F. The implementation steps are as follows: (2.2.1) The core tensor G generated by the tensor decomposition network: Among them, r1 represents the spatial dimension after H compression; r2 represents the spatial dimension after W compression; r3 is the channel dimension; (2.2.2) Directly regard r3 as the number of tensor blocks, take each tensor block as the input of a self-attention unit, and use the Transformer's multi-head self-attention mechanism MHSA to operate on the tensor block to obtain the global feature tensor F: Where d is the number of output channels of Transformer.

8. The method according to claim 7, characterized in that: Step (2.2.2) uses Transformer's multi-head self-attention mechanism MHSA to operate on the tensor block, where the calculation formula for each attention head is as follows: Where Q, K, and V represent the matrices of query, key, and value, respectively; d k Represents the normalization factor of the attention weight.

9. The method according to claim 1, characterized in that: Step (3) iteratively trains the tensor Transformer model, specifically by minimizing the predicted change detection map Y and the actual labeled change detection map Y through supervised learning. true The error between is optimized; the steps are as follows: (3.1) Load the training dataset, including multi-temporal remote sensing image pairs [X1, X2] and corresponding change detection labels Y true ; (3.2) The input image pair [X1, X2] is passed through the tensor neural network and Transformer in turn, and finally a change detection map P is generated through the classifier; (3.3) Through the cross entropy loss function L BCE Calculate the predicted graph P and the labeled graph Y true The error between; the cross entropy loss function L BCE The binary cross entropy loss BCE is used as the objective function, which is defined as follows: Where N is the number of pixels; Y true (i) is the annotated change detection label, which takes the value of 0 or 1; P(i) is the output probability of the model, which takes the value between [0,1]; (3.4) Back-propagate the error to the weights of the tensor neural network and Transformer, and use the back-propagation algorithm to adjust the parameters of each layer; (3.5) Use the Adam optimizer or SGD optimizer to update the trainable parameters of the network and minimize the loss function.

10. The method according to claim 1, characterized in that: The binary change detection map in step (4) is obtained by the following operation: (4.1) Use the convolution classifier to directly convert the feature tensor F into a change probability map P of the same size as the input image: P=σ(Conv 1×1 (F))∈R H×W , Among them, Conv 1×1 is a 1×1 convolution operation, which is used to compress the number of channels d to 1; σ is the Sigmoid activation function, which is used to map the output of the convolution to the interval [0,1]; (4.2) The probability map P is converted into a binary change detection map Y using a binarization operation: Among them, P(x,y) represents the change probability of pixel (x,y) in the probability map, τ is a threshold; 1 represents the changed area, and 0 represents the unchanged area.

Citation Information

Patent Citations

  • Remote sensing image segmentation method for enhancing global features based on matrix decomposition

    CN116310339A

  • Remote sensing image change detection method combining convolutional neural network and Transform

    CN116402766A

  • Remote sensing image scene classification method based on convolutional neural network and multilayer perceptron

    CN116563683A

  • Remote sensing image change detection method based on hierarchical cross-scale global feature fusion deep network

    CN117853897A