A remote sensing image change detection method, system, device and medium based on transformer architecture search
Through the method based on transformer architecture search, the problems of long training time of remote sensing image change detection model and large computing resource consumption are solved, and efficient detection adapted to remote sensing image changes are achieved, reducing the calculation cost and missed detection rate.
Patent Information
- Application Number
- CN202310698148.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-13
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-06-13
AI Technical Summary
The existing remote sensing image change detection model has a long training time and consumes a lot of computing resources. The rectangular structure of traditional convolution kernels is difficult to achieve remote pixel interaction. Neural architecture search NAS has problems with double-layer optimization and huge search space, resulting in unstable training and high computing costs.
Using a transformer architecture search method, we design a decomposed architecture search framework to automatically find the architecture settings suitable for the current detection task, and achieve stable and fast transformer architecture search.
It reduces the computational cost of training and inference, adapts to the rapid changes in remote sensing images, improves the accuracy and efficiency of change detection, and reduces the false detection and missed detection rates.
Smart Images

Figure CN117011697B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly relates to a remote sensing image change detection method, system, device and medium based on transformer architecture search. Background Art
[0002] Remote sensing technology is a method for indirectly obtaining target state information and is currently widely used in various surface monitoring tasks. Remote sensing technology observes the ground from a distance, uses the propagation and reception of electromagnetic waves to sense a certain feature of the observed target and analyzes and applies it, and has the advantages of real-time transmission, fast processing, instantaneous imaging, large monitoring range and little influence by the ground. At present, the images obtained by remote sensing technology mainly include high-resolution optical remote sensing images, hyperspectral remote sensing images, and polarimetric SAR remote sensing images, and these images can all obtain rich surface information. In the fields of land cover detection, resource exploration, and urban expansion, the government, environmental protection departments, and research institutions can use existing remote sensing image change detection algorithms to obtain the change information of the current regional surface, so as to provide important guidance for urban planning, disaster prevention, environmental protection, etc.
[0003] Among them, remote sensing image change detection is one of the important remote sensing tasks. Change detection is to identify the changes in two different images. For the change detection of remote sensing images, it is first necessary to collect images, that is, a series of images at the same location obtained at different times. Since the images may be affected by factors such as the atmosphere and noise, correction is required to eliminate these effects. Then, a change detection algorithm is applied to identify the significant changes in the images, and a segmentation algorithm is applied to divide the significant changes into change regions and unchanged regions. With the increase in the availability of high-resolution satellite images and the improvement of remote sensing imaging systems, remote sensing data is developing in the direction of high resolution, multi-sensor and multi-temporal. Multi-temporal remote sensing data such as satellite images and aerial images can provide rich information. However, the rich data sources and the requirements for change accuracy in actual production activities also pose severe challenges to change detection technology.
[0004] Remote sensing image change detection methods can be divided into traditional change detection methods and deep learning-based methods. Traditional remote sensing image change detection methods can be divided into three categories: methods based on image arithmetic, methods based on image transformation, and post-classification methods. Methods based on image arithmetic directly classify each pixel in the image. These methods usually use methods based on algebraic operations to calculate the pixel values of the image and classify the pixels in the difference map by setting appropriate thresholds; Methods based on image transformation suppress relevant information and highlight change information through the statistics and transformation of the image. This change detection method involves using transformations on image pixels to detect changes in the image. Its main idea is to apply the principal component analysis (PCA) method to analyze and transform two-temporal images; The post-classification method is a commonly used supervised method. It classifies two-temporal images by using a classifier. It can compare the obtained feature maps according to the corresponding positions and then obtain the changed areas. The post-classification method can obtain better change detection results, but the whole process is relatively complex. At the same time, they are very sensitive to the classification results and largely depend on the performance of the selected classifier.
[0005] In recent years, due to the excellent representation learning ability of deep learning technology, the substantial improvement in the performance of image processors, the continuous improvement of earth observation technology, and the emergence of massive remote sensing data, deep learning technology has been widely applied to remote sensing image change detection tasks. Among them, because the convolutional neural network (CNN) has achieved excellent results in computer vision tasks, CNN has attracted extensive attention in the remote sensing field. CNN has strong robustness, can extract deep features of high-resolution images, obtain rich semantic information and clear texture information. The ability of CNN to extract deep features helps to detect the changed areas and obtain more complete and obvious change information. Deep learning methods can be divided into encoder-decoder-based networks and classical image classification-based networks according to the different backbone networks constructed. Currently, these deep learning change detection algorithms generally implement change detection tasks on remote sensing images based on manually built networks. However, manually designing a neural network is a very time-consuming and cumbersome process, which largely depends on expert domain knowledge. Therefore, many studies adopt neural architecture search (NAS) to automatically search for network architectures, and the detection effect is better than that of manually built networks. However, since the NAS strategy involves double-layer optimization, it usually leads to unstable training and expensive computational costs. And NAS also involves the type of architecture search, such as the transformer architecture search technology, which is gradually popular in the field of computer vision, but its research in the field of remote sensing change detection is still blank. These are the problems in current remote sensing image change detection.
[0006] The architecture search based on transformers was first proposed by Su et al. in June 2021, (ViTAS: Vision Transformer Architecture Search, Xiu Su, Shan You, Jiyang Xie, Mingkai Zheng, Fei Wang, Chen Qian, Changshui Zhang, Xiaogang Wang, Chang Xu, ArXiv-Computer Science-Computer Vision and Pattern Recognition) who tried to apply NAS methods to transformer search. However, during the search process, the entire training process was unstable for the supernetwork. Therefore, they proposed a cyclic weight sharing mechanism for each input token, enabling each channel to learn all candidate frameworks more evenly. In July 2021, Chen et al. proposed AutoFormer, (AutoFormer: Searching Transformers for Visual Recognition, M. Chen, H. Peng, J. Fu and H. Ling, 2021 IEEE / CVF International Conference on Computer Vision (ICCV)) AutoFormer uses a one-shot approach for architecture search. When training the supernetwork, AutoFormer updates the weights of different regions in the same layer using the weight entanglement strategy. Thanks to this strategy, the trained supernetwork can effectively train the remaining subnets. In September 2021, Liao et al. proposed ViT-ResNAS, (Searching for Efficient Multi-Stage Vision Transformers, Yi-Lun Liao, Sertac Karaman, Vivienne Sze, ArXiv-Computer Science -Computer Vision and Pattern Recognition) ViT-ResNAS proposed residual space reduction to reduce the sequence length in the deeper layers. When reducing the length, skip connections are added to improve performance and stabilize the training of deeper networks. Secondly, weight sharing NAS with multi-architecture sampling was also proposed. Finally, to train the supernetwork more effectively, forward and backward channels were proposed to sample and train multiple subnets.In March 2022, Zhang et al. studied the structural characteristics of transformers and convolutions and proposed a method for searching the architecture of a transformer with a convolutional structure (VTCAS) (Vision Transformer with Convolutions Architecture Search, Haichao Zhang, Kuangrong Hao, Witold Pedrycz, Lei Gao, Xuesong Tang, Bing Wei, ArXiv - Computer Science - Computer Vision and Pattern Recognition). The high - performance backbone network searched by VTCAS introduces the features of convolutional neural networks into the transformer architecture while maintaining the advantages of the multi - head attention mechanism. The backbone network based on search blocks can extract feature maps of different scales. The topological structure based on multi - head self - attention and CNN adaptively correlates the relational features of pixels with the multi - scale features of objects, enhancing the robustness of the neural network for object recognition.
[0007] However, the existing technologies have the following problems:
[0008] 1. The training of traditional change detection models requires a large amount of time and computing resources. The remote sensing image data volume is huge and changing rapidly, and it is difficult for traditional change detection models to adapt to this change;
[0009] 2. The rectangular structure of traditional convolution kernels limits their receptive fields to local contexts, making it difficult to achieve long - range interactions between pixels at different positions;
[0010] 3. Traditional neural architecture search NAS involves double - layer optimization and a huge search space, which can lead to unstable training and expensive computing costs. Summary of the Invention
[0011] In order to overcome the above - mentioned disadvantages of the existing technologies, the purpose of the present invention is to provide a method, system, device and medium for remote sensing image change detection based on transformer architecture search. By designing a decomposition architecture search framework and using this framework to automatically focus on and find the architecture settings suitable for the current detection task, stable and fast transformer architecture search can be achieved, which has the characteristics of strong adaptability, fast and stable search process.
[0012] In order to achieve the above purpose, the technical solutions adopted by the present invention are as follows:
[0013] A method for remote sensing image change detection based on transformer architecture search, comprising the following steps:
[0014] Step 1, set the search space: Set the search space for neural architecture search, where the search space includes search space Ω and search space Ψ;
[0015] Step 2, build the search framework: Build the search framework of the change detection network SCTN using the search space set in Step 1;
[0016] Step 3, search for the optimal structure factors: Use the factorization architecture search FAS to determine the optimal structure factor α of search space Ω and the optimal structure factor θ of search space Ψ in Step 1 * in Step 1, and use the optimal structure factor α * and the optimal structure factor θ * to replace search space Ω and search space Ψ in the search framework of the change detection network SCTN built in Step 2 respectively; * Step 4, construct the change detection network: Use the optimal structure factor α
[0017] obtained in Step 3 and the optimal structure factor θ * to construct the final change detection network SCTN; * Step 5, train the change detection network: Train the change detection network SCTN constructed in Step 4 on the training set {X
[0018] , Y tra , Y tra} until the network converges, and test the trained change detection network SCTN on the test set {X tes , Y tes} to obtain the final change result.
[0019] The specific process of Step 1 is as follows:
[0020] Step 1.1, set search space Ω: The search space Ω is the attention search space Attention SearchSpace, including spatial and channel attention modules, and the spatial and channel attention modules are different combinations α of three hierarchical orders of the spatial attention module and the channel attention module. The different combinations α are respectively:
[0021] CSA module: It means that the feature map is first input into the channel attention module and then into the spatial attention module;
[0022] SCA module: It means that the feature map is first input into the spatial attention module and then into the channel attention module;
[0023] S&C module: It means that the feature map is input into the spatial attention module and the channel attention module in parallel and then added together;
[0024] In each combined module, the weighted output feature can be calculated by the following formula:
[0025]
[0026] In the formula, α represents an operation in the search space Ω, and β α represents the trainable weight of each unit in the combination module, and F out represents the output tensor, and F in represents the input tensor, and exp{β α} represents the weight value of each operator, β represents the trainable architecture parameter, and ∑ α' exp{β α'} represents the normalization term;
[0027] Step 1.2, set the search space Ψ: The search space Ψ is a multi-scale search space Multi-scale Search Space, including a multi-scale fusion module, and the multi-scale fusion module is three different combinations θ of multi-scale modules. The different combinations θ are respectively:
[0028] CSAP module: It represents a multi-scale module that first inputs a channel attention module and then inputs a spatial attention module;
[0029] SCAP module: It represents a multi-scale module that first inputs a spatial attention module and then inputs a channel attention module;
[0030] S&CP module: It represents a multi-scale module that inputs a spatial attention module and a channel attention module in parallel.
[0031] The specific process of Step 2 is as follows:
[0032] Step 2.1, input the image and the image Concatenate the image and the image on the channel and then input them into the ResNet18 network to extract preliminary features. The image and the image represent two temporal phase images at different times i1 and i2;
[0033] Step 2.2, input the preliminary features extracted in Step 2.1 into the spatial and channel attention modules in the search space Ω set in Step 1 to obtain output features;
[0034] Step 2.3, input the output features of Step 2.2 into the multi-scale fusion module in the search space Ψ set in Step 1 to extract different scale features;
[0035] Step 2.4, concatenate the output features of Step 2.2 and the different scale features extracted in Step 2.3, and input them into the classifier to obtain the final change map.
[0036] Step 3 uses the decomposition architecture search FAS to determine the optimal structure factors, including two sub - processes: FAS1 and FAS2. Specifically, it is as follows:
[0037] Step 3.1, FAS1: Select the multi - scale fusion module combinations in the search space Ψ Fix different combinations θ of the multi - scale fusion modules in the search space Ψ as the combination Find the optimal combination α of the spatial and channel attention modules in the search space Ω * , that is, the optimal structure factor α * , and the formula is as follows:
[0038]
[0039]
[0040] In the formula, ω represents the network parameters to be trained, {X D , Y D} represents the development set, {X val , Y val} represents the validation set, λ1 represents the learning rate of the FAS1 process, represents the fixed structure factor set in the network architecture, α represents the initialized network architecture factor, represents the network architecture factor during the search process, represents the loss function The gradient of with respect to ω, α * represents an optimal factor in the network architecture. By training the network parameters ω on the development set X D , find the optimal combination α on the validation set X val ; * ;
[0041] Step 3.2, FAS2: Fix different combinations α of the spatial and channel attention modules in the search space Ω as the optimal combination α obtained in Step 3.1 * , and find the optimal multi - scale fusion module combination θ in the search space Ψ * , that is, the optimal structure factor θ * , and the formula is as follows:
[0042]
[0043]
[0044] In the formula, ω represents the network parameters to be trained, {X D , Y D} represents the development set, {X val , Y val} represents the validation set, λ2 represents the learning rate of the FAS2 process, α * represents the optimal factor of the network architecture searched by FAS1, θ represents the initialized network architecture factor, represents the network architecture factor during the search process, represents the loss function gradient of ω, θ * represents an optimal factor in the network architecture. By training the network parameters ω on the development set X D and searching for the optimal combination θ on the validation set X val ; * ;
[0045] Step 3.3: Determine the optimal combination of the spatial and channel attention modules, the S&C module, in the search space Ω according to the optimal structure factor α obtained in Step 3.1; determine the optimal combination of the multi-scale fusion module, the S&CP module, in the search space Ψ according to the optimal structure factor θ obtained in Step 3.2; * Determine the optimal combination of the spatial and channel attention modules, the S&C module, in the search space Ω according to the optimal structure factor α obtained in Step 3.1; determine the optimal combination of the multi-scale fusion module, the S&CP module, in the search space Ψ according to the optimal structure factor θ obtained in Step 3.2; * Determine the optimal combination of the multi-scale fusion module, the S&CP module, in the search space Ψ;
[0046] Step 3.4: Eliminate the search space Ω and replace it with the S&C module obtained in Step 3.3; eliminate the search space Ψ and replace it with the S&CP module obtained in Step 3.3.
[0047] The specific process of Step 4 is as follows:
[0048] Step 4.1, Input the image and the image Concatenate the images and the image in the channel dimension and input them into the ResNet18 network to extract preliminary features;
[0049] Step 4.2, Input the features extracted in Step 4.1 into the S&C module, that is, input the image into the spatial attention module and the channel attention module in parallel and then add them together to obtain the output features;
[0050] Step 4.3, Input the output features of Step 4.2 into the S&CP module to extract features of different scales through the multi-scale module;
[0051] Step 4.4, Concatenate the output features of Step 4.2 and the features of different scales extracted in Step 4.3 in the channel dimension and input them into the classifier to obtain the final change map.
[0052] The spatial attention module in Step 1.1 is implemented by a two-dimensional convolutional kernel of size 1×1. The specific construction method is as follows:
[0053] (1) Calculate query, key, and value, and the formulas are as follows:
[0054]
[0055]
[0056]
[0057] In the formula, g(·) represents the use of a two-dimensional 1×1 convolution, and W Q , W K and W V respectively represent the matrices randomly initialized for query, key, and value. c, w, h, and c' respectively represent the number of channels, height, width of the input feature Figure X and the number of channels of Q and K;
[0058] (2) Reshape the sizes of Q and K into c'×hw for subsequent matrix multiplication. The final spatial attention output is as follows:
[0059] Output = V·softmax(Q T K)
[0060] In the formula, the softmax function represents normalizing the attention weights. Q T represents the transpose of the query matrix Q, and Output represents the output feature. Multiply the normalized attention weight matrix softmax(Q T K) and V to obtain the output feature Output.
[0061] The construction method of the channel attention module in step 1.1 is as follows:
[0062] (1) Reshape the input feature into where N = H×W, then perform matrix multiplication on the transpose of F and F, and obtain the channel attention map through the softmax function The formula is as follows:
[0063]
[0064] In the formula, represents the influence of the i-th channel on the j-th channel. The stronger the correlation between the two channels, the larger the value; F represents the input feature, F i represents the i-th channel feature after the input feature is Reshaped, and F j represents the j-th channel feature after the input feature is Reshaped;
[0065] (2) Multiply the result obtained in step (1) by the scale parameter λ and perform an element-wise summation operation on F to obtain the final channel attention output. The formula is as follows:
[0066]
[0067] In the formula, λ represents the scale parameter, which is initialized to 0 and gradually learns to allocate more weights; represents the influence of the i-th channel on the j-th channel, F i represents the i-th channel feature after reshaping the input feature, F j represents the j-th channel feature after reshaping the input feature, represents the j-th channel feature of the output.
[0068] A remote sensing image change detection system based on transformer architecture search, comprising:
[0069] Search space setting module: Set the search space for neural architecture search, where the search space includes search space Ω and search space Ψ;
[0070] Search framework building module: Use the search space to build the search framework of the change detection network SCTN;
[0071] Optimal structure factor search module: Use the factorized architecture search FAS to determine the optimal structure factor α of search space Ω * and the optimal structure factor θ of search space Ψ * , and use the optimal structure factor α * and the optimal structure factor θ * to replace search space Ω and search space Ψ in the search framework of the change detection network SCTN respectively;
[0072] Change detection network construction module: Used to construct the final change detection network SCTN;
[0073] Change detection network training module: Train the change detection network SCTN constructed by the change detection network construction module on the training set {X tra , Y tra} until the network converges, and test the trained change detection network SCTN on the test set {X tes , Y tes} to obtain the final change result.
[0074] A remote sensing image change detection device based on transformer architecture search, comprising:
[0075] Memory: Used to store the computer program for implementing the remote sensing image change detection method based on transformer architecture search;
[0076] Processor: When executing the computer program, it implements a remote sensing image change detection method based on Transformer architecture search.
[0077] A computer-readable storage medium, comprising:
[0078] The computer-readable storage medium stores a computer program, which can implement a remote sensing image change detection method based on Transformer architecture search when executed by a processor.
[0079] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0080] 1. The present invention uses Transformer-based architecture search for remote sensing image change detection. Compared with existing change detection models, the Transformer-based architecture search method can automatically find an efficient model architecture, thereby reducing the computational cost of training and inference; at the same time, the amount of remote sensing image data is huge and changing rapidly, and existing change detection models are difficult to adapt to this change, while the Transformer-based architecture search method can adaptively learn the patterns in the change process, so as to better meet the task requirements of remote sensing image change detection.
[0081] 2. The spatial attention mechanism adopted by the present invention can adaptively learn the spatial distribution of features, overcome the constraints of the convolutional kernel network structure, and compared with the prior art, reduce the high computational cost of processing large spatial size images.
[0082] 3. The channel attention mechanism adopted by the present invention can adaptively learn the channel importance of features, thereby strengthening the channels with important information and weakening the channels with noise or useless information compared with the prior art.
[0083] 4. The present invention adopts a multi-scale fusion module to fuse feature maps of different scales to obtain more comprehensive change information. Compared with the prior art, multi-scale fusion makes full use of change information of different scales, improves the accuracy of change detection; reduces the false detection rate caused by image scale changes, thereby improving the robustness of change detection; at the same time, by fusing feature maps of different scales, the computational complexity can be reduced to a lower scale, improving the efficiency of change detection.
[0084] 5. The present invention designs a decomposition architecture search strategy FAS. Using FAS, the search space can be decomposed into two independent search spaces, the network settings of traditional NAS can be decomposed into two network structure system factors, and then constraints are imposed on the two search spaces respectively to search for the optimal network structure system factors. Compared with the prior art, FAS decouples the search and training stages and has the characteristics of a faster and more stable search process.
[0085] In summary, compared with the prior art, the present invention designs a decomposition architecture search framework, which automatically focuses on and finds out the architecture settings suitable for the current detection task by using this framework, so as to realize stable and fast transformer architecture search, and has the characteristics of strong adaptability, fast and stable search process. Brief Description of the Drawings
[0086] Figure 1 It is a flowchart of the method of the present invention.
[0087] Figure 2 It is the overall flowchart of the decomposition architecture search of the present invention.
[0088] Figure 3 It is the architecture diagram of the spatial attention mechanism in the present invention.
[0089] Figure 4 It is the architecture diagram of the channel attention mechanism in the present invention.
[0090] Figure 5 It is the architecture diagram of the multi-scale search space of the attention module search space in the present invention.
[0091] Figure 6 It is the architecture diagram of the change detection network SCTN of the present invention.
[0092] Figure 7 It is the detailed architecture diagram of the change detection network SCTN of the present invention.
[0093] Figure 8 It is the change detection result diagram of the change detection network SCTN of the present invention and its comparison algorithm on the LEVIR-CD dataset.
[0094] Figure 9 It is the change detection result diagram of the change detection network SCTN of the present invention and its comparison algorithm on the CCD dataset.
[0095] Figure 10 It is the ablation experiment result diagram of the change detection network SCTN of the present invention on the LEVIR-CD.
[0096] Figure 11 It is the ablation experiment result diagram of the change detection network SCTN of the present invention on the CCD. Detailed Embodiments
[0097] The technical solutions of the present invention will be described in detail below with reference to the drawings and simulations.
[0098] Refer to Figure 1 and Figure 2 , a remote sensing image change detection method based on transformer architecture search, includes the following steps:
[0099] Step 1, set the search space: Set the search space for neural architecture search, where the search space includes search space Ω and search space Ψ;
[0100] Step 2, build the search framework: Build the search framework of the change detection network SCTN using the search space set in Step 1;
[0101] Step 3, search for the optimal structure factors: Use the factorization architecture search FAS to determine the optimal structure factor α of search space Ω in Step 1 * and the optimal structure factor θ of search space Ψ * , and use the optimal structure factor α * and the optimal structure factor θ * to replace search space Ω and search space Ψ in the search framework of the change detection network SCTN built in Step 2, respectively;
[0102] Step 4, construct the change detection network: Use the optimal structure factor α obtained in Step 3 * and the optimal structure factor θ * to construct the final change detection network SCTN;
[0103] Step 5, train the change detection network: Train the change detection network SCTN constructed in Step 4 on the training set {X tra , Y tra} until the network converges, and test the trained change detection network SCTN on the test set {X tes , Y tes} to obtain the final change result.
[0104] See Figure 3 , Figure 4 and Figure 5 , in Step 1, set the search space for neural architecture search, specifically:
[0105] Define the spatial attention module. In the present invention, a two-dimensional 1×1 convolutional kernel is used to construct the spatial attention module. The spatial attention mechanism is as Figure 3 shown. First, calculate the query, key, and value, and the formulas are as follows:
[0106]
[0107]
[0108]
[0109] where g(·) represents using a two-dimensional 1×1 convolution; W Q , W K and WV represent matrices randomly initialized for query, key, and value respectively; c, w, h, and c' represent the number of channels of the input feature Figure X , height, width, and the number of channels of Q and K respectively.
[0110] Then reshape the sizes of Q and K to c'×hw for subsequent matrix multiplication, and the final spatial attention output is as follows:
[0111] Output = V·softmax(Q T K)
[0112] where the softmax function represents normalizing the attention weights, Q T represents the transpose of the query matrix Q, and Output represents the output feature; multiply the normalized attention weight matrix softmax(Q T K) and V to obtain the output feature Output.
[0113] Each high-level feature channel mapping can be regarded as a response to the ground object. By using the correlation between channel mappings, the mutually dependent feature mappings can be enhanced, and the feature representation with specified semantics can be improved, so as to better distinguish changes. Therefore, the present invention constructs a channel attention module for constructing the relationship between channels. As Figure 4 shown, compared with the spatial attention module, no convolutional operation is used in the channel attention module to obtain new features. The input feature is reshaped to where N = H×W. Then multiply the transpose of F by F, and obtain the channel attention map through the softmax function The formula is as follows:
[0114]
[0115] where, can be used to measure the influence of the i-th channel on the j-th channel. Similarly, the stronger the correlation between two channels, the larger the value of. Reshape F to and multiply it by F x to get the result. Finally, multiply the result of the previous step by the scale parameter λ and perform an element-wise summation operation on F to obtain the final output, and the formula is as follows:
[0116]
[0117] Among them, λ is initialized to 0 and gradually learns to assign more weights. From the above formula, it can be obtained that the final feature of each channel is the result of the weighted sum of the features of all channels and the original features, which models the long-distance semantic dependencies between feature maps. It enhances the recognizability of features and highlights the feature representation of the changing regions.
[0118] The present invention introduces a decomposition architecture search FAS to search for network architectures, decomposing the search space into two search subspaces Ω and Ψ, denoted as Attention Search Space and Multi-scale SearchSpace respectively, as Figure 5 shown. The search space Ω contains three hierarchical orders of spatial and channel attention blocks. Among them, CSA means that the feature map first passes through the channel attention module and then through the spatial attention module, SCA means that the feature map first passes through the spatial attention module and then through the channel attention module, and S&C means that the feature map passes through the channel attention module and the spatial attention module in parallel and then adds them. In each combined module, the weighted output feature can be calculated by the following formula:
[0119]
[0120] Among them, α represents the operation in Ω, and β α represents the trainable weight of each unit in the combined module. Then, the output tensor F of the hierarchical unit in the transformer block is calculated by the weighted sum of all operators in the transformer block through the intermediate tensor out . The weight value of each operator is equal to exp{β α}, where β is a trainable architecture parameter, and then divided by a normalization term ∑ α' exp{β α'} to compare multiple modules in this way.
[0121] The search space Ψ contains different combinations of multi-scale modules. Among them, CSAP represents a multi-scale module that first passes through the channel attention module and then through the spatial attention module, SCAP represents a multi-scale module that first passes through the spatial attention module and then through the channel attention module, and S&CP represents a multi-scale module that inputs the spatial and channel attention modules in parallel. The present invention limits the block-level search of FAS to only operate on 3 candidate operators shown in the Ψ space.
[0122] See Figure 6 , Step 2 of building the search framework is specifically as follows:
[0123] Step 2.1, input image and image Input image and image Input them in series on the channel into the ResNet18 network to extract preliminary features, where the image and the image represent two-phase images at different times i1 and i2;
[0124] Step 2.2, input the extracted preliminary features into the spatial and channel attention modules in the search space Ω set in Step 1 to obtain output features;
[0125] Step 2.3, input the output features into the multi-scale fusion module in the search space Ψ set in Step 1 to extract features at different scales; the selection of these two types of modules is determined by using FAS to search for two architecture factors α and θ in the search spaces Ω and Ψ respectively, where α represents the combination of spatial and channel attention blocks, and θ represents the combination of multi-scale modules;
[0126] Step 2.4, concatenate the output features in Step 2.2 and the features at different scales extracted in Step 2.3, and input them into the classifier to obtain the final change map.
[0127] Step 3, searching for the optimal architecture factor specifically:
[0128] The network can be expressed as Net(X, α, θ, ω). FAS searches for the optimal architecture factor in one search space by fixing the architecture factor in the other space, so effective and efficient searches can be performed in each space. FAS is divided into two sub-processes, FAS1 and FAS2, and both processes are implemented on the development set {X D , Y D} and the validation set {X val , Y val}.
[0129] Step 3.1, FAS1: Select the combination of multi-scale fusion modules in the search space Ψ Fix different combinations θ of the multi-scale fusion modules in the search space Ψ as the combination Search for the optimal combination α of spatial and channel attention modules in the search space Ω * , that is, the optimal architecture factor α * , the formula is as follows:
[0130]
[0131]
[0132] Among them, calculate the loss function of the development set {X D , Y D} on the current network structure Then use gradient optimization to train and update the network parameters ω, and then according to the validation set {X val,Y val The loss function on the current network structure to determine the optimal structure factor α of the current network * 。
[0133] In the formula, ω represents the network parameters to be trained, {X D ,Y D} represents the development set, {X val ,Y val} represents the validation set, λ1 represents the learning rate of the FAS1 process, represents the fixed structure factor set in the network architecture, α represents the initialized network architecture factor, represents the network architecture factor during the search process, represents the loss function the gradient of with respect to ω, α * represents an optimal factor in the network architecture. By training the network parameters ω on the development set X D , find the optimal combination α on the validation set X val 。 * 。
[0134] Step 3.2, FAS2: Fix the different combinations α of the spatial and channel attention modules in the search space Ω to the optimal combination α obtained in Step 3.1 * , and find the optimal multi-scale fusion module combination θ * in the search space Ψ, that is, the optimal structure factor θ * , the formula is as follows:
[0135]
[0136]
[0137] Among them, calculate the loss function of the current network structure on the development set {X D ,Y D} Then use gradient optimization to train and update the network parameters ω, and then according to the loss function of the current network structure on the validation set {X val ,Y val} to determine the optimal structure factor θ of the current network * 。
[0138] In the formula, ω represents the network parameters to be trained, {X D ,Y D} represents the development set, {X val ,Y val} represents the validation set, λ2 represents the learning rate of the FAS2 process, α *Let $\alpha$ denote the optimal factor of the network architecture searched by FAS1, and $\theta$ denote the initialized network architecture factor. Denote the network architecture factor during the search process. Denote the loss function The gradient of $\omega$, $\theta$ * Represents an optimal factor in the network architecture. By training the network parameter $\omega$ on the development set $X$ D and searching for the optimal combination $\theta$ on the validation set $X$ val . * .
[0139] Step 3.3: The change detection network SCTN is similar to the search framework of the change detection network SCTN in Step 2. Remove the original Attention Search Space and Multi-scale Search Space, and use the searched optimal structure factors $\alpha$ * and $\theta$ * to replace them, that is, the S&C module and the S&CP module.
[0140] See Figure 7 , Step 4. The specific construction of the change detection network is as follows:
[0141] Step 4.1: Input the image and the image Concatenate the images and the image in channels and input them into the ResNet18 network to extract preliminary features;
[0142] Step 4.2: Input the features extracted in Step 4.1 into the S&C module, that is, input the image into the spatial attention module and the channel attention module in parallel and then add them to obtain the output features;
[0143] Step 4.3: Input the output features of Step 4.2 into the S&CP module to extract features of different scales through the multi-scale module;
[0144] Step 4.4: Concatenate the output features of Step 4.2 and the features of different scales extracted in Step 4.3 in channels and input them into the Classifier classifier to obtain the final change map.
[0145] Step 5. The specific training of the change detection network is as follows:
[0146] Train the change detection network SCTN on the training set $\{X$ tra , $Y$ tra} until the network converges. Finally, on the test set $\{X$ tes , $Y$ tes} Test the trained change detection network SCTN to obtain the final change result.
[0147] A remote sensing image change detection system based on transformer architecture search, comprising:
[0148] Search space setting module: Set the search space for neural architecture search, where the search space includes search space Ω and search space Ψ;
[0149] Search framework building module: Use the search space to build the search framework of the change detection network SCTN;
[0150] Optimal structure factor search module: Use the decomposition architecture search FAS to determine the optimal structure factor α of search space Ω * and the optimal structure factor θ of search space Ψ * , and use the optimal structure factor α * and the optimal structure factor θ * to replace search space Ω and search space Ψ in the search framework of the change detection network SCTN respectively;
[0151] Change detection network construction module: Used to construct the final change detection network SCTN;
[0152] Change detection network training module: Train the change detection network SCTN constructed by the change detection network construction module on the training set {X tra , Y tra} until the network converges, and test the trained change detection network SCTN on the test set {X tes , Y tes} to obtain the final change result.
[0153] A remote sensing image change detection device based on transformer architecture search, comprising:
[0154] Memory: Used to store the computer program for implementing the remote sensing image change detection method based on transformer architecture search;
[0155] Processor: Used to implement the remote sensing image change detection method based on transformer architecture search when executing the computer program.
[0156] A computer-readable storage medium, comprising:
[0157] The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement a remote sensing image change detection method based on transformer architecture search.
[0158] The application effect of the present invention will be described in detail below in conjunction with simulations.
[0159] 1. Simulation conditions
[0160] The simulation experiments of the present invention are all carried out on a workstation with an NVIDIA RTX 3090 graphics card based on the Pytorch deep learning framework, and six of the most popular deep learning change detection networks are selected for comparative experiments.
[0161] 2. Simulation parameter settings: For the change detection network SCTN proposed by the present invention, in the structure search stage, as Figure 5 shown, two options are set for each level of each combination in the search space Ω. Then the warm-up mode is used for the first 15 epochs, and a total of 100 epochs are searched. During this period, the Adam optimizer is used. For the LEVIR-CD and CCD datasets, the learning rate is set to 0.01 in the search stage. The momentum decay and weight decay are set to 0.9 and 0.0005 respectively. In the network training stage, the present invention adopts the Figure 7 optimal architecture settings shown. In this stage, the present invention uses the Stochastic Gradient Descent (SGD) optimizer to train the proposed network. The SGD hyperparameter momentum is set to 0.99, the learning rate is set to 0.01, the weight decay is set to 0.0005, and the linear decay is set to 0. The present invention trains the entire network for 200 epochs until the network converges. During training and testing, the batchsize is set to 8. The selected comparison algorithms are SNUNet, BiT, DTCDSCN, IFN, STANet, and MSPSNet. The index settings of these six comparison algorithms are the same as the corresponding original text.
[0162] 3. Simulation content
[0163] Simulation 1: To verify the effectiveness of the change detection network SCTN proposed by the present invention, six popular deep learning change detection algorithms are selected as comparison algorithms, namely SNUNet, BiT, DTCDSCN, IFN, STANet, and MSPSNet. The detection results of these seven algorithms on the LEVIR-CD dataset and the CCD dataset are obtained in this experiment, and these methods are evaluated and compared based on numerical and visualization results.
[0164] On the LEVIR-CD dataset, Table 1 shows the numerical results of these methods. Figure 8The visualization effect is shown. In the table, IFN, SNUNet, STANet, BiT, MSPSNet, DTCDSCN, and SCTN respectively represent different algorithms. Pre(%) represents precision, Rec(%) represents recall rate, F1(%) represents an index that comprehensively considers Pre and Rec, IoU(%) represents the intersection over union, which is used to evaluate the overlapping degree between the predicted change region and the true change region, and OA(%) represents the overall precision.
[0165] Table 1 Index statistical results of the present invention and other comparative algorithms on the LEVIR-CD dataset
[0166]
[0167] As can be seen from Table 1 and Figure 8 it can be known that the edge of the change graph of IFN is relatively rough and there are certain misdetections. Although it can detect some change regions, it may not be able to detect the complete edge. Compared with IFN, the detection results of the present invention and SNUNet are more complete, and there are fewer misdetections and missed detections. Among them, STANet embeds a spatio-temporal attention module in the feature extractor, emphasizing the features of the change region, so the precision value is the highest. And BiT avoids misdetections and missed detections through a transformer module that can emphasize the connection between high-level semantic features, but it is not sensitive to the edge features of the change object, which may cause the edges, corners, and objects to be detected as a whole. SNUNet adopts dense connections, multi-level feature integration, and attention methods, and has good detection performance. Although DTCDSCN can obtain good change detection results, there are still problems such as being easily detected as one target when the background targets are too close, resulting in misdetections. Although the present invention does not achieve the best index results, this method has certain effectiveness in multi-scale feature extraction and aggregation. Generally speaking, different methods have their own advantages and limitations, and the specific choice of which method depends on the specific application scenario and requirements.
[0168] On the CCD dataset, the target objects are sparse and small, which makes the detection results vulnerable to many factors, such as noise. Due to the small size of the targets, some algorithms may have missed detections and misdetections. Table 2 shows the numerical results of all algorithms on the CCD dataset, Figure 8 and gives their visualization effects.
[0169] Table 2 Index statistical results of the present invention and other comparative algorithms on the CDD dataset
[0170]
[0171] As can be seen from Table 2, the present invention performs best, achieving the highest scores in terms of Rec (95.21%), F1-score (96.10%), IoU (92.49%), and OA (99.09%). Compared with the other five networks, the detection results of the present invention have more complete edges and fewer missed and false detections. Through Figure 8 it can be seen that IFN can obtain the change maps of most large target change regions, but it does not perform well in capturing small targets. SNUNet can obtain the entire change region in most cases, but it is not sensitive to some edge information. In addition, IFN and SNUNet may miss and misdetect small target objects. BiT can obtain complete change maps, but there are still problems of false detection and missed detection. Its detection performance for irregular objects is not good. STANet enhances the depth feature extraction ability by using the proposed spatio-temporal attention module to emphasize the change region, and the network performance has been significantly improved. However, when capturing the edge information in large-size targets, it is not sensitive to the weak changes in edge information. Although DTCDSCN can obtain a good F1-score, there are still problems of missed detection in its detection of small targets. When DTCDSCN detects large targets, there are problems such as edge information loss and detection errors at the boundaries of large targets. The present invention can almost capture the complete variation region and provide fewer error and omission regions. It has the highest F1 score, can identify irregular objects, and obtain complete change maps of large and small objects. Generally speaking, the present invention performs best on the CCD dataset and has the fewest problems of false detection and missed detection.
[0172] Simulation 2. This experiment belongs to an ablation simulation experiment. The effects of different combinations of the spatial and channel attention modules of the present invention were tested on the CCD dataset and the LEVIR-CD dataset, so as to verify that the proposed network architecture of the present invention is the optimal framework.
[0173] In this simulation experiment, the SCTN network proposed by the present invention was used as the benchmark network. The one that sequentially inputs the feature map into the spatial and channel attention modules is called SCATN, and the one that sequentially inputs the feature map into the channel and spatial attention modules is called CSATN. The ablation experiments of SCATN and CSATN as the benchmark network were used to verify the effectiveness of the present invention. The numerical indicators of the ablation experiments on the two datasets are shown in Tables 3 and 4.
[0174] Table 3 Statistical results of the indicators of the ablation simulation experiment of the present invention on the LEVIR-CD dataset
[0175]
[0176]
[0177] As shown in Table 3, the Rec value, F1 value, and IoU value of the present invention have obtained the best results in the LEVIR-CD dataset. Compared with SCATN, they have increased by 2.52%, 1.34%, and 2.18% respectively. Compared with CSATN, they have increased by 2.12%, 0.95%, and 1.55% respectively. This shows that the design of the present invention can obtain richer image features, thus obtaining better change detection results. As Figure 10 shown, the present invention can obtain a more complete change detection result, and there are fewer missed and misdetected targets.
[0178] Table 4 Statistical results of the ablation simulation experiment of the present invention on the CCD dataset
[0179]
[0180] As shown in Table 4, the Rec value, F1 value, and IoU value of the present invention have obtained the best results in the LEVIR-CD dataset. Compared with SCATN, they have increased by 0.61%, 0.14%, and 0.96% respectively. Compared with CSATN, they have increased by 0.54%, 0.25%, and 0.73% respectively. As Figure 11 shown, the present invention shows better results for small targets and edge targets, and can obtain complete change information for large targets, and the occurrences of missed and misdetected are much less than the remaining two structures. The present invention has obtained the best results on two of the most popular datasets, proving the effectiveness of the present invention obtained through architecture search.
[0181] In summary, compared with several currently popular change detection methods, the present invention shows better performance indicators on the LEVIR-CD and CCD datasets, and the present invention can obtain a more complete change map. The ablation experiment shows the effectiveness of attention modules of different scales.
[0182] Remote sensing change detection is to detect changes in the surface or ground objects by comparing multi-temporal remote sensing images. This technology has broad application prospects in various fields. The following are some of the main application prospects:
[0183] 1. Monitoring of land use and land cover change
[0184] Monitoring of farmland expansion: By comparing multi-temporal remote sensing images, the change in farmland area can be detected, helping agricultural managers understand the trend of land use to formulate appropriate agricultural policies and plans.
[0185] Monitoring of urbanization: Remote sensing change detection can be used to monitor urban expansion and construction activities, providing basic data for urban planning and land management, and helping to optimize land use and promote sustainable urban development.
[0186] Forest Change Monitoring: Through change detection of remote sensing data, deforestation, forest fires, and vegetation degradation can be monitored, providing timely information support for forest resource protection and management.
[0187] 2. Disaster Monitoring and Assessment
[0188] Flood Monitoring: Remote sensing change detection can help monitor surface changes after floods, including flood extent and changes in land types affected by floods, providing important information for flood risk assessment and rescue operations.
[0189] Earthquake and Geological Disaster Monitoring: Through change detection of remote sensing images, geological disasters such as surface ruptures, landslides, and rock layer movements caused by earthquakes can be quickly identified, contributing to post-disaster assessment and risk management.
[0190] Fire Monitoring: Remote sensing change detection can be used to monitor the spread of forest fires and changes in burned areas, providing important data support for fire monitoring and forest management.
[0191] 3. Environmental Monitoring and Ecosystem Assessment
[0192] Water Pollution Monitoring: Through remote sensing change detection, changes in water body color, suspended solid content, and algal blooms can be detected, helping to monitor and assess the degree of water pollution to support water resource management and protection.
[0193] Wetland Degradation Monitoring: Through change detection of remote sensing data, wetland reduction, wetland vegetation degradation, and water body changes can be monitored, providing key information for wetland protection and ecosystem restoration.
[0194] Vegetation Cover Change Monitoring: By comparing multi-temporal remote sensing images, changes in vegetation cover can be monitored, such as forest degradation, vegetation stress, and vegetation restoration, providing data support for ecosystem assessment and biodiversity protection.
[0195] 4. Urban Planning and Land Management
[0196] Urban Expansion Monitoring: Through remote sensing change detection, urban expansion and changes in newly built areas can be monitored, providing a decision-making basis for urban planning and land management, and contributing to the rational planning of urban development and the optimization of land use.
[0197] Building Renewal Monitoring: By comparing multi-temporal remote sensing images, changes in buildings can be detected, including building construction, demolition, and renovation, providing data support for urban renewal and infrastructure planning.
[0198] Land Use Change Monitoring: Remote sensing change detection can be used to monitor changes in land use types, such as the conversion of farmland to industrial land and the degradation of cultivated land to grassland, providing information support for land management and sustainable development.
[0199] 5. Agriculture and Food Security
[0200] Monitoring of farmland changes: Through change detection of remote sensing data, the changes in farmland can be monitored, including changes in cultivated land area, transformation of crop types, and changes in land use, providing decision-making support for agricultural management and food security.
[0201] Monitoring of crop growth: By comparing multi-temporal remote sensing images, the growth status, disaster-affected conditions, and yield estimation of crops can be monitored, providing data support for agricultural production management and farmers' decision-making.
[0202] 6. Climate Change Research
[0203] Monitoring of glacier changes: Through remote sensing change detection, the retreat of glaciers and the expansion of glacial lakes can be monitored, providing important data for climate change research and glacier monitoring.
[0204] Monitoring of sea-level rise: Remote sensing change detection can be used to monitor the changes in the coastline, the disappearance of islands, and the situation of marine erosion, providing data support for the impact assessment and response measures of sea-level rise.
[0205] In summary, remote sensing change detection has broad application prospects in the fields of land use, disaster monitoring, environmental protection, urban planning, agriculture, and climate change, and can provide key information support for decision-making and sustainable development in various fields.
Claims
1. A remote sensing image change detection method based on transformer architecture search, characterized in that, It includes the following steps: Step 1, setting the search space: Set the search space for neural architecture search, where the search space includes search space Ω and search space Ψ; Step 2, building the search framework: Build the search framework of the change detection network SCTN using the search space set in Step 1; Step 3, Search for the optimal structure factors: Use the decomposition architecture to search for the FAS to determine the optimal structure factor α of the search space Ω in Step 1 * and the optimal structure factor θ of the search space Ψ * , and use the optimal structure factor α * and the optimal structure factor θ * to replace the search space Ω and the search space Ψ in the search framework of the change detection network SCTN built in Step 2, respectively; Step 4, construct a change detection network: utilize the optimal structure factor α obtained in Step 3 * and the optimal structure factor θ * , to construct the final change detection network SCTN; Step 5, training the change detection network: Train the change detection network SCTN constructed in Step 4 on the training set {X tra , Y tra} until the network converges, and test the trained change detection network SCTN on the test set {X tes , Y tes} to obtain the final change result.
2. The remote sensing image change detection method based on transformer architecture search according to claim 1, characterized in that The specific process of Step 1 is as follows: Step 1.1, setting search space Ω: The search space Ω is the Attention Search Space, including spatial and channel attention modules. The spatial and channel attention modules are three different hierarchical order combinations α of the spatial attention module and the channel attention module. The different combinations α are respectively: CSA module: It means that the feature map is first input into the channel attention module and then into the spatial attention module; SCA module: It means that the feature map is first input into the spatial attention module and then into the channel attention module; S&C module: It means that the feature map is input into the spatial attention module and the channel attention module in parallel and then added together; In each combined module, the weighted output feature can be calculated by the following formula: where α represents an operation in the search space Ω, and β α represents the trainable weight of each unit in the combination module, F out represents the output tensor, F in represents the input tensor, exp{β α} represents the weight value of each operator, β represents the trainable architecture parameter, ∑ α' exp{β α'} represents the normalization term; Step 1.2, setting search space Ψ: The search space Ψ is the Multi-scale Search Space, including a multi-scale fusion module. The multi-scale fusion module is three different combinations θ of the multi-scale module. The different combinations θ are respectively: CSAP module: It means a multi-scale module that first inputs the channel attention module and then the spatial attention module; SCAP module: It means a multi-scale module that first inputs the spatial attention module and then the channel attention module; S&CP module: It means a multi-scale module that inputs the spatial attention module and the channel attention module in parallel.
3. A remote sensing image change detection method based on transformer architecture search according to claim 1, characterized in that, The specific process of Step 2 is as follows: Step 2.1, input images and the image Input the image and the image After concatenating the images on the channel, input them into the ResNet18 network to extract preliminary features. The image and the image represent two temporal images at different times i1 and i2; Step 2.2, input the preliminary features extracted in Step 2.1 into the spatial and channel attention modules in search space Ω set in Step 1 to obtain output features; Step 2.3, input the output features of Step 2.2 into the multi-scale fusion module in search space Ψ set in Step 1 to extract features of different scales; Step 2.4, concatenate the output features of Step 2.2 and the features of different scales extracted in Step 2.3, and input them into the classifier to obtain the final change map.
4. A remote sensing image change detection method based on transformer architecture search according to claim 1, characterized in that, Step 3 uses the Factorized Architecture Search (FAS) to determine the optimal structure factors, including two sub-processes, FAS1 and FAS2. Specifically: Step 3.1, FAS1: Select the multi-scale fusion module combinations in the search space Ψ Fix different combinations θ of the multi-scale fusion modules in the search space Ψ as the combination Find the optimal combination α of the spatial and channel attention modules in the search space Ω * , that is, the optimal structure factor α * , the formula is as follows: Where, ω represents the network parameters to be trained, {X D , Y D} represents the development set, {X val , Y val} represents the validation set, λ1 represents the learning rate of the FAS1 process, represents the fixed structure factor set in the network architecture, α represents the initialized network architecture factor, represents the network architecture factor in the search process, ▽ ω represents the loss function about the gradient of ω, α * represents an optimal factor in the network architecture. By training the network parameters ω on the development set X D , find the optimal combination α val on the validation set X * ; Step 3.2, FAS2: Fix different combinations α of the spatial and channel attention modules in the search space Ω to the optimal combination α obtained in Step 3.1 * , and find the optimal combination θ of multi-scale fusion modules in the search space Ψ * , that is, the optimal structure factor θ * , and the formula is as follows: ω := ω - λ2▽ ω L(Y D , Net(X D , α * , θ, ω)) Where, ω represents the network parameters to be trained, {X D , Y D} represents the development set, {X val , Y val} represents the validation set, λ2 represents the learning rate of the FAS2 process, α * represents the optimal factor of the network architecture searched by FAS1, θ represents the initialized network architecture factor, represents the network architecture factor during the search process, ▽ ω represents the loss function about the gradient of ω, θ * represents an optimal factor in the network architecture. By training the network parameters ω on the development set X D , the optimal combination θ val is found on the validation set X * ; Step 3.3: Based on the optimal structure factor α obtained in Step 3.1 * Determine the optimal combination S&C module of the spatial and channel attention modules in the search space Ω; Based on the optimal structure factor θ obtained in Step 3.2 * Determine the optimal combination S&CP module of the multi-scale fusion modules in the search space Ψ; Step 3.4: Remove search space Ω and replace it with the S&C module obtained in Step 3.3; Remove search space Ψ and replace it with the S&CP module obtained in Step 3.
3.
5. A remote sensing image change detection method based on transformer architecture search according to claim 1, characterized in that, The specific process of Step 4 is as follows: Step 4.1, input images and the image Concatenate the images and the image on the channel and then input them into the ResNet18 network to extract preliminary features; Step 4.2, input the features extracted in Step 4.1 into the S&C module, that is, input the image into the spatial attention module and the channel attention module in parallel and then add them together to obtain output features; Step 4.3, input the output features of Step 4.2 into the S&CP module to extract features of different scales through the multi-scale module; Step 4.4, concatenate the output features of Step 4.2 and the features of different scales extracted in Step 4.3 on the channel, and input them into the classifier to obtain the final change map.
6. The remote sensing image change detection method based on transformer architecture search according to claim 2, characterized in that The spatial attention module in step 1.1 is implemented by a two-dimensional convolutional kernel of size 1×1. The specific construction method is as follows: (1) Calculate query, key, and value, and the formulas are as follows: where \(g(\cdot)\) represents the use of a two-dimensional \(1\times1\) convolution, and \(W\) Q , \(W\) K and \(W\) V respectively represent the matrices randomly initialized for query, key, and value. \(c\), \(w\), \(h\), and \(c'\) respectively represent the number of channels, height, width of the input feature map \(X\), and the number of channels of \(Q\) and \(K\); (2) Reshape the sizes of Q and K to c'×hw for subsequent matrix multiplication. The final spatial attention output is as follows: Output = V·softmax(Q T K) wherein, the softmax function represents normalizing the attention weights, and Q T represents the transpose of the query matrix Q, and Output represents the output feature; Multiply the normalized attention weight matrix softmax(Q T K) and V to obtain the output feature Output.
7. A remote sensing image change detection method based on transformer architecture search according to claim 2, characterized in that, (2) Multiply the result obtained in step (1) by the scale parameter λ and perform an element-wise summation operation on F to obtain the final channel attention output. The formula is as follows: (1) Reshape the input feature into where N = H × W, then perform matrix multiplication on the transpose of F and F, and obtain the channel attention map through the softmax function The formula is as follows: In the formula, represents the influence of the i-th channel on the j-th channel. The stronger the correlation between the two channels, the larger the value; F represents the input feature, and F i represents the i-th channel feature after the input feature is reshaped, and F j represents the j-th channel feature after the input feature is reshaped; (2) Multiply the result obtained in step (1) by the scale parameter λ and perform an element-wise summation operation on F to obtain the final channel attention output. The formula is as follows: Where λ represents the scale parameter, which is initialized to 0 and gradually learns to assign more weights; represents the influence of the i-th channel on the j-th channel, F i represents the i-th channel feature after reshaping the input feature, F j represents the j-th channel feature after reshaping the input feature, represents the j-th channel feature of the output.
8. A remote sensing image change detection system based on transformer architecture search, characterized in that, including: Search space setting module: Set the search space for neural architecture search, where the search space includes search space Ω and search space Ψ; Search framework building module: Use the search space to build the search framework of the change detection network SCTN; Optimal Structure Factor Search Module: Use the decomposition architecture to search for the FAS to determine the optimal structure factor α of the search space Ω * and the optimal structure factor θ of the search space Ψ * , and use the optimal structure factor α * and the optimal structure factor θ * to replace the search space Ω and the search space Ψ in the search framework of the change detection network SCTN, respectively; Change detection network building module: Used to build the final change detection network SCTN; Change detection network training module: Train the change detection network SCTN constructed by the change detection network construction module on the training set {X tra , Y tra} until the network converges, and test the trained change detection network SCTN on the test set {X tes , Y tes} to obtain the final change result.
9. A remote sensing image change detection device based on transformer architecture search, characterized in that, including: Memory: Used to store the computer program of a remote sensing image change detection method based on transformer architecture search implementing claims 1-7; Processor: When executing the computer program, it implements a remote sensing image change detection method based on transformer architecture search of claims 1-7.
10. A computer-readable storage medium, characterized in that, including: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it can implement a remote sensing image change detection method based on transformer architecture search of claims 1-7.
Citation Information
Patent Citations
Hyperspectral change detection method based on collaborative analysis autonomous sensing network structure
CN114842328A
Remote sensing image change detection method based on neural network structure search
CN115601660A