Hyperspectral Image Change Detection Method Based on Enhancement of Local Information
Through the combination of Graph-transformer and LIEG blocks, the global and local feature extraction problems in hyperspectral image change detection are solved, efficient change detection is achieved, and detection accuracy is improved.
Patent Information
- Application Number
- CN202211184372.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-27
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-09-27
AI Technical Summary
The existing hyperspectral image change detection methods are difficult to effectively extract global and local features, and the calculation complexity is high, and the traditional methods rely on manual feature representation, resulting in insufficient detection accuracy.
Graph-transformer is used to model global feature, and local information is enhanced through convolutional operations. The LIEG block is designed to build a dual-branch change detection network, combining SLIC superpixel segmentation and multi-layer LIEG block to extract the local and global features of multi-time phase HSI, and using the cross entropy loss function to optimize network parameters.
The accuracy of change detection is improved, the calculation cost is reduced, and the local and global features of the image are effectively extracted, achieving more efficient change detection.
Smart Images

Figure CN115564721B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a hyperspectral image change detection method based on local information enhancement. Background Art
[0002] Change detection is an important technology that uses multi-temporal images of the same area to identify changes in natural scenes. With the increasing popularity of remote sensing satellite images, change detection has been developed and widely applied in large-scale remote sensing research activities, including natural disasters, urban expansion research, land cover change, and water resource management. As one of many remote sensing data, HSI has hundreds of continuous spectral bands from ultraviolet to mid-infrared wavelengths, which can show the multi-spectral characteristics of land covers, and HSI has become an effective tool for land cover change detection.
[0003] Change detection has currently become a hot research direction, and many scholars have proposed some classic change detection methods. These methods can basically be divided into four categories, namely image algebra methods, image transformation methods, classification detection methods, and other classic algorithms. Image algebra methods calculate the image differences between multi-temporal data to obtain change detection results, such as sequential spectral change vector analysis (S2CVA). Methods based on image transformation, such as principal component analysis (PCA) and iterative reweighted multivariate alteration detection (IR-MAD), distinguish changed and unchanged regions by transforming HSI into other feature spaces. Classification methods use specific classifiers, such as support vector machines (SVM), to distinguish HSI of two different time phases respectively. Other classic algorithms include Markov random fields and random forests, which have also been proven to be promising in hyperspectral image change detection. However, the above methods rely on manually extracted feature representations and currently face many challenges.
[0004] Recently, deep learning methods with the ability to automatically extract deep features have been widely applied in change detection tasks. Although CNN-based methods have been proven to be effective in hyperspectral image change detection tasks, there are still some serious problems. Specifically, the receptive field of CNN is severely limited by the size of the convolutional kernel, which makes it only focus on local features and difficult to model the global context information in complex scenes.
[0005] Transformer with excellent learning ability can well focus on the global information of images and shows good performance in many image processing tasks, but it still has some disadvantages. Specifically, in terms of computational complexity, processing high-dimensional HSI data is a huge challenge for Transformer.
[0006] Summarize the defects of the above methods: (1) Traditional methods only manually extract shallow features, which limits their ability to express high-level features and thus loses some details. (2) The receptive field of the CNN method is severely restricted by the size of the convolutional kernel, which makes it only focus on local features and difficult to model the global context information in complex scenes; (3) The huge computational cost is a major challenge for Transformer in processing high-dimensional HSI data.
[0007] To address the above defects, the problems solved by the present invention are as follows
[0008] (1) The present invention designs a Graph-transformer that can accept graph node sequences to model global features and fully considers the spatial-spectral correlation between pixels, greatly reducing the computational cost.
[0009] (2) Enhance the local information of the Graph-transformer through convolution, which can fully extract the local-global features of the image. Summary of the Invention
[0010] The object of the present invention is to solve the problems existing in the prior art, and propose a hyperspectral image change detection method based on local information enhancement.
[0011] To achieve the above object, the present invention adopts the following technical solutions:
[0012] The hyperspectral image change detection method based on local information enhancement includes:
[0013] Input two multi-temporal hyperspectral images acquired at the same area at different times, preprocess the images, and select a training sample set;
[0014] Introduce the simple linear iterative clustering (SLIC) method to perform superpixel segmentation on the images to obtain graph nodes;
[0015] Construct a Graph-transformer to model the global context of the images;
[0016] Design a LIEG block that can effectively extract local and global feature information simultaneously, and obtain local information through convolution operations to inject into the Graph-transformer;
[0017] Multiple layers of LIEG blocks are used to construct a two-branch change detection network to obtain a D-LIEG network structure to fully extract the local and global features of multi-temporal HSI;
[0018] Take the difference between the outputs of the two branches to obtain differential features, and obtain a change detection prediction map after classification;
[0019] The constructed D-LIEG network model is trained in a supervised manner to obtain network parameters suitable for the model.
[0020] Furthermore, for the hyperspectral image change detection method based on local information enhancement, two multi-temporal hyperspectral images acquired at the same area at different times are input, and the images are preprocessed by maximum-minimum normalization. A training sample set is selected, and the normalization formula is:
[0021]
[0022] where x i represents a pixel in the hyperspectral image, x min and x max represent the maximum and minimum values of the hyperspectral image respectively, is a pixel after normalization. 1% and 0.5% of the total samples are randomly selected as the training sample set.
[0023] As a further technical solution of the present invention, the simple linear iterative clustering (SLIC) method is introduced to perform superpixel segmentation on the image to obtain graph nodes.
[0024] To construct graph nodes, a region segmentation-based method called simple linear iterative clustering (SLIC) is introduced. This method gradually grows local clusters iteratively through the K-means algorithm until the iteration reaches the optimum to complete the segmentation operation. In this way, pixels with hyperspectral spatial similarity are usually divided into the same image region (i.e., graph nodes). To ensure consistent graph node distributions for the two segmented HSIs, the input multi-temporal hyperspectral images T1 and T2 are concatenated along the channel dimension and then divided into a series of compact regions. The average spectral vector value of the pixels included in the segmented region is determined as the feature vector value of the corresponding graph node, and the original image graph nodes and N is the number of graph nodes. The mapping relationship between the original graph and the graph nodes can be expressed as M(·);
[0025] G = M(Concat(T1, T2)) = A T Concat(T1, T2)
[0026] where is the association matrix between the segmentation result and the original image, and Concat(·) is cross-channel feature concatenation.
[0027] As a further technical solution of the present invention, a Graph-transformer is designed to model the global context of the image.
[0028] First, perform an input embedding operation on the feature vectors of the input graph nodes and add position encoding to increase position information, which serves as the input to the Graph-transformer. The Graph-transformer consists of three cascaded encoders, and each encoder layer is composed of a multi-head self-attention MHSA, a multi-layer perceptron MLP, layer normalization LN, and a residual connection;
[0029] (1) Pass the input graph nodes through a fully connected layer. The number of output nodes of the fully connected layer is set to 256. The position encoder uses learnable position encoding, and the output after position encoding is used as the input to the Graph-transformer.
[0030] (2) The MHSA attention operation Att(·) is defined as:
[0031]
[0032] where Q represents the query matrix, K represents the key matrix to be queried, and V represents the value matrix of the output.
[0033] In this paper, the multi-head self-attention mechanism is used, and its formula is:
[0034] MHSA(LN(Z i )) = Concat(h1, h2,..., h s )W 0
[0035] where:
[0036]
[0037] where i = 1, 2, the number of heads s of the multi-head attention mechanism is set to 8, B ii = 1 / ∑ j A ij is a diagonal matrix, and its size performs normalization processing on the feature matrix. is used to unify the dimensions of the input feature matrix; is used to add position encoding to increase position information. In this paper, learnable position encoding is adopted. is a learnable parameter matrix, and LN(·) represents layer normalization. The result is sent to the MLP for feature integration.
[0038] (3) The multi-layer perceptron consists of two linear transformation layers and a Gelu activation function to further transform all the learned features of the heads. The number of output nodes of the two linear layers is set to 128 and 256 respectively. Before entering the fully connected layer, first perform layer normalization on the output of Att(·).
[0039] (4) In addition, in order to avoid the loss of feature information, residual connections are adopted at the outputs of Att(·) and the fully connected layer respectively.
[0040] (5) The forward propagation process of the entire Graph-transformer can be described as:
[0041] G (l+1) = f gt (G (l) )
[0042] where G (l+1) is the output of the i-th layer of the Graph-transformer. f gt (·) represents the Graph-transformer composed of n encoders.
[0043] As a further technical solution of the present invention, an LIEG block that can effectively extract local and global feature information simultaneously is designed, and local information is obtained through convolution operations and injected into the Graph-transformer.
[0044] (1) Since the feature vector of each graph node is the average value of the spectral vectors of the pixels included in the node, this results in partial loss of local information. Therefore, local features of the HSI are extracted through convolution operations and mapped to the feature matrix.
[0045] (2) The feature matrix obtained after convolution and the initial input feature matrix are concatenated along the channel dimension and fed into a fully connected layer to unify the input dimension, obtaining the input of the Graph-transformer with enhanced local information. The formula is:
[0046]
[0047] where L (l) and G (l) respectively represent the input feature map and the feature matrix of the l-th Graph-transformer. G (0) is G, L (0) is T. Conv(·) represents the local information extraction convolution operation, is the input of the l-th Graph-transformer layer. The size of the convolution kernel is set to 3, and the output channel dimension is 256. The number of output nodes of the fully connected layer is set to 256.
[0048] (3) The forward model of LIEG can be simplified as:
[0049]
[0050] L(l+1) = Conv(L (l) )
[0051] Among them, and L (l+1) are the output feature matrix and feature map of the l-th LIEG block respectively.
[0052] As a further technical solution of the present invention, multiple LIEG blocks are used to construct a dual-branch change detection network, resulting in D-LIEG, to fully extract the local-global features of multi-temporal HSI;
[0053] (1) D-LIEG adopts a dual-branch structure to obtain sufficient features of multi-temporal HSI to distinguish different objects, and each branch consists of multiple LIEG blocks to extract complementary local and global features. The propagation process of multiple LIEG is described as follows:
[0054]
[0055] where i = 1, 2, representing two branches. m is the number of LIEG blocks in each branch, and represent the output feature matrix and feature map respectively, f LIEG (·) is a simplified description of the forward propagation process of LIEG. It is verified in experiments that m = 3 has the best performance.
[0056] (2) To fully retain local information, the convolutional output of the last LIEG block is fed into a convolutional layer and then cascaded with the output of the Graph-transformer. It is defined as:
[0057]
[0058] where B i is the output of the i-th branch.
[0059] As a further technical solution of the present invention, the outputs of the two branches are subtracted to obtain differential features, and after classification, the final change detection prediction map is obtained:
[0060] The outputs of the two branches containing sufficient features are subtracted to obtain differential features, and then they are converted into a feature map using an association matrix. Finally, the feature map is sent to a classifier composed of two fully connected layers, a Relu activation function, and a softmax non-linear activation function, and the prediction result of change detection is further obtained as:
[0061]
[0062] where f c1 and f c2They respectively represent two fully connected layers. The number of output nodes is set to 128 and 256 respectively. It is the predicted map of the output.
[0063] As a further technical solution of the present invention, the established D-LIEG network model is trained in a supervised manner to obtain network parameters suitable for the model;
[0064] (1) Input the labeled training samples into the network model to be trained, and output the label prediction for the training samples;
[0065] (2) Use the following cross-entropy loss function to calculate the loss function between the predicted label and the true label of the reference image:
[0066]
[0067] where w is the number of samples, Y represents the reference image, is the output predicted map. E can quantitatively reflect the difference between the model prediction result and the true label, and the optimal network model can be obtained by minimizing E.
[0068] (3) Use the stochastic gradient descent method to train the network parameters until the network converges, and save the optimal network parameters to complete the discrimination of the two categories of changed and unchanged.
[0069] The beneficial effects of the present invention:
[0070] 1. The present invention proposes Graph-transformer, which realizes the application of Transformer in the HIS change detection task. It should be noted that Graph-transformer fully considers the spatial-spectral correlation between pixels and greatly reduces the computational cost.
[0071] 2. The present invention innovatively proposes the LIEG block composed of Graph-transformer with global representation ability and convolutional operation with local acquisition ability, which can use convolution to enhance the local information of the image and effectively extract local and global feature information simultaneously.
[0072] 3. The double-branch structure composed of multiple LIEG blocks in the present invention fully extracts the features of multi-temporal HIS and realizes the distinction of different features. Description of the Drawings
[0073] Figure 1 It is the flowchart of the hyperspectral image change detection method provided by the embodiment of the present invention.
[0074] Figure 2 It is the structural schematic diagram of LIEG provided by the embodiment of the present invention.
[0075] Figure 3 It is the D-LIEG network structure diagram provided by the embodiments of the present invention.
[0076] Figure 4 Among them: (a) is the Groud-truth standard diagram; (b) is the result diagram of the CVA method; (c) is the result diagram of the PCA method; (d) is the result diagram of the IR-MAD method; (e) is the result diagram of the SVM method; (f) is the result diagram of the ReCNN method; (g) is the result diagram of the present invention. Detailed implementation manners
[0077] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined purpose, the following will, in conjunction with the accompanying drawings and preferred embodiments, detail the specific implementation manners, structures, features and their effects of the present invention as follows.
[0078] Refer to Figures 1-4 , a hyperspectral image change detection method based on local information enhancement, and the present invention will be described in detail below with reference to the accompanying drawings.
[0079] As Figure 1 shown, the hyperspectral image change detection method based on local information enhancement provided by the present invention includes the following steps:
[0080] S101 Input two multi-temporal hyperspectral images acquired at different times in the same area, preprocess the images, and select a training sample set;
[0081] S102 Introduce the simple linear iterative clustering (SLIC) method to perform superpixel segmentation on the images to obtain graph nodes;
[0082] S103 Design a Graph-transformer to model the global context of the images;
[0083] S104 Design a LIEG block that can effectively extract local and global feature information at the same time, and obtain local information injection into the Graph-transformer through convolution operations;
[0084] S105 Multiple LIEG blocks are used to construct a dual-branch change detection network to obtain a D-LIEG network structure, so as to fully extract the local-global features of multi-temporal HSI;
[0085] S106 Take the difference between the outputs of the two branches to obtain differential features, and obtain a change detection prediction map after classification;
[0086] S107 Perform supervised training on the built D-LIEG network model to obtain network parameters suitable for the model.
[0087] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0088] As Figure 1 shown, the hyperspectral image change detection method based on local information enhancement provided by the embodiment of the present invention is implemented as follows:
[0089] (1) Input two multi-temporal hyperspectral images acquired at the same area at different times, preprocess the images by maximum-minimum normalization, and select a training sample set. The normalization formula is:
[0090]
[0091] where x i represents a pixel in the hyperspectral image, x min and x max represent the maximum and minimum values of the hyperspectral image respectively, is a pixel after normalization. Randomly select 1% and 0.5% of the total samples as the training sample set.
[0092] (2) Introduce the simple linear iterative clustering (SLIC) method to perform superpixel segmentation on the image to obtain graph nodes;
[0093] To construct graph nodes, a region segmentation-based method called simple linear iterative clustering (SLIC) is introduced. This method iteratively grows local clusters through the K-means algorithm until the iteration reaches the optimum to complete the segmentation operation. In this way, pixels with hyperspectral spatial similarity are usually divided into the same image region (i.e., graph nodes). To ensure consistent graph node distribution for the two segmented HSIs, the input multi-temporal hyperspectral images T1 and T2 are concatenated along the channel dimension and then divided into a series of compact regions. The average spectral vector value of the pixels included in the segmented region is determined as the feature vector value of the corresponding graph node, and the original image graph nodes and N is the number of graph nodes. The mapping relationship between the original graph and the graph nodes can be expressed as M(·);
[0094] G = M(Concat(T1,T2)) = A T Concat(T1,T2)
[0095] where is the association matrix between the segmentation result and the original image, and Concat(·) is the cross-channel feature connection.
[0096] (3) Design a Graph-transformer to model the global context of the image;
[0097] First, perform an input embedding operation on the feature vectors of the input graph nodes and add positional encoding to increase positional information, which serves as the input to the Graph-transformer. The Graph-transformer consists of 3 cascaded encoders. Each encoder layer is composed of multi-head self-attention (MHSA), a multi-layer perceptron (MLP), layer normalization (LN), and residual connections;
[0098] (3a) Pass the input graph nodes through a fully connected layer. The number of output nodes of the fully connected layer is set to 256. The positional encoder uses learnable positional encoding, and the output after positional encoding is used as the input to the Graph-transformer.
[0099] (3b) The MHSA attention operation Att(·) is defined as:
[0100]
[0101] where Q represents the query matrix, K represents the key matrix to be queried, and V represents the value matrix of the output.
[0102] In this paper, the multi-head self-attention mechanism is used, and its formula is:
[0103] MHSA(LN(Z i )) = Concat(h1, h2,..., h s )W 0
[0104] where:
[0105]
[0106] where i = 1, 2, the number of heads s of the multi-head attention mechanism is set to 8, B ii = 1 / ∑ j A ij is a diagonal matrix, and its size performs normalization on the feature matrix. is used to unify the dimensions of the input feature matrix; is used to add positional encoding to increase positional information. In this paper, learnable positional encoding is adopted. is a learnable parameter matrix, and LN(·) represents layer normalization. The result is fed into the MLP for feature integration.
[0107] (3c) The multi-layer perceptron consists of two linear transformation layers and a Gelu activation function to further transform all head learning features. The number of output nodes of the two linear layers is set to 128 and 256 respectively. Before entering the fully connected layer, layer normalization is performed on the output of Att(·).
[0108] (3d) In addition, to avoid the loss of feature information, residual connections are adopted at the outputs of Att(·) and the fully connected layer respectively.
[0109] (3e) The forward propagation process of the entire Graph-transformer can be described as:
[0110] G (l+1) =f gt (G (l) )
[0111] where G (l+1) is the output of the i-th layer of the Graph-transformer. f gt (·) represents the Graph-transformer composed of n encoders.
[0112] (4) As Figure 2 shown, design the LIEG block that can effectively extract local and global feature information simultaneously, obtain local information through convolution operations and inject it into the Graph-transformer;
[0113] (4a) Since the feature vector of each graph node is the average value of the spectral vectors of the pixels contained in the node, which leads to partial loss of local information. Therefore, local features of the HSI are extracted through convolution operations and mapped to the feature matrix.
[0114] (4b) The convolved feature matrix and the initial input feature matrix are concatenated along the channel dimension and fed into a fully connected layer to unify the input dimension, obtaining the input of the Graph-transformer with enhanced local information. The formula is:
[0115]
[0116] where L (l) and G (l) represent the input feature map and the feature matrix of the l-th Graph-transformer respectively. G (0) is G, L (0) is T. Conv(·) represents the local information extraction convolution operation, is the input of the l-th Graph-transformer layer. The size of the convolution kernel is set to 3, and the output channel dimension is 256. The number of output nodes of the fully connected layer is set to 256.
[0117] (4c) The forward model of LIEG can be simplified as:
[0118]
[0119] L (l+1) = Conv(L (l) )
[0120] where and L (l+1) are the output feature matrix and feature map of the l-th LIEG block.
[0121] (5) As Figure 3 shown, multiple LIEG blocks are used to construct a dual-branch change detection network to obtain the D-LIEG network structure, so as to fully extract the local and global features of multi-temporal HSI;
[0122] (5a) D-LIEG adopts a dual-branch structure to obtain sufficient features of multi-temporal HSI to distinguish different objects, and each branch consists of multiple LIEG blocks to extract complementary local and global features. The propagation process of multiple LIEG blocks is described as follows:
[0123]
[0124] where i = 1, 2, representing two branches. m is the number of LIEG blocks in each branch, and represent the output feature matrix and feature map respectively, f LIEG (·) is a simplified description of the forward propagation process of the LIEG block. It is verified in the experiment that m = 3 has the best performance.
[0125] (5b) To fully retain local information, the last convolutional output of the LIEG block is fed into a convolutional layer and then cascaded with the output of the Graph-transformer. It is defined as:
[0126]
[0127] where B i is the output of the i-th branch.
[0128] (6) The outputs of the two branches are subtracted to obtain the differential features, and after classification, the final change detection prediction map is obtained:
[0129] The outputs of two branches containing sufficient features are differentiated to obtain differential features, and then they are converted into a feature map using an association matrix. Finally, the feature map is sent to a classifier composed of two fully connected layers, a Relu activation function, and a softmax non-linear activation function to further obtain the prediction result of change detection as follows:
[0130]
[0131] where f c1 and f c2 respectively represent two fully connected layers. The number of output nodes is 128 and 256 respectively. is the predicted map of the output.
[0132] (7) Supervised training is performed on the built D-LIEG network model to obtain network parameters suitable for the model;
[0133] (7a) The labeled training samples are input into the network model to be trained, and the label prediction for the training samples is output;
[0134] (7b) Using the following cross-entropy loss function, calculate the loss function between the predicted label and the true label of the reference image:
[0135]
[0136] where w is the number of samples, Y represents the reference image, is the output predicted map. E can quantitatively reflect the difference between the model prediction result and the true label, and the optimal network model can be obtained by minimizing E.
[0137] (7c) Use the stochastic gradient descent method to train the network parameters until the network converges, and save the optimal network parameters to complete the discrimination of the two categories of changed and unchanged. We use the Adam optimizer with a learning rate of 1e-5 and complete the learning process after 800 epochs.
[0138] The technical effects of the present invention are described in detail below in combination with simulation experiments:
[0139] 1. Simulation experiment conditions:
[0140] The hardware platform for the simulation experiment of the present invention is: NVIDIA AGTX 3090 GPU
[0141] The software platform for the simulation experiment of the present invention is: Linux18.06 operating system, python3.7, and pytorch1.12.
[0142] The hyperspectral images used in the simulation experiment of the present invention are Santa Barbara images, which are captured using an airborne visible / infrared imaging spectrometer (AVIRIS) sensor. The HIS datasets of two temporal phases captured in the Santa Barbara area in 2013 and 2014 have 250×250 pixels, including 224 bands, with a spectral range of 0.4 to 2.5 μm, and the number of graph nodes N is set to 300.
[0143] 2. Experimental Content and Result Analysis
[0144] To verify the effectiveness of the proposed D-LIEG method, we selected five widely used hyperspectral image change detection methods, including CVA, PCA, IR-MAD, SVM, and ReCNN. The input Santa Barbara hyperspectral images were respectively subjected to change detection to obtain the final change detection result map.
[0145] The prior art comparison change detection method used in the present invention refers to:
[0146] The prior art change vector analysis CVA change detection method refers to the change detection method proposed by Malila et al. in the literature "Change-vector analysis in multitemporal space: a tool to detect and categorize land-cover change processes using high temporal-resolution satellite data [J]. Remote Sensing of Environment, 1994, 48(2): 231-244."
[0147] The prior art principal component analysis PCA change detection method refers to the change detection method proposed by Deng et al. in the literature "PCA-based land-use change detection and analysis using multitemporal and multisensory satellite data [J]. Int. J. Remote Sens., vol. 29, no. 16, pp. 4823–4838, 2008."
[0148] The existing iterative reweighted multivariate change detection method IR-MAD refers to the change detection method proposed by Nielsen et al. in the literature "The Regularized Iteratively Reweighted MAD Method for Change Detection in Multi and Hyperspectral Data".
[0149] The existing support vector machine SVM classification method refers to the hyperspectral image classification method proposed by Hearst et al. in "Support vector machines. IEEE Intelligent Systems and Their Applications, 13(4):18–21.", which is abbreviated as the support vector machine SVM classification method.
[0150] The existing recurrent convolutional neural network ReCNN refers to the hyperspectral image classification method proposed by Mou, Bruzzone, and Zhu et al. in "Learning Spectral-Spatial-Temporal Features via a Recurrent Convolutional Neural Network for Change Detection in Multispectral Imagery", which is abbreviated as the recurrent convolutional neural network ReCNN classification method.
[0151] The following Figure 4 is combined with the change detection result diagram in
[0152] From Figure 4 (b) of
[0153] it can be seen that the CVA method has some noise points in the change map because it starts from the Euclidean distance between pixels and is sensitive to the noise in the input image. And there are a large number of false detections in the changed area. Figure 4 From
[0154] (c) and (d) of Figure 4 it can be seen that PCA and IR-MAD improve performance by reducing redundant information in the spectrum, but still do not solve the problem of noise sensitivity and cannot effectively detect some boundary regions.
[0155] From Figure 4As can be seen from (f), when the training sample size accounts for 1%, ReCNN detects some changing pixels, but a large number of unchanged regions are wrongly detected as changing regions.
[0156] From Figure 4 As can be seen from (g), the results obtained by the proposed D-LIEG change detection method are closest to the ground truth of changes. It has fewer noise points and smaller misdetection regions.
[0157] The change detection result maps obtained by six methods are objectively evaluated using two evaluation metrics: overall accuracy (OA) and Kappa coefficient. The overall accuracy OA represents the proportion of correctly classified samples in the total samples. The closer the OA value is to 1, the higher the detection accuracy. The Kappa coefficient characterizes the consistency between the obtained results and the reference map. The closer the Kappa value is to 1, the better the performance of the method. The values of various evaluation metrics are tabulated in Table 1.
[0158] Table 1 Quantitative analysis table of change detection results of the present invention and existing inventions for Santa Barbara hyperspectral images
[0159]
[0160] Combined with Table 1, it can be seen that when the training set is selected at 1%, the overall accuracy OA of the present invention reaches 97.27% and the Kappa value reaches 0.9439, which are respectively 1.67% and 4.79% higher than those of the best-performing method (ReCNN) among the currently compared methods, and are significantly higher than the existing technology methods, proving that the present invention can better detect the changing regions and its performance is significantly better than the existing methods.
[0161] The above simulation experiments show that the present invention proposes a Dual-branch Local Information Enhanced Graph-transformer (D-LIEG) change detection network for hyperspectral image change detection tasks. The transformer that can model global features is introduced to solve the problem of hyperspectral image change detection. The Graph-transformer is innovatively designed, which not only improves the computational efficiency but also maintains the spatial and spectral correlations between pixels. In order to reduce the loss of local information in the Graph-transformer, a local information enhancement module is proposed to inject the local information obtained by convolution into the Graph-transformer to fully extract local-global features. Each branch of the dual-branch network structure D-LIEG consists of multiple LIEG blocks, which are used to extract sufficient features of multi-temporal HSI and send them to the classifier to realize the discrimination and prediction of changed and unchanged regions. A large number of experiments show that the present invention has achieved excellent performance in both quantitative and qualitative results, effectively improving the accuracy of change detection.
[0162] As described above, it is only a preferred embodiment of the present invention and does not impose any form of limitation on the present invention. Although the present invention has been disclosed as above with the preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or equivalent changes by using the disclosed technical content within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention, any brief modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A hyperspectral image change detection method based on local information enhancement, characterized in that The specific steps are as follows: S101. Input two multi-temporal hyperspectral images acquired at different times in the same area, preprocess the images, and select a training sample set; S102. Introduce the simple linear iterative clustering (SLIC) method to perform superpixel segmentation on the images to obtain graph nodes; S103. Construct a Graph-transformer to model the global context of the images; S104. Design a LIEG block that can effectively extract local and global feature information simultaneously, and obtain local information through convolution operations and inject it into the Graph-transformer; S105. Multiple LIEG blocks are used to construct a dual-branch change detection network to obtain a D-LIEG network structure for extracting local and global features of multi-temporal HSI; S106. Take the difference between the outputs of the two branches to obtain differential features, and obtain a change detection prediction map after classification; S107. Perform supervised training on the constructed D-LIEG network model to obtain network parameters suitable for the model.
2. The hyperspectral image change detection method based on local information enhancement according to claim 1, characterized in that In step S101, two multi-temporal hyperspectral images acquired at different times in the same area are input, the images are preprocessed by maximum-minimum normalization, and a training sample set is selected. The normalization formula is: where x i represents a pixel in the hyperspectral image, x min and x max represent the maximum and minimum values of the hyperspectral image, respectively, is a pixel after normalization; 1% or 0.5% of the total samples are randomly selected as the training sample set.
3. The hyperspectral image change detection method based on local information enhancement according to claim 1, characterized in that, In step S102, the simple linear iterative clustering (SLIC) method is introduced to perform superpixel segmentation on the images to obtain graph nodes. Specifically: The simple linear iterative clustering SLIC method gradually grows local clusters iteratively through the K-means algorithm until the iteration reaches the optimum to complete the segmentation operation, dividing pixels with high spectral-spatial similarity into the same image regions, i.e., graph nodes; cascading the input hyperspectral images of two time phases, T1 and T2, along the channel dimension, and then dividing them into a series of compact regions; determining the average spectral vector value of the pixels included in the segmentation regions as the eigenvector value corresponding to the graph nodes of the corresponding graphs, and establishing the graph nodes of the original images and N is the number of graph nodes; the mapping relationship between the original graph and the graph nodes can be expressed as M(·): G = M(Concat(T1, T2)) = A T Concat(T1, T2) Among them is the association matrix between the segmentation result and the original image, and Concat(·) is the feature connection across channels.
4. The hyperspectral image change detection method based on local information enhancement according to claim 1, characterized in that In step S103, constructing a Graph-transformer to model the global context of the images is specifically: First, perform an input embedding operation on the feature vectors of the input graph nodes and add positional encoding to increase positional information as the input of the Graph-transformer; the Graph-transformer consists of 3 cascaded encoders, and each encoder layer is composed of a multi-head self-attention (MHSA), a multi-layer perceptron (MLP), layer normalization (LN), and a residual connection; (1) Pass the input graph nodes through a fully connected layer, set the number of output nodes of the fully connected layer to 256, and the positional encoder uses learnable positional encoding, and use the output after positional encoding as the input of the Graph-transformer; (2) The MHSA attention operation Att(·) is defined as: Among them Q represents the Query matrix, K represents the Key matrix to be queried, and V represents the Value matrix of the output value; In the present invention, a multi-head self-attention mechanism is used, and its formula is: MHSA(LN(Z i )) = Concat(h1, h2,..., h s )W 0 where where \(i = 1, 2\), the number of heads \(s\) of the multi-head attention mechanism is set to 8, \(B\) ii = 1 / ∑ j A ij is a diagonal matrix, and its size performs normalization on the feature matrix; is used to unify the dimensions of the input matrix; is used to add positional encoding to increase positional information; is a learnable parameter matrix, \(LN(\cdot)\) represents layer normalization, and the output result is fed into the MLP for feature integration; (3) The multi-layer perceptron (MLP) consists of two linear transformation layers and a Gelu activation function to further learn the features of all heads; The number of output nodes of the two linear layers are set to 128 and 256 respectively; in addition, before entering the fully connected layer, first perform layer normalization on the output of Att(·); (4) In addition, to avoid loss of feature information, residual connections are used at the outputs of Att(·) and the fully connected layer respectively; (5) The forward propagation process of the entire Graph-transformer can be described as: G (l+1) = f gt (G (l) ) Among which G (l+1) is the output of the i-th layer of the Graph-transformer; f gt (·) represents the Graph-transformer composed of n encoders.
5. The hyperspectral image change detection method based on local information enhancement according to claim 1, characterized in that In step S104, design a LIEG block that can effectively extract local and global feature information simultaneously, obtain local information through convolution operations and inject it into the Graph-transformer. Specifically: (1) Extract the local features of HSI through convolution operation and map them to the feature matrix; (2) Concatenate the feature matrix obtained after convolution and the initial input feature matrix along the channel dimension, and feed them into a fully connected layer to unify the input dimension, obtaining the input of Graph-transformer with enhanced local information. The formula is: where L (l) and G (l) represent the input feature map and the feature matrix of the l-th Graph-transformer respectively; G (0) is G, and L (0) is T; Conv(·) represents the local information extraction convolution operation, is the input of the l-th Graph-transformer layer; the size of the convolution kernel is set to 3, and the output channel dimension is 256; the number of output nodes of the fully connected layer is set to 256; (3) The forward model of LIEG can be simplified as: L (l+1) = Conv(L (l) ) Among them, and L (l+1) are the output feature matrix and feature map of the l-th LIEG block.
6. The hyperspectral image change detection method based on local information enhancement according to claim 1, characterized in that, In step S105, the multi-layer LIEG blocks are used to construct a dual-branch change detection network to obtain the D-LIEG network structure, so as to fully extract the local-global features of multi-temporal HSI; Specifically: (1) D-LIEG adopts a dual-branch structure to obtain sufficient features of multi-temporal HSI to distinguish different objects. Each branch consists of multi-layer LIEG blocks to extract complementary local and global features; The propagation process of multi-layer LIEG is described as follows: where \(i = 1, 2\) represents two branches; \(m\) is the number of LIEG blocks in each branch, and represent the output feature matrix and the feature map respectively, f LIEG (·) is a simplified description of the LIEG forward propagation process; (2) To fully retain local information, the output of the last convolution of LIEG is fed into a convolution layer, and then concatenated with the output of Graph-transformer; defined as: Among which B i is the output of the i-th branch.
7. The hyperspectral image change detection method based on local information enhancement according to claim 1, characterized in that In step S106, the outputs of the two branches are subtracted to obtain the differential features, and the final change detection prediction map is obtained after classification. Specifically: Subtract the outputs of the two branches containing sufficient features to obtain the difference features, and then convert them into a feature map using the association matrix; finally, send the feature map to a classifier composed of two fully connected layers, a Relu activation function, and a softmax non-linear activation function, and further obtain the prediction result of change detection as: where f c1 and f c2 represent two fully connected layers respectively; the number of output nodes is set to 128 and 256 respectively; is the predicted map of the output.
8. The hyperspectral image change detection method based on local information enhancement according to claim 1, wherein In step S107, the constructed D-LIEG network model is supervised training to obtain the network parameters suitable for the model. Specifically: (1) Input the labeled training samples into the network model to be trained, and output the label prediction of the training samples; (2) Use the following cross-entropy loss function to calculate the loss function between the predicted label and the true label of the reference image: where \(w\) is the number of samples, \(Y\) represents the reference image, is the output prediction map; \(E\) can quantitatively reflect the difference between the model prediction result and the true label, and the optimal network model can be obtained by minimizing \(E\); (3) Use the stochastic gradient descent method to train the network parameters until the network converges, and save the optimal network parameters to complete the discrimination of the two categories of changed and unchanged.
Citation Information
Patent Citations
Local similarity preserving-based hyperspectral image extreme learning machine clustering method
CN108197650A
Hyperspectral image change detection method based on multistage cyclic convolution self-encoding network
CN112733725A