Hyperspectral Image Change Detection Method Based on Multi-Head Self-Cross Hybrid Attention
Through the hyperspectral image change detection method of multi-head self-cross hybrid attention, the abundance matrix learning network and multi-level multi-head self-cross hybrid attention mechanism are used to solve the problem of low accuracy caused by mixed cell phenomena, and achieve higher change detection accuracy.
Patent Information
- Application Number
- CN202211428908.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-15
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-11-15
AI Technical Summary
The existing hyperspectral image change detection algorithms have low accuracy and fail to fully utilize the correlation between bi-time phase images when facing mixed cell phenomena.
The hyperspectral image change detection method of multi-head self-cross hybrid attention is used to reduce the influence of mixed cells through the abundance matrix learning network and multi-level multi-head self-cross hybrid attention mechanism, and the difference information is extracted using the correlation between the two-time phase images.
The accuracy of change detection is significantly improved, the change area can be better identified and the interference of unchanged information can be reduced.
Smart Images

Figure CN115731462B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of hyperspectral remote sensing image change detection, and specifically relates to a hyperspectral image change detection method based on multi-head self-cross hybrid attention. Background Art
[0002] A hyperspectral image is a data cube composed of rich spectral dimension information and two-dimensional spatial information. Compared with multi-spectral images, hyperspectral images cover a wider wavelength range and can capture more refined ground object features, so they are widely used in remote sensing image processing technology. Hyperspectral image change detection refers to detecting changes between bi-temporal hyperspectral images obtained by the same sensor at different times, and has been widely applied in many fields such as mineral exploration, coastal wetland monitoring, and land cover monitoring.
[0003] Classic change detection algorithms can be roughly divided into: 1) image algebra methods; 2) image transformation methods; 3) classification detection methods; 4) other traditional methods; and 5) deep learning-based methods.
[0004] Image algebra methods obtain change feature maps through simple algebraic operations, such as image difference method, image ratio method, and change vector analysis (CVA). CVA is one of the most classic algorithms. By performing difference operations, the change vector between bi-temporal images is obtained, where the Euclidean distance of the change vector represents the change intensity, and the change content is represented by the direction of the change vector. The key to this type of method is to find the change threshold, and as the number of bands increases, the selection of the threshold becomes increasingly difficult.
[0005] Image transformation methods enhance change features and suppress invariant features by mapping the image to a new feature domain, while reducing data redundancy. Principal component analysis, slow feature analysis, and iterative weighted multivariate change are the most representative algorithms in image transformation methods. However, spectral information will inevitably be lost during the transformation process.
[0006] Classification-based change detection mainly includes post-classification comparison method and joint classification method. The principle of post-classification comparison is to classify the two temporal images separately, and then compare the classified regions pixel by pixel to determine the location and type of change information; the joint classification method is to stack multi-temporal images together to form a dataset for direct classification. Although this type of method detects the change categories, the classification results are too dependent on the performance of the classifier.
[0007] In addition to the above three classification methods, there are also some other methods for change detection that have achieved excellent results, such as those based on unmixing ideas, low-rank sparse representations, and tensor regression. However, the high-dimensional and redundant characteristics of hyperspectral data still pose great challenges to change detection.
[0008] In recent years, as a powerful tool for feature extraction, deep learning has become an important branch of remote sensing image processing. Although many advanced change detection algorithms have achieved excellent results at present, there are still some deficiencies: (1) Due to the low spatial resolution of spectral sensors and other external interference factors, there is a phenomenon of mixed pixels in hyperspectral images. The mixed pixels contain information of multiple ground objects, resulting in the phenomena of different objects with the same spectrum and the same object with different spectra, which greatly affects the accuracy of change detection; (2) Most of the deep learning-based change detection algorithms use a dual-branch structure, that is, the spatial-spectral features of the corresponding single-temporal image are independently extracted for each branch, and the correlation between the two-temporal images is not fully utilized. Summary of the Invention
[0009] In order to overcome the above-mentioned deficiencies of the prior art, the purpose of the present invention is to provide a hyperspectral image change detection method based on multi-head self-cross hybrid attention. The abundance matrix learning network learns the hyperspectral image abundance matrix based on endmember sharing, effectively weakening the influence of mixed pixels while mapping the differences between images to the corresponding abundance matrix; in addition, the difference information extraction module composed of a multi-level multi-head self-cross hybrid attention mechanism makes full use of the correlation between the two-temporal images, and the information of the unchanged regions in the obtained change feature map is suppressed, making the changed regions easier to identify.
[0010] In order to achieve the above purpose, the technical solution adopted by the present invention is:
[0011] A hyperspectral image change detection method based on multi-head self-cross hybrid attention, comprising the following steps;
[0012] S101: Input the two-temporal hyperspectral images into the vertex component analysis method and the abundance matrix learning network respectively, extract the endmember matrix, and learn the corresponding abundance matrix S based on endmember sharing i ;
[0013] S102: Reconstruct the image according to the extracted abundance matrix S i and the endmember matrix, and continuously reduce the difference between the reconstructed image and the original image to optimize the abundance matrix learning network;
[0014] S103: After dividing the abundance matrix S in S101 i into blocks, extract its shallow features, project them to different subspaces through a convolutional layer, and obtain the corresponding Q i,h 、K i,h 、V i,h matrix, and then input them into the multi-head self-cross hybrid attention mechanism;
[0015] S104: The multi-head self-crossing hybrid attention mechanism uses the correlation difference to weight and output its input, suppressing the invariant information of the two-phase images and retaining the differential information mapping of the changing information;
[0016] S105: Through the stacking of multiple multi-head self-crossing hybrid attention mechanisms, that is, the multi-level multi-head self-crossing hybrid attention mechanism, the differential information is continuously strengthened;
[0017] S106: The differential information extraction module composed of the multi-level multi-head self-crossing hybrid attention mechanism is supervised and trained. After determining the optimal network parameters, the predicted change detection map is output.
[0018] Furthermore, the specific content of step S101 is as follows:
[0019] The hyperspectral image is a linear combination of an endmember matrix and an abundance matrix, and its matrix form is expressed as:
[0020]
[0021] where represents the hyperspectral image, B is the number of spectral bands, M is the number of pixels, represents the endmember matrix, represents the abundance matrix, N represents the number of endmembers, represents the additive noise, is the constraint that the sum of the column vectors of the abundance matrix is 1, and 1 N is an all-ones vector of N×1; this constraint condition ensures the sparsity of the abundance matrix, better reflects the contribution degree of each endmember to the pixel spectrum, and at the same time reduces the computational complexity;
[0022] Let represent the two-phase hyperspectral images obtained at different times in the same area, where H and W represent the height and width of the image respectively; M = H×W represents the number of image pixels;
[0023] The hyperspectral images obtained at different times in the same scene will share the endmember matrix. First, the two input hyperspectral images are merged in the spatial dimension into Then, the vertex component analysis method is used to extract the endmember matrix of the merged image. This algorithm iteratively performs projections along the direction perpendicular to the subspace formed by the known endmembers. The projection extreme value of each iteration is the new endmember; when the number of endmembers meets the set value, the iteration ends, and the complete endmember matrix A is output.
[0024] Utilizing the powerful learning and data fitting capabilities of deep neural networks, an abundance matrix learning network was designed based on an L-layer fully connected network. To meet the constraint conditions of the linear mixing model of hyperspectral images, the Softmax activation function was used to constrain the values of each element in the column vectors of the abundance matrix to be between 0 and 1 while making their sum equal to 1, respectively satisfying the non-negativity constraint and the constraint that the sum of abundances is 1; the abundance matrix S was extracted i (i = 1, 2) can be expressed as:
[0025] S i = σ(f L (···f l (···f 1 (Y i ))), i = 1, 2
[0026] Among them, f j (·) represents the function of the j-th layer composed of a fully connected layer and a ReLU activation function, and σ(·) represents the Softmax activation function.
[0027] Furthermore, the specific steps of S102 are as follows:
[0028] To obtain the optimal abundance matrix, the endmember matrix and two abundance matrices were used to reconstruct the two-temporal hyperspectral images respectively. When the reconstructed image Y i (i = 1, 2) and the original image Y i (i = 1, 2) have the smallest difference, the optimal abundance matrix learning network is obtained. Therefore, the L1 norm is used as the loss function L1 to evaluate the optimization degree of the abundance matrix:
[0029]
[0030] Among them, Y i = AS i , i = 1, 2;
[0031] The parameters in the abundance matrix learning network are updated through gradient backpropagation. When the number of iterations is greater than the set value, the iteration ends and the abundance matrix S that best fits the original image is output i .
[0032] Furthermore, the specific steps of S103 are as follows:
[0033] The shallow features of the abundance matrix can be extracted and expressed as:
[0034] S shai = g shallow (S i ), i = 1, 2
[0035] Among them, g shallow(·) It includes a convolutional layer with a convolutional kernel size of 3×3 and a ReLu activation function to extract shallow features;
[0036] The query (queries, Q), key (keys, K), and value (values, V) matrices corresponding to the shallow features are obtained by a convolutional layer with a convolutional kernel size of 3×3. To improve the feature expression ability, features are extracted from multiple dimensions to improve the performance of the module. The query matrix, key matrix, and value matrix are projected into different subspaces to implement the multi-head self-cross hybrid attention mechanism, as shown in the following formula:
[0037]
[0038] where h ∈ [1, ···, H] represents the h-th head, and H represents the number of projection spaces. represents the projection matrix. represents the convolution with a convolutional kernel of 3×3.
[0039] Get Q i,h , K i,h , V i,h After that, the matrix is input into the multi-head self-cross hybrid attention mechanism composed of the self-cross hybrid attention mechanism to extract differential information.
[0040] Further, the specific content of step S104 is as follows:
[0041] The self-matching result of Q i , K i in the single-temporal hyperspectral image is called the self-correlation matrix (SCM), and the cross-matching result of Q1 and K2 between the double-temporal images is called the cross-correlation matrix (CCM), which can be calculated by the following formula:
[0042] SCM i = Q i K i T , i = 1, 2
[0043] CCM = Q1K2 T
[0044] The stronger the correlation between the abundance matrices, the higher the similarity between the double-temporal images; according to this theory, the closer the values of SCM i and CCM are, the higher the similarity between the double-temporal images; calculate the difference between SCM i and CCM, that is, the correlation difference between the two images, to obtain the weight of the image features, and use a fixed value d kScale the difference value to make the gradient more stable. The obtained result is normalized by the Softmax function to obtain the attention mechanism weights, and then the weight matrix is multiplied by the V matrix to suppress the invariant information, obtaining distinguishable change information; the single-head self-cross hybrid attention mechanism for dual-temporal images is defined as:
[0045]
[0046]
[0047] where σ i (·)(i = 1, 2) represents the Softmax activation function, and d k represents the dimension of the key matrix (Key);
[0048] The SCM and CCM calculated by the multi-head self-cross hybrid attention mechanism of the h-th head are:
[0049] SCM i,h = Q i,h K i,h T , i = 1, 2
[0050] CCM h = Q 1,h K 2,h T
[0051] The difference information mapping finally output by the multi-head self-cross hybrid attention mechanism is obtained by combining the calculation results of each single-head attention mechanism. The calculation process for Y1 is expressed as:
[0052]
[0053] where
[0054]
[0055] where F1 represents the result corresponding to the input of image Y1 to the multi-head self-cross hybrid attention mechanism, D k represents the fixed value selected to maintain the gradient stability, and W1 O represents the projection matrix learned by the 1×1 convolutional kernel;
[0056] Similarly, the result corresponding to the input of image Y2 to the multi-head self-cross hybrid attention mechanism can be expressed as:
[0057]
[0058] where
[0059]
[0060] Further, the specific implementation of step S105 is as follows:
[0061] The multi-level attention mechanism is stacked to form a difference information extraction module, and the whole process is expressed as:
[0062] [G1, G2] = g t (g t-1 ···(g 1 (S sha1 , S sha2 )))
[0063] where G1 and G2 represent two enhanced difference feature maps output by the multi-level attention mechanism for the shallow features, and g t (·) represents the t-th layer of the multi-head self-cross hybrid attention mechanism.
[0064] Further, the specific implementation of step S106 is as follows:
[0065] The two enhanced difference feature maps G1 and G2 are concatenated and then input into the fully connected layer, and the Softmax activation function is used to constrain the classification result in the probability space. The obtained result can identify whether the sub-pixels have changed, and the final change feature map can be expressed as:
[0066] CD = g cd (Concat(G1, G2))
[0067] where g cd (·) is a classifier composed of a fully connected network;
[0068] To make the predicted value as close as possible to the label value, a loss function is used to measure the difference between the predicted value and the ground-truth label, so as to further optimize the model. The cross-entropy loss function used in this method is:
[0069]
[0070] Advantages of the present invention:
[0071] 1. The present invention proposes an abundance matrix correlation analysis network, which effectively weakens the influence of mixed pixels and fully utilizes the correlation between the two-temporal images.
[0072] 2. The present invention proposes an abundance matrix learning network based on endmember sharing, which maps the change information of the two-temporal images to the corresponding abundance matrix while automatically learning the abundance matrix.
[0073] 3. The present invention proposes a multi-level multi-head self-cross hybrid attention mechanism, which weights image features by using the correlation differences between dual-temporal images, thereby gradually suppressing invariant information and retaining changing information, significantly improving the change detection accuracy. Description of the Drawings
[0074] Figure 1 is a flowchart of the hyperspectral image change detection method provided by an embodiment of the present invention.
[0075] Figure 2 is a schematic diagram of the abundance matrix learning network and the extraction of the endmember matrix provided by an embodiment of the present invention.
[0076] Figure 3 is a schematic diagram of the extraction of the difference feature map by the multi-head self-cross attention mechanism provided by an embodiment of the present invention.
[0077] Figure 4 is a schematic diagram of the single-head self-cross attention mechanism provided by an embodiment of the present invention.
[0078] Figure 5 is a comparison of the change detection result maps provided by an embodiment of the present invention.
[0079] Figure 5 Among them: (a)-(h) are the Ground-truth standard map, the result map of the CVA method, the result map of the PCA method, the result map of the IR-MAD method, the result map of the SVM method, the result map of the ReCNN method, the result map of the BCNNs method, and the result map of the present invention in sequence. Detailed Embodiment
[0080] The present invention will be further described in detail below with reference to the drawings.
[0081] Refer to Figures 1 - 5 , a hyperspectral image change detection method based on multi-head self-cross hybrid attention, and the present invention will be described in detail below with reference to the drawings.
[0082] As Figure 1 shown, the hyperspectral image change detection method based on multi-head self-cross hybrid attention provided by the present invention includes the following steps:
[0083] S101: Input the dual-temporal hyperspectral images into the vertex component analysis method and the abundance matrix learning network respectively, extract the endmember matrix, and learn the corresponding abundance matrix S based on endmember sharing i ;
[0084] S102: Reconstruct the image according to the extracted abundance matrix S i and the endmember matrix, and continuously reduce the difference between the reconstructed image and the original image to optimize the abundance matrix learning network;
[0085] S103: Extract the shallow features of the abundance matrix S in S101, project them into different subspaces through a convolutional layer after partitioning, and obtain the corresponding Q i , K i,h , and V i,h matrices, and then input them into the multi-head self-cross hybrid attention mechanism; i,h After obtaining the matrix, input it into the multi-head self-cross hybrid attention mechanism;
[0086] S104: The multi-head self-cross hybrid attention mechanism weights its input using the correlation difference, and outputs a difference information map that suppresses the invariant information of the two temporal images while retaining the changed information;
[0087] S105: Through the stacking of multiple multi-head self-cross hybrid attention mechanisms, that is, the multi-level multi-head self-cross hybrid attention mechanism, continuously strengthen the difference information;
[0088] S106: Conduct supervised training on the difference information extraction module composed of the multi-level multi-head self-cross hybrid attention mechanism, determine the optimal network parameters, and then output the predicted change detection map. The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0089] As Figure 1 shown, the hyperspectral image change detection method based on multi-head self-cross hybrid attention provided by the embodiment of the present invention is implemented as follows:
[0090] (1) As Figure 2 shown, input the dual-temporal hyperspectral images into the vertex component analysis method and the abundance matrix learning network respectively, extract the endmember matrix, and learn the corresponding abundance matrix S based on endmember sharing i ;
[0091] The linear mixing model of hyperspectral images indicates that the pixels in hyperspectral images are linear combinations of endmembers and abundances, and its matrix form can be expressed as:
[0092]
[0093] where represents the hyperspectral image, B is the number of spectral bands, M is the number of pixels, represents the endmember matrix, represents the abundance matrix, N represents the number of endmembers, represents the additive noise, is the constraint that the sum of the column vectors of the abundance matrix is 1, and 1 N is an all-ones vector of N×1; this constraint condition ensures the sparsity of the abundance matrix, better reflects the contribution degree of each endmember to the pixel spectrum, and at the same time reduces the computational complexity;
[0094] Let Denote the two - temporal - phase hyperspectral images obtained at different times in the same region, where \(H\) and \(W\) represent the height and width of the image respectively; \(M = H\times W\) represents the number of image pixels. We assume that the hyperspectral images obtained at different times in the same scene share the end - member matrix;
[0095] This method first combines the two input hyperspectral images in the spatial dimension into Then, the vertex component analysis method is used to extract the end - member matrix of the combined image. This algorithm iteratively performs projections along the direction perpendicular to the subspace formed by the known end - members. The extreme value of each projection iteration is the new end - member. When the number of end - members meets the set value, the iteration ends, and the complete end - member matrix \(A\) is output;
[0096] Utilizing the powerful learning and data - fitting capabilities of the deep neural network, an abundance matrix learning network is designed based on the \(L\) - layer fully - connected network. To meet the constraint conditions of the hyperspectral image linear mixing model, the Softmax activation function is used to constrain the values of the elements in the column vectors of the abundance matrix to be between 0 and 1 and their sum to be 1, respectively satisfying the non - negative constraint and the constraint that the sum of abundances is 1; the abundance matrix \(S\) is extracted i (\(i = 1,2\)) can be expressed as:
[0097] \(S\) i =\(\sigma(f\) L (\cdots f\) l (\cdots f\) 1 (Y\) i ))), \(i = 1,2\)
[0098] where \(f\) j (\cdot)\) represents the function of the \(j\) - th layer composed of a fully - connected layer and a ReLU activation function, and \(\sigma(\cdot)\) represents the Softmax activation function.
[0099] (2) Reconstruct the image according to the extracted abundance matrix \(S\) i and the end - member matrix, and continuously reduce the difference between the reconstructed image and the original image to optimize the abundance matrix learning network;
[0100] (2a) To obtain the optimal abundance matrix, the two - temporal - phase images are reconstructed using the end - member matrix and the abundances respectively. When the difference between the reconstructed image \(Y\) i (\(i = 1,2\)) and the original image \(Y\) i (\(i = 1,2\)) is the smallest, the optimal abundance matrix learning network is obtained. Therefore, the \(L1\) norm is used as the loss function \(L1\) to evaluate the optimization degree of the abundance matrix:
[0101]
[0102] where \(Y\) i =AS\) i , \(i = 1,2\);
[0103] (2b) Update the parameters in the abundance matrix learning network through gradient backpropagation. When the number of iterations is greater than the set value, end the iteration and output the abundance matrix S that best fits the original image. i .
[0104] (3) As Figure 3 shown, after partitioning the abundance matrix S of S101, extract its shallow features, project them into different subspaces through a convolutional layer, and obtain the corresponding Q i , K i,h , i,h , V i,h matrices, and then input them into the multi-head self-cross hybrid attention mechanism. Specifically:
[0105] (3a) The extraction of shallow features of the abundance matrix can be expressed as:
[0106] S shai = g shallow (S i ), i = 1, 2
[0107] where g shallow (·) includes a convolutional layer with a convolutional kernel size of 3×3 and a ReLu activation function to extract shallow features;
[0108] (3b) The query (queries, Q), key (keys, K), and value (values, V) matrices corresponding to the shallow features are obtained by a convolutional layer with a convolutional kernel size of 3×3. To improve the feature expression ability, features are extracted from multiple dimensions to improve the performance of the module. The query matrix, key matrix, and value matrix are projected into different subspaces to implement the multi-head self-cross hybrid attention mechanism, as shown in the following formula:
[0109]
[0110] where h ∈ [1, ···, H] represents the h-th head, H represents the number of projection spaces, represents the projection matrix, represents the convolution with a convolutional kernel of 3×3.
[0111] After obtaining Q i,h , K i,h , V i,h matrices, input them into the multi-head self-cross hybrid attention mechanism composed of the self-cross hybrid attention mechanism to extract differential information.
[0112] (4) As Figure 4 shown, the multi-head self-cross hybrid attention mechanism weights the image features using the correlation difference and outputs a differential information map that suppresses the invariant information of the two-temporal-phase images while retaining the changing information. Specifically:
[0113] (4a) The self-matching results of Q i and K i in the single-temporal hyperspectral image are called the self-correlation matrix (SCM), and the cross-matching results of Q1 and K2 between the double-temporal images are called the cross-correlation matrix (CCM), which can be calculated by the following formula:
[0114] SCM i = Q i K i T , i = 1, 2
[0115] CCM = Q1K2 T
[0116] (4b) The stronger the correlation between the abundance matrices, the higher the similarity between the double-temporal images; according to this theory, the closer the values of SCM i and CCM are, the higher the similarity between the double-temporal images; calculate the difference between SCM i and CCM, that is, the correlation difference between the two images, to obtain the weight of the image features, and use a fixed value d k to scale this difference to make the gradient more stable. The obtained result is normalized by the Softmax function to obtain the attention mechanism weight, and then the weight matrix is multiplied by the V matrix to suppress the invariant information to obtain the change information that is easy to distinguish; the single-head self-cross hybrid attention mechanism of the double-temporal image can be defined as:
[0117]
[0118]
[0119] where σ i (·)(i = 1, 2) represents the Softmax activation function, and d k represents the dimension of the key matrix.
[0120] (4c) The SCM and CCM calculated by the multi-head self-cross hybrid attention mechanism of the h-th head are:
[0121] SCM i,h = Q i,h K i,h T , i = 1, 2
[0122] CCM h = Q 1,h K 2,h T
[0123] (4d) The differential information map finally output by the multi-head self-cross hybrid attention mechanism is obtained by combining the calculation results of each single-head attention mechanism. The calculation process for Y1 can be expressed as:
[0124]
[0125] where
[0126]
[0127] where F1 represents the result corresponding to the input of image Y1 to the multi-head self-cross hybrid attention mechanism, and D k represents a fixed value selected to maintain gradient stability, and W1 O represents the projection matrix learned by a 1×1 convolutional kernel;
[0128] (4f) Similarly, the result corresponding to the input of image Y2 to the multi-head self-cross hybrid attention mechanism can be expressed as:
[0129]
[0130] where
[0131]
[0132] (5) Through the stacking of multiple multi-head self-cross hybrid attention mechanisms, that is, the multi-level multi-head self-cross hybrid attention mechanism, the differential information is continuously strengthened. Specifically:
[0133] The multi-level attention mechanism stacking forms a differential information extraction module, and its entire process can be expressed as:
[0134] [G1, G2] = g t (g t-1 ···(g 1 (S sha1 , S sha2 )))
[0135] where G1 and G2 represent the differential feature maps output by the multi-level attention mechanism for the shallow features, and g t (·) represents the t-th multi-head self-cross hybrid attention mechanism.
[0136] (6) The differential information extraction module composed of the multi-level multi-head self-cross hybrid attention mechanism is trained in a supervised manner. After determining the optimal network parameters, the predicted change detection map is output. Specifically:
[0137] The two enhanced difference feature maps G1 and G2 are concatenated and then input into a fully connected layer, and the Softmax activation function is used to constrain the classification results in the probability space. The obtained results can identify whether sub-pixels have changed. The final change feature map can be expressed as:
[0138] CD = g cd (Concat(G1, G2))
[0139] where g cd (·) is a classifier composed of a fully connected network.
[0140] To make the predicted value as close as possible to the label value, a loss function is used to measure the difference between the predicted value and the ground-truth label, thereby further optimizing the model. The cross-entropy loss function used in this method is:
[0141]
[0142] The technical effects of the present invention will be described in detail below in combination with simulation experiments:
[0143] 1. Simulation experiment conditions and data set:
[0144] The hardware platform for the simulation experiment of the present invention is: NVIDIA AGTX 3090 GPU
[0145] The software platform for the simulation experiment of the present invention is: Linux 18.06 operating system, python 3.7 and pytorch 1.12.
[0146] The Farmland data set obtained by the EO-1 satellite is used in the simulation experiment of the present invention. This data set represents the changes in farmland near Yancheng City, Jiangsu Province, China. The two images were obtained on March 3, 2006 and April 23, 2007 respectively, and their resolutions are both 420×140. In addition, the hyperspectral data sets of the two time phases contain 242 bands, the spectral range is 0.4 to 2.5 μm, and the spatial resolution is 30 m. After denoising and preprocessing, the remaining 154 bands can be used for change detection. Among all the labeled pixels, 20% of the samples are randomly selected as the training set, and the remaining 80% are used as the test set.
[0147] 2. Evaluation indicators
[0148] Two indicators are selected to evaluate the change detection results, namely the overall accuracy OA and the kappa coefficient KC. OA represents the proportion of correctly classified samples in the total samples, and KC represents the consistency between the predicted value and the reference value. The OA and KC values are positively correlated with the accuracy of the change detection results.
[0149] 3. Comparative algorithms and result analysis
[0150] To verify the effectiveness of the proposed method, seven representative methods were selected for comparative experiments, including CVA, PCA, IR-MAD, SVM, ReCNN, and BCNNs.
[0151] The detection results obtained by different methods are as Figure 5 shown. The changed regions generated on this dataset are relatively concentrated and the shapes are relatively regular. Observing Figure 5 (b) and Figure 5 (c), it can be obtained that the CVA and PCA methods cannot well detect the main changed regions. There is a large amount of noise in the results of CVA, and the most obvious changed region is not fully detected. As a dimensionality reduction algorithm, PCA inevitably leads to the loss of useful information. As Figure 5 (d) shows, compared with the first two methods, IR-MAD improves the detection ability of significantly changed regions, but there are still many unchanged regions misdetected as changed regions. SVM reduces the generation of noise, but the missed detection rate is still very high. As Figure 5 (f) and Figure 5 (g) show, the detection results of the two deep learning-based methods have been greatly improved, but the boundary regions are relatively rough. It can be obtained from the supervisor's results that the detection results obtained by the method of the present invention are the most accurate.
[0152] The quantitative evaluation of the experimental results is shown in Table 1. The overall accuracy OA of the first three classical methods CVA, PCA, and IR-MAD is all lower than 0.9. Although the overall accuracy OA of SVM is greater than 0.9, its Kappa coefficient KC is lower than 0.8 like the first three unsupervised algorithms. It can be seen from this that the traditional classical methods cannot well detect the changed regions. The overall accuracy OA and Kappa coefficient KC of the two deep learning-based methods have been significantly improved. Compared with the machine learning-based SVM method, the KC of the ReCNN method has increased by 19.5%, and the result of BCNNs has increased by nearly 20%. As shown in bold in Table 1, the method of the present invention has the most superior performance compared with other comparative methods, with OA reaching 0.9799 and KC reaching 0.9533.
[0153] Table 1 Quantitative evaluation of the method of the present invention and other representative methods
[0154]
[0155] The above are only the preferred embodiments of the present invention and do not impose any form of limitation on the present invention. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to equivalent embodiments by using the above-disclosed technical content within the scope of the technical solution of the present invention. However, as long as it does not depart from the content of the technical solution of the present invention, any brief modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. Hyperspectral image change detection method based on multi-head self-cross hybrid attention, characterized in that Including the following steps; S101: Input the dual-temporal hyperspectral images into the vertex component analysis method and the abundance matrix learning network respectively, extract the endmember matrix, and learn the corresponding abundance matrix S based on endmember sharing i ; S102: Reconstruct an image based on the extracted abundance matrix S i and the endmember matrix, and continuously reduce the difference between the reconstructed image and the original image to optimize the abundance matrix learning network; S103: For the abundance matrix S of S101 i After partitioning, extract its shallow features, project them onto different subspaces through a convolutional layer, and obtain the corresponding Q i,h , K i,h , V i,h matrices, and then input them into the multi-head self-cross hybrid attention mechanism; S104: The multi-head self-cross hybrid attention mechanism weights its input using the correlation difference, and outputs a difference information map that suppresses the invariant information of the two-phase images while retaining the changing information; S105: Through the stacking of multiple multi-head self-cross hybrid attention mechanisms, namely the multi-level multi-head self-cross hybrid attention mechanism, the difference information is continuously strengthened; S106: The difference information extraction module composed of the multi-level multi-head self-cross hybrid attention mechanism is supervised and trained. After determining the optimal network parameters, the predicted change detection map is output.
2. The hyperspectral image change detection method based on multi-head self-cross hybrid attention according to claim 1, wherein The specific steps of S101 are as follows: The hyperspectral image is divided into a linear combination of an endmember matrix and an abundance matrix, and its matrix form is expressed as: wherein represents a hyperspectral image, B is the number of spectral bands, M is the number of pixels, represents an endmember matrix, represents an abundance matrix, N represents the number of endmembers, represents additive noise, is the constraint that the sum of the column vectors of the abundance matrix is 1, and 1 N is an all-ones vector of N×1; Let denote the dual-temporal hyperspectral images obtained at different times in the same region, where H and W represent the height and width of the image respectively; M = H × W represents the number of image pixels; Hyperspectral images obtained at different times in the same scene will share an endmember matrix. First, the two input hyperspectral images are merged in the spatial dimension into Then, the vertex component analysis method is used to extract the endmember matrix of the merged image. This algorithm iteratively performs projections along the direction perpendicular to the subspace formed by the known endmembers. The extreme value of each projection iteration is the new endmember. When the number of endmembers meets the set value, the iteration ends, and the complete endmember matrix A is output; An abundance matrix learning module is designed based on an L-layer fully connected network. To meet the constraint conditions of the linear mixing model of hyperspectral images, the Softmax activation function is used to constrain the values of the elements in the column vectors of the abundance matrix to be between 0 and 1 and at the same time make their sum equal to 1, respectively meeting the non-negativity constraint and the constraint that the sum of abundances is 1; the abundance matrix S is extracted i (i = 1, 2) can be expressed as: S i = σ(f L (···f l (···f 1 (Y i )))), i = 1, 2 Among them, f j (·) represents the function of the j-th layer composed of a fully connected layer and a ReLU activation function, and σ(·) represents the Softmax activation function.
3. The hyperspectral image change detection method based on multi-head self-cross hybrid attention according to claim 1, wherein The specific steps of S102 are as follows: To obtain the optimal abundance matrix, the endmember matrix and abundances are used to reconstruct the bi-temporal images respectively. When the reconstructed image Y i (i = 1, 2) has the smallest difference from the original image Y i (i = 1, 2), the optimal abundance matrix learning network is obtained. Therefore, the L1 norm is adopted as the loss function L1 to evaluate the optimization degree of the abundance matrix: where Y i = AS i , i = 1, 2; Update the parameters in the abundance matrix learning network through gradient backpropagation. When the number of iterations is greater than the set value, end the iteration and output the abundance matrix S that best fits the original image. i .
4. The hyperspectral image change detection method based on multi-head self-cross hybrid attention according to claim 1, wherein The specific steps of S103 are as follows: The shallow features of the abundance matrix can be expressed as: S shai = g shallow (S i ), i = 1, 2 where g shallow (·) includes a convolutional layer with a convolutional kernel size of 3×3 and a ReLu activation function to extract shallow features; The query, key, and value matrices corresponding to the shallow features are obtained by a convolutional layer with a convolutional kernel size of 3×3. To improve the feature expression ability, features are extracted from multiple dimensions to improve the performance of the module. The query matrix (queries, Q), key matrix (keys, K), and value matrix (values, V) are projected into different subspaces to implement the multi-head self-cross hybrid attention mechanism, as shown in the following formula: where \(h\in[1,\cdots,H]\) represents the \(h\)-th head, and \(H\) represents the number of projection spaces. represents the projection matrix. represents a convolution with a convolution kernel of \(3\times3\). Obtain Q i,h , K i,h , V i,h After that, the matrix inputs them into the multi-head self-cross hybrid attention mechanism composed of the self-cross hybrid attention mechanism to extract differential information.
5. The hyperspectral image change detection method based on multi-head self-cross hybrid attention according to claim 1, characterized in that The specific steps of S104 are as follows: The self-matching results of Q i and K i in the single-temporal hyperspectral image are called the self-correlation matrix (SCM). The cross-matching results of Q1 and K2 between the double-temporal images are called the cross-correlation matrix (CCM), which can be calculated by the following formula: SCM i = Q i K i T , i = 1, 2 CCM = Q1K2 T The stronger the correlation between the abundance matrices, the higher the similarity between the dual-temporal images; according to this theory, SCM i and the closer the values of CCM are, the higher the similarity between the dual-temporal images; calculate SCM i The difference between CCM, that is, the correlation difference between the two images, is obtained to get the weight of the image features. Use a fixed value d k Scale this difference to make the gradient more stable. The obtained result is normalized by the Softmax function to get the attention mechanism weight, and then the weight matrix is multiplied by the V matrix to suppress the invariant information, obtaining the change information that is easy to distinguish; the single-head self-cross hybrid attention mechanism of the dual-temporal images is defined as: where σ i (·)(i = 1, 2) represents the Softmax activation function, d k represents the dimension of the key matrix (Key); The SCM and CCM calculated by the multi-head self-cross hybrid attention mechanism of the h-th head are: SCM i,h = Q i,h K i,h T , i = 1, 2 CCM h = Q 1,h K 2,h T The difference information map finally output by the multi-head self-cross hybrid attention mechanism is obtained by combining the calculation results of each single-head attention mechanism. The calculation process for Y1 is expressed as: F1 = MultiHead1(Q1, K1, V1, K2) = Concat(head 1,1 , ··· head 1,h , ··· head 1,H )W1 O where Among them, F1 represents the result corresponding to the input of image Y1 into the multi-head self-cross hybrid attention mechanism, D k represents a fixed value selected to maintain gradient stability, W1 O represents a projection matrix learned by a convolutional kernel of size 1×1; Similarly, the result corresponding to the input of the multi-head self-cross hybrid attention mechanism for the image Y2 can be expressed as: F2 = MultiHead2(Q2, K2, V2, K1) = Concat(head 2,1 , ··· head 2,h , ··· head 2,H )W2 O where 6. The hyperspectral image change detection method based on multi-head self-cross hybrid attention according to claim 1, characterized in that The specific steps of S105 are as follows: The multi-level attention mechanism is stacked to form a difference information extraction module, and its entire process is expressed as: [G1,G2] = g t (g t-1 ···(g 1 (S sha1 ,S sha2 ))) Among them, G1 and G2 represent two enhanced differential feature maps output by the shallow features through the multi-level attention mechanism, and g t (·) represents the multi-head self-cross hybrid attention mechanism of the t-th layer.
7. The hyperspectral image change detection method based on multi-head self-cross hybrid attention according to claim 1, wherein The specific steps of S106 are as follows: The two enhanced difference feature maps G1 and G2 are concatenated and then input into the fully connected layer, and the Softmax activation function is used to constrain the classification result in the probability space. The obtained result can identify whether the sub-pixel has changed. The final change feature map can be expressed as: CD = g cd (Concat(G1, G2)) where g cd (·) is a classifier composed of a fully connected network; To make the predicted value as close as possible to the label value, a loss function is used to measure the difference between the predicted value and the ground-truth label, thereby further optimizing the model. The cross-entropy loss function used in this method is:
Citation Information
Patent Citations
End-to-end hyperspectral image change detection method based on depth convolution neural network
CN109102529A
Remote sensing image weak and small target detection method based on form and multi-example learning
CN113887652A