A Hyperspectral Anomaly Detection Method Based on Siamese Graph Attention Encoding
Through the combination of the spatial-spectral feature dynamic composition module and the twin graph attention automatic encoder, the problem of inaccurate construction of hyperspectral image background is solved, and more efficient abnormal detection effect is achieved.
Patent Information
- Application Number
- CN202310650917.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-03
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-06-03
AI Technical Summary
In the absence of prior information, the construction accuracy of hyperspectral image background is low, and the spectral classification characteristics of a few cells are not obvious, resulting in poor reconstruction effect and reducing the accuracy of detection.
The dynamic composition module design of the space-spectral feature is adopted, and the pseudo-label is obtained using the DBSCAN algorithm and the graph structure data set is constructed through ERS. Combined with the twin graph attention automatic encoder, network training and abnormal detection are performed, and the network is optimized through node feature reconstruction loss, graph structure reconstruction loss and comparison loss, and abnormal detection is performed using Mahayana distance.
It improves the accuracy of background distribution estimation, purifies the reconstruction of background, improves detection efficiency and accuracy, and effectively solves the problem of low background construction accuracy.
Smart Images

Figure CN116597231B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hyperspectral anomaly detection, and specifically to a hyperspectral anomaly detection method based on twin graph attention coding. Background Technique
[0002] The anomaly detection technology of hyperspectral images has become an important research content in the field of intelligent processing and analysis of remote sensing data because it can locate abnormal targets in images by analyzing spectral characteristic differences under the condition of lacking prior information of unknown ground objects, providing a reference for subsequent accurate identification and analysis of targets.
[0003] Currently, a large number of research works have been dedicated to developing various efficient and accurate anomaly detection algorithms and technologies, such as anomaly detection algorithms based on statistical methods, representation models, and deep neural networks. With the complexity of scenarios and applications, the background of hyperspectral images has become complex, which limits the detection performance of traditional methods such as statistical methods and representation models. The method based on generative spectral reconstruction has become the mainstream research direction, but most of these methods only utilize the spectral information of hyperspectral data and ignore the spatial information, resulting in poor expression ability of the network. In addition, the spectral classification characteristics of a few pixels are not obvious, causing the reconstruction network to learn abnormal expression methods, resulting in poor reconstruction effects. These problems have greatly reduced the detection accuracy.
[0004] Based on the above, the present invention proposes a hyperspectral anomaly detection method based on twin graph attention coding to solve the above problems. Summary of the Invention
[0005] 1. Technical Problems to be Solved by the Invention
[0006] The purpose of the present invention is to propose a hyperspectral anomaly detection method based on twin graph attention coding to solve problems such as low accuracy of background construction and unclear spectral classification characteristics of a few pixels in the case of lacking prior information.
[0007] 2. Technical Solution
[0008] To achieve the above purpose, the present invention provides the following technical solution:
[0009] A hyperspectral anomaly detection method based on twin graph attention coding specifically includes the following:
[0010] S1. Design of Spatial-Spectral Feature Dynamic Composition Module:
[0011] ① Unsupervised Background Category Search: Use the DBSCAN algorithm to obtain pseudo-labels of image pixels without prior information, and divide the original pixels into "anomaly", "background", and "uncertain" pixels;
[0012] ② Construction of spatial spectral feature map dataset:
[0013] 1) Use ERS to segment the original data into superpixels, take each superpixel region as a node in the graph structure, and establish the adjacency relationship between superpixels;
[0014] 2) Based on the representation method described in step 1), convert the hyperspectral data into graph structure data to achieve more accurate feature embedding;
[0015] 3) Combine superpixels with abnormal pixels. If the number of abnormal pixels in a superpixel exceeds half of the number of pixels in the superpixel, mark the superpixel as an abnormal superpixel, otherwise mark it as a normal superpixel; use the random walk algorithm to sample superpixels on the graph to form a graph dataset;
[0016] S2. Design of Siamese graph attention autoencoder:
[0017] Use the graph attention autoencoder network as the backbone of the Siamese network. The encoder takes two graph attention autoencoders with the same parameters as the main body, and measures the similarity of latent space features between the latent spaces encoded in the middle of the two autoencoders; during training, the two graph attention encoders are randomly input with a type of graph data each time, and selective background reconstruction is performed according to the labels of the graph data; after each training, the two graph autoencoders synchronously update the same network parameters;
[0018] S3. Network training and anomaly detection:
[0019] Use the reconstruction loss of node features between input and output, the reconstruction loss of the graph structure, and the contrast loss in the feature expression differentiation mechanism to train the network. After training, obtain the output of the network, perform a difference operation on the output and the original data, and use the Mahalanobis distance to perform anomaly detection on the difference image to obtain the final detection map.
[0020] Preferably, the unsupervised background category search described in S1 specifically includes the following content:
[0021] Use the DBSCAN algorithm to obtain the initial classification of the original pixels, divide the pixels into two clusters of "background" and "anomaly", and calculate the silhouette coefficient of each pixel respectively. The specific calculation formula is as follows:
[0022]
[0023] Where SC(i) is the pixel silhouette coefficient. The closer SC(i)∈[-1, 1] is to 1, the more reasonable the clustering of sample i is. The closer SC(i) is to -1, the more sample i needs to be classified into another cluster. If SC(i) is approximately 0, it means that sample i is on the boundary of two clusters. Pixels with SC(i)∈[-0.1, 0.1] are selected to construct an "uncertain" cluster. The "uncertain" cluster pixels do not participate in reconstruction and feature separation.
[0024] Preferably, the construction of the spatial-spectral feature map dataset in S1 specifically includes the following:
[0025] A1. ERS is used to obtain densely homogeneous and similarly sized uniform regions to learn the intrinsic low-dimensional features of different regions of hyperspectral data, thereby obtaining superpixels that represent the spatial-spectral information of the image. Each superpixel region is used as a node in the graph structure, and the adjacency relationship between superpixels is established.
[0026] A2. Based on the representation described in A1, convert the hyperspectral data into graph-structured data to achieve more accurate feature embedding. Specifically, the following contents are included:
[0027] For the hyperspectral image X, entropy rate superpixel segmentation is performed, and the superpixel segmentation problem is defined as a graphics optimization problem. The image is mapped to a graph G = (V, E), where V represents the vertex set of the graph and E represents a series of vertices v i and vertex v j The edge between i,j The edge set composed of i,j Represents vertex v i and v j The similarity of + {0}Calculated; then the graph partitioning problem is transformed into a subset selection problem, using the objective function to find a set of edges Make the graph G'=(V, A) contain all connected subgraphs; finally, remove the edges in the original edge set E through the objective function to achieve the purpose of image segmentation;
[0028] In ERS, the random walk entropy rate H(A) on the graph G'=(V, A) is calculated as shown in formula (2):
[0029]
[0030] Where, is the stationary distribution of random walk, ω i represents all the nodes passing through vertex v i The sum of the edge weights, is the regularization constant, where |V| is the number of vertices in the graph; p i,j(A) is the probability of random walk, and its calculation process is shown in formula (3):
[0031]
[0032] The objective function of ERS satisfies the relationship shown in formula (4) with the random walk entropy rate term H(A) and the balance term B(A):
[0033]
[0034] Where λ≥0 is a variable weight factor to coordinate the proportional relationship between the random walk entropy rate term H(A) and the balance term B(A), and the superpixel segmentation of the image is realized by maximizing the objective function;
[0035] Taking the first principal component map as the base image for superpixel segmentation, the first principal component F of the hyperspectral image X is extracted by PCA, and ERS superpixel segmentation is performed on F through formula (5):
[0036]
[0037] Where y is the number of superpixels, and X k (1≤k≤y) is the k-th superpixel;
[0038] For the superpixel regions representing the spatial-spectral information of the image, each superpixel region X i is regarded as a graph node, the adjacency relationship between superpixels is established, and the HSI is converted into an undirected graph G=(F, ε), where F is the vertex set representing the obtained superpixel set; ε represents the edges between superpixels and is represented by an adjacency matrix; the adjacency matrix of graph G is defined as:
[0039]
[0040] A3. Combine the pseudo-labels obtained by unsupervised background class search with the graph structure data. In the normal pixel set and the abnormal pixel set, respectively use the random walk algorithm to construct several randomly generated graph data and send them into the subsequent reconstruction network;
[0041] For each graph structure data, if the corresponding number of abnormal pixels exceeds 50% of the number of pixels contained in the superpixel, the superpixel is marked as an abnormal superpixel, otherwise it is marked as a normal superpixel; use the random walk algorithm to sample graph structure data on graph G to construct a dataset where Nw represents the number of sampled random walk sequences; s represents the number of steps of random walk, and each superpixel selected in each random walk is the same as the first sampled superpixel.
[0042] Preferably, in S2, the graph attention autoencoder network uses a self-attention mechanism to learn the relationships and representations between nodes in the graph; the design of the graph attention encoder network specifically includes the following:
[0043] ① Determine the importance of each node by calculating the attention scores between each node and its neighbor nodes, and use the attention scores to perform weighted averaging on the nodes to generate the representation of the nodes;
[0044] ② During the background estimation process, the graph dataset obtained by the spatio-spectral feature dynamic graph construction module is sent into the encoder of S2BAE to obtain the latent feature embedding vector. At the k-th encoder layer, the correlation between adjacent node j and node i is shown in formula (7):
[0045]
[0046] where σ(·) represents the sigmoid activation function, are the trainable parameters of the k-th encoder layer, and respectively represent the node features of node i and node j in the (k - 1)-th encoder layer; for the first layer, the node feature is the input sample, i.e., Subsequently, obtain the normalized k-th attention coefficient:
[0047]
[0048] where NBR i represents the neighbors of node i, that is, a set of nodes connected to node i according to the adjacency matrix A m including node i itself; generate the embedding of node i in the k-th encoder layer
[0049]
[0050] After passing through L encoder layers, take the output of the last layer as the final node representation, i.e.,
[0051] ③ Construct a latent feature comparator between the outputs of the two-way encoders to perform similarity measurement on the encoded latent feature embedding vectors; finally, send the latent feature embedding vectors into the decoder to obtain the reconstructed attribute information and connection information.
[0052] ④ Use a decoder with the same number of layers as the encoder. Each decoder layer is the reverse of the structure of its corresponding encoder layer, that is, each decoder layer reconstructs the representation of a node using the representations of its neighbors according to the correlation of the nodes; in the k-th decoder layer, the attention coefficient of adjacent node j to node i is defined as follows:
[0053]
[0054] Wherein, σ(·) represents the sigmoid activation function, are the trainable parameters of the k-th encoder layer, and respectively represent the node features of node i and node j in the k-th encoder layer; the output of the last encoder layer is the input of the decoder, that is, where L is the number of encoder layers;
[0055] The normalized attention coefficient between adjacent nodes j and i in the k-th decoder layer is calculated as follows:
[0056]
[0057] where NBR i is the neighbor of node i, and the reconstructed node representation of the decoder at the (k - 1)-th layer is as follows:
[0058]
[0059] After passing through L decoder layers, the output of the last decoder layer is used as the reconstructed node attribute, that is,
[0060] Preferably, the siamese network described in S2 is a neural network structure for calculating the similarity or distance between two inputs; the specific design of the siamese network includes the following:
[0061] The siamese network accepts a sample pair composed of two samples as input, and inputs them into the feature extraction network respectively, mapping the inputs into two feature vectors G w (x1), G w (x2) in a new space, where the two feature networks GATE1 and GATE2 share weights, and the specific formula is expressed as follows:
[0062]
[0063] where x1 and x2 represent two inputs; f(·) represents the function of the sub-network; θ represents the weight parameter of the sub-network; g represents the distance metric function between features; D represents the output distance.
[0064] Preferably, S3 specifically includes the following:
[0065] S3.1. Minimize the reconstruction loss of node features, and define the reconstruction loss of node features as the l2 loss between the input of the encoder and the output of the decoder:
[0066]
[0067] Among them, ||·||2 represents the l2 norm;
[0068] S3.2. Measure whether the representations of the reconstructed adjacent nodes in S3.1 are similar, and define the reconstruction loss of the graph structure as follows:
[0069]
[0070] Among them, σ(·) represents the sigmoid activation function;
[0071] S3.3. Use the contrast loss in the feature expression differentiation mechanism to measure whether two similar samples in the feature space are still similar and whether two dissimilar samples are still dissimilar after feature extraction. The specific contrast loss is:
[0072]
[0073] Among them, d is the Euclidean distance between two embedding vectors; y represents whether the two samples have the same label. y = 1 means the two samples have the same label, and y = 0 means the labels are different; margin is the set threshold;
[0074] S3.4. Combining the contents of S3.1 to S3.3, the overall loss function of the network can be obtained:
[0075] L = L att + L str + L CEDM (17)
[0076] In the formula, L att represents the reconstruction loss of node features; L str represents the reconstruction loss of the graph structure; L CEDM represents the contrast loss;
[0077] S3.5. Use the overall loss function obtained in S3.4 to train the network. After training, the output of the network is obtained. Perform a difference operation on the output and the original hyperspectral image. The specific calculation formula is:
[0078]
[0079] In the formula, X is the original hyperspectral image; is the reconstructed data obtained after the background distribution model tests the original hyperspectral image X; ΔX is the difference data;
[0080] S3.6. Based on the differential data obtained in S3.5, use Mahalanobis distance to perform anomaly detection on the differential image, calculate the anomaly score of the pixel under test, and then detect the abnormal pixels in the differential image. The calculation process is shown in formula (19):
[0081] M D (ΔX) = (Δx p -μ) T Γ -1 (Δx p -μ) (19)
[0082] In the formula, represents the p-th vector of ΔX; respectively represent the mean value and covariance matrix of ΔX.
[0083] 3. Beneficial effects
[0084] (1) The present invention proposes a spatial-spectral feature dynamic composition module, which includes unsupervised background category search and spatial-spectral feature map dataset construction. Among them, the unsupervised background category search is based on the Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm, which realizes the acquisition of pseudo-labels of image pixels without prior information, and the pixels are respectively labeled as "background", "abnormal" and "uncertain" pixels. Based on this, the accuracy of background distribution estimation is further improved in the reconstruction network; the spatial-spectral feature map dataset construction uses Entropy Rate Superpixel (ERS) segmentation to obtain superpixels representing the spatial-spectral information of the image, and each superpixel region is used as a node in the graph structure to establish the adjacency relationship between superpixels. Using this representation method, the hyperspectral data is converted into graph structure data, thus realizing more accurate feature embedding;
[0085] (2) In the reconstruction process of the present invention, it is proposed that only the data with the pseudo-label of "background" is used for background reconstruction in the siamese graph attention autoencoder, while the "abnormal" and "uncertain" do not participate in this process. Through this design, a purer reconstructed background can be obtained; at the same time, through the siamese graph attention autoencoder model, while improving the spatial perception ability of the model, the detection efficiency of the algorithm is effectively improved; finally, the Mahalanobis distance is used to calculate the anomaly score in the hyperspectral image, and the performance of anomaly detection is improved. Description of the drawings
[0086] Figure 1 is the overall framework diagram of a hyperspectral anomaly detection method based on siamese graph attention encoding proposed by the present invention;
[0087] Figure 2 This is a graph showing the anomaly detection results of the anomaly detection method proposed in this invention on five data sets. DETAILED DESCRIPTION
[0088] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0089] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0090] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.
[0091] Embodiment 1:
[0092] See also Figure 1 This paper proposes a hyperspectral anomaly detection method based on twin graph attention encoding, which mainly includes three parts: spatial-spectral feature dynamic composition module design, twin graph attention autoencoder design, and network training and anomaly detection:
[0093] Part I - Design of spatial-spectral feature dynamic composition module:
[0094] ① Unsupervised background category search: Using the DBSCAN algorithm, pseudo labels of image pixels are obtained without prior information, and the original pixels are classified into "abnormal", "background" and "uncertain" pixels;
[0095] ② Construction of spatial spectral feature map dataset:
[0096] 1) Use ERS to split the original data into superpixels, treat each superpixel region as a node in the graph structure, and establish the adjacency relationship between superpixels;
[0097] 2) Based on the representation method described in step 1), the hyperspectral data is converted into graph structure data to achieve more accurate feature embedding;
[0098] 3) Combine the superpixels with the abnormal pixels. If the number of abnormal pixels in a superpixel exceeds half of the number of pixels in that superpixel, mark the superpixel as an abnormal superpixel; otherwise, mark it as a normal superpixel. Use the random walk algorithm to sample superpixels on the graph to form a graph dataset;
[0099] The second part - the design of the Siamese graph attention autoencoder:
[0100] Adopt the graph attention autoencoder network as the backbone of the Siamese network. The encoder takes two graph attention autoencoders with the same parameters as the main body, and measures the similarity of latent space features between the latent spaces encoded in the middle of the two autoencoders. During the training process, the two graph attention encoders are randomly input with a type of graph data each time, and selective background reconstruction is performed according to the labels of the graph data. After each training, the two graph autoencoders synchronously update the same network parameters;
[0101] The third part - network training and anomaly detection:
[0102] Use the reconstruction loss of node features between input and output, the reconstruction loss of the graph structure, and the contrast loss in the feature expression differentiation mechanism to train the network. After the training is completed, obtain the output of the network, perform a difference operation on the output and the original data, and use the Mahalanobis distance to perform anomaly detection on the difference image to obtain the final detection graph.
[0103] More specifically, it includes the following content:
[0104] A. Spatial-spectral feature dynamic graph construction module
[0105] In order to construct a reliable background spectral graph dataset without prior information, the present invention proposes a spatial-spectral feature dynamic graph construction module, which includes unsupervised background class search and spatial-spectral feature graph dataset construction.
[0106] 1) Unsupervised background class search
[0107] Unsupervised background class search is to obtain the pseudo-labels of image pixels based on the DBSCAN algorithm without prior information.
[0108] DBSCAN is a density-based clustering algorithm. The core idea of this algorithm is to determine the category through the sample distribution density. In hyperspectral images, abnormal targets have the characteristics of sparsity and low probability distribution. The inter-class difference between abnormal samples and background samples is much larger than the intra-class difference of background samples. This data property enables the DBSCAN algorithm to classify background classes and abnormal classes in hyperspectral images.
[0109] The initial classification of the original pixels is obtained using the DBSCAN algorithm. The pixels are divided into two clusters: "background" and "anomaly". The Silhouette Coefficient of each pixel is calculated, as shown in formula (1).
[0110]
[0111] In the formula, SC(i) is the pixel silhouette coefficient. The closer SC(i) ∈ [-1, 1] is to 1, the more reasonable the clustering of sample i is. The closer SC(i) is to -1, the more sample i should be classified into another cluster. If SC(i) is approximately 0, it means that sample i is on the boundary of the two clusters. Pixels with SC(i) ∈ [-0.1, 0.1] are selected to construct an "uncertain" cluster. The pixels included in this cluster are difficult to classify because their spectral characteristics are somewhat similar to those of the other two clusters. Therefore, they should be excluded during the background reconstruction process to establish a more "pure" reconstructed background. At the same time, due to the classification uncertainty of the pixels in the "uncertain" cluster, if their feature representation is separated from the background during the training process, there is also a risk of increasing the false detection rate of anomalies. Therefore, the pixels in the "uncertain" cluster do not participate in the reconstruction and also do not participate in the feature separation.
[0112] 2) Construction of the spatial-spectral feature map dataset
[0113] Hyperspectral images usually contain a large amount of spectral information, which is likely to increase the computational complexity during node embedding and classification. And in hyperspectral anomaly detection tasks, most previous algorithms ignore the local structure information of hyperspectral data and cannot capture the correlation between hyperspectral data, resulting in a decline in detection performance. Superpixel segmentation is an effective method for generating homogeneous regions and extracting spatial information. Pixels within a homogeneous region are more likely to have the same spectral characteristics. It makes each superpixel region have a smaller spatial range and a certain spatial continuity, which helps to detect anomalies more accurately.
[0114] Therefore, the present invention uses ERS to obtain tightly homogeneous and similarly sized uniform regions to learn the intrinsic low-dimensional features of different regions of hyperspectral data, thereby obtaining superpixels that represent the spatial-spectral information of the image. Each superpixel region is used as a node in the graph structure, and the adjacency relationship between superpixels is established. Using this representation, hyperspectral data is converted into graph structure data, thereby achieving more accurate feature embedding.
[0115] Specifically, entropy rate superpixel segmentation is performed on the hyperspectral image X. This method defines the superpixel segmentation problem as an optimization problem in graphics, mapping the image to a graph G = (V, E), where V represents the vertex set of the graph and E represents a series of edges e i between vertex v j and vertex vi,j The set of edges that makes up the graph, where the weight of each edge is ω i,j Denote vertex v i and v j The similarity between vertices, which is calculated by the weight function ω: E → + {0}; then transform the graph partitioning problem into a subset selection problem, and use the objective function to find a set of edges such that the graph G’=(V, A) contains all connected subgraphs; finally, remove the edges in the original edge set E through the objective function to achieve the purpose of image segmentation. In ERS, the calculation method of the graph random walk entropy rate H(A) for the graph G’=(V, A) is shown in formula (2).
[0116]
[0117] In the formula, is the stationary distribution of the random walk, ω i represents the sum of the weights of all edges passing through vertex v i ; is the regularization constant, where |V| is the number of vertices in the graph; p i , j (A) is the probability of the random walk, and its calculation process is shown in formula (3).
[0118]
[0119] The objective function of ERS, the random walk entropy rate term H(A), and the balance term B(A) satisfy the relationship shown in formula (4).
[0120]
[0121] In the formula, λ≥0 is a variable weight factor, which is used to coordinate the proportional relationship between the random walk entropy rate term H(A) and the balance term B(A), and realizes the superpixel segmentation of the image by maximizing the objective function.
[0122] Since the first principal component in the hyperspectral image covers the main information in the image, the present invention uses the first principal component image as the basic image for superpixel segmentation, extracts the first principal component F of the hyperspectral image X by PCA, and performs ERS superpixel segmentation on F through formula (5):
[0123]
[0124] In the formula, y is the number of superpixels, and X k (1≤k≤y) is the k-th superpixel.
[0125] For the superpixel regions that represent the spatial-spectral information of the image, for each superpixel region X iRegard it as a graph node, establish the adjacency relationship between superpixels, and convert HSI into an undirected graph G=(F, ε), where F is the vertex set, representing the set of obtained superpixels; ε represents the edges between superpixels, which is represented by an adjacency matrix; define the adjacency matrix of graph G as:
[0126]
[0127] The construction of the graph dataset combines the pseudo-labels obtained from unsupervised background class search with the graph structure data. In the normal pixel set and the abnormal pixel set, the random walk algorithm is used to construct several randomly generated graph data, and then they are fed into the subsequent reconstruction network.
[0128] For each graph structure data, if the corresponding number of abnormal pixels exceeds 50% of the number of pixels contained in this superpixel, then we mark this superpixel as an abnormal superpixel, otherwise mark it as a normal superpixel. Use the random walk algorithm to sample graph structure data on graph G to construct the dataset where Nw represents the number of random walk sequences sampled; s represents the number of steps of the random walk, and the superpixel selected each time during the random walk is the same as the first sampled superpixel.
[0129] B. Siamese Graph Attention Encoder
[0130] The present invention applies the graph attention autoencoder network as the backbone of the siamese network. Figure 1 Part B specifically shows the overall architecture of the selective siamese background encoder. This encoder takes two graph attention autoencoders with the same parameters as the main body, and measures the similarity of the latent space features between the latent spaces encoded in the middle of the two autoencoders. During the training process, the two graph attention encoders are randomly input with a type of graph data each time, and selective background reconstruction is performed according to the label of the graph data. After each training, the two graph autoencoders synchronously update the same network parameters.
[0131] 1). Graph Attention Autoencoder
[0132] The graph self-attention network uses the self-attention mechanism to learn the relationships and representations between nodes in the graph. Specifically, the graph attention network determines the importance of each node by calculating the attention scores between each node and its neighbor nodes, and uses these scores to perform weighted averaging on the nodes to generate the representation of the nodes. This method enables the graph self-attention network to dynamically model the relationships between different nodes and shows good generalization ability when dealing with different graph structures.
[0133] During the background estimation process, the present invention sends the graph data set obtained by the spatio-spectral feature dynamic composition module into the encoder of S2BAE to obtain the latent feature embedding vector. At the k-th encoder layer, the correlation between adjacent node j and node i is shown in Equation (7).
[0134]
[0135] In the formula, σ(·) represents the sigmoid activation function. are the trainable parameters of the k-th encoder layer. and represent the node features of node i and node j in the (k - 1)-th encoder layer respectively; for the first layer, the node feature is the input sample, that is Subsequently, the normalized k-th attention coefficient is obtained.
[0136]
[0137] In the formula, NBR i represents the neighbors of node i, that is, a group of nodes connected to node i according to the adjacency matrix A m including node i itself; generate the embedding of node i in the k-th encoder layer
[0138]
[0139] After passing through L encoder layers, the output of the last layer is used as the final node representation, that is
[0140] A latent feature comparator is constructed between the outputs of the two-way encoders to perform similarity measurement on the encoded latent feature embedding vectors. Finally, the latent feature embedding vectors are sent into the decoder to obtain the reconstructed attribute information and connection information. During the entire reconstruction process, in order to obtain a purer reconstructed background, only the data with the pseudo-label "background" is used for background reconstruction, while "anomaly" and "uncertain" do not participate in this process.
[0141] We use a decoder with the same number of layers as the encoder. Each decoder layer is the reverse of the structure of its corresponding encoder layer, that is, each decoder layer reconstructs the representation of a node by using the representations of its neighbors according to the correlation of the nodes; in the k-th decoder layer, the attention coefficient of adjacent node j to node i is defined as follows:
[0142]
[0143] In the formula, σ(·) represents the sigmoid activation function. is the trainable parameter of the k-th encoder layer, and respectively represent the node features of node i and node j in the k-th encoder layer; the output of the last encoder layer is the input of the decoder, that is where L is the number of encoder layers.
[0144] The normalized attention coefficient between adjacent nodes j and i in the k-th decoder layer is calculated as follows:
[0145]
[0146] where NBR i is the neighbor of node i, and the reconstructed node representation of the decoder in the (k - 1)-th layer is as follows:
[0147]
[0148] After passing through L decoder layers, the output of the last decoder layer is used as the reconstructed node attribute, that is
[0149] 2) Siamese network
[0150] A siamese network is a neural network structure mainly used to calculate the similarity or distance between two inputs. In hyperspectral anomaly detection, there are often a small amount of abnormal data and a large amount of normal data, so it belongs to the problem of small sample learning. The shared weight mechanism of the siamese network can effectively improve the generalization ability of the model, thus avoiding the overfitting problem caused by insufficient data volume. The siamese network can also capture the similarities and differences in hyperspectral data, thereby improving the accuracy of anomaly detection. During the training process, the siamese network will try to map similar data points to close positions in the space and dissimilar data points to far positions in the space, making it easier to detect abnormal data points.
[0151] The siamese network takes a sample pair composed of two samples as input, and inputs them into the feature extraction network respectively, mapping the inputs into two feature vectors G w (x1), G w (x2) in the new space, where the two feature networks (graph attention autoencoder networks) GATE1 and GATE2 share weights.
[0152]
[0153] where x1 and x2 represent the two inputs; f(·) represents the function of the sub-network; θ represents the weight parameter of the sub-network; g represents the distance metric function between features; D represents the output distance.
[0154] C. Training Loss and Anomaly Detection
[0155] The graph-structured data includes node features and connection information. We first minimize the reconstruction loss of the node features, where the reconstruction loss is defined as the l2 loss between the input of the encoder and the output of the decoder, as follows:
[0156]
[0157] where ||·||2 represents the l2 norm. In addition, it is also necessary to measure whether the representations of the reconstructed adjacent nodes are similar, and the reconstruction loss of the graph structure is defined as follows:
[0158]
[0159] where σ(·) represents the sigmoid activation function.
[0160] The role of the contrastive loss is to make the distance between similar samples as small as possible, and the distance between dissimilar samples as large as possible, so as to improve the accuracy of the siamese network. The following contrastive loss metric is used to measure whether two similar samples are still similar and whether two dissimilar samples are still dissimilar in the feature space after feature extraction:
[0161]
[0162] where d is the Euclidean distance between two embedding vectors; y indicates whether the two samples have the same label, y = 1 means the two samples have the same label, and y = 0 means the labels are different; margin is the set threshold.
[0163] The overall loss function of the network is
[0164] L = L att + L str + L CEDM (17)
[0165] During the testing process, there may be a very small number of abnormal pixels in the above-mentioned constructed background spectral samples. Directly using the reconstruction error as the anomaly detection result cannot obtain the optimal detection effect. According to the design principle of the background representation network and the observation of the spectral characteristics and spatial characteristics of the reconstructed hyperspectral image, it is found that the background pixels of the reconstructed hyperspectral image are almost exactly the same as those of the original hyperspectral image, while the abnormal pixels are quite different, and the distribution of abnormal pixels tends to the background distribution. Therefore, the difference between the original hyperspectral image and the reconstructed hyperspectral image can be used to suppress the background of the hyperspectral image, so as to achieve the purpose of separating the background and anomalies. In addition, the reconstructed hyperspectral image is more in line with the distribution of background data. By using the difference between the original hyperspectral image and the reconstructed image, the shape consistency of abnormal targets can be guaranteed. Formula (18) represents the calculation process of the difference.
[0166]
[0167] In the formula, X is the original hyperspectral image; is the reconstructed data obtained after the background distribution model tests the original hyperspectral image X; ΔX is the difference data.
[0168] After obtaining the difference data with a significant difference between the background and anomalies, the present invention uses the Mahalanobis distance to calculate the anomaly score of the pixel to be tested, and then detects the abnormal pixels in the difference image. The calculation process is shown in formula (19):
[0169] M D (ΔX)=(Δx p -μ) T Γ -1 (Δx p -μ)(19)
[0170] In the formula, represents the p-th vector of ΔX; respectively represent the average value and covariance matrix of ΔX.
[0171] In summary, the present invention proposes a twin graph attention encoding hyperspectral anomaly detection network. In the twin graph attention encoding hyperspectral anomaly detection network, in order to construct a reliable background spectral map dataset in the absence of prior information, the present invention proposes a spatio-spectral feature dynamic graph construction module, which includes unsupervised background class search and spatio-spectral feature map dataset construction. The unsupervised background class search is based on the Density-Based Spatial Clustering of Applications with Noise (DBSCAN) algorithm to obtain pseudo-labels of image pixels without prior information, and the pixels are respectively labeled as "background", "anomaly" and "uncertain" pixels. Based on this, the accuracy of background distribution estimation is further improved in the reconstruction network; the spatio-spectral feature map dataset construction uses Entropy Rate Superpixel (ERS) segmentation to obtain superpixels representing the spatio-spectral information of the image, and each superpixel region is used as a node in the graph structure to establish the adjacency relationship between superpixels. Using this representation, the hyperspectral data is converted into graph structure data, thereby achieving more accurate feature embedding; and the pseudo-labels are combined with the graph structure data, and in the normal pixel set and the abnormal pixel set, several randomly generated graph datasets are constructed using the random walk algorithm and sent to the subsequent reconstruction network.
[0172] Meanwhile, during the background estimation process, the present invention applies a graph attention autoencoder network as the backbone of the twin network, transforming hyperspectral anomaly detection into a similarity discrimination problem under the twin network framework in small sample learning. Specifically, the graph dataset obtained based on the spatio-spectral feature dynamic graph construction module is sent into the twin graph attention autoencoder to obtain latent feature embedding vectors. A latent feature comparator is constructed between the outputs of the two encoders to measure the similarity of the encoded latent feature embedding vectors. Finally, the latent feature embedding vectors are sent into the decoder to obtain the reconstructed attribute information and connection information. During the entire reconstruction process, in order to obtain a purer reconstructed background, the algorithm proposes to use only the data with the pseudo-label of "background" for background reconstruction in the twin graph attention autoencoder, while "anomaly" and "uncertain" do not participate in this process. Through the twin graph attention autoencoder model, while improving the spatial perception ability of the model, the detection efficiency of the algorithm is effectively improved. Finally, the Mahalanobis distance is used to calculate the anomaly score in the hyperspectral image to improve the performance of anomaly detection.
[0173] As described above, it is only the preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and its improved concept of the present invention, making equivalent substitutions or changes, shall be covered by the protection scope of the present invention.
Claims
1. A hyperspectral anomaly detection method based on twin graph attention encoding, characterized in that, Specifically, it includes the following content: S1. Design of Spatial-Spectral Feature Dynamic Composition Module: ① Unsupervised Background Category Search: Using the DBSCAN algorithm, pseudo-labels of image pixels are obtained without prior information, and the original pixels are divided into "anomaly", "background", and "uncertain" pixels; ② Construction of Spatial-Spectral Feature Map Dataset: 1) The ERS is used to segment the original data into superpixels. Each superpixel region is used as a node in the graph structure, and the adjacency relationship between superpixels is established; 2) Based on the representation method in step 1), the hyperspectral data is converted into graph structure data to achieve more accurate feature embedding; 3) The superpixels are combined with the anomaly pixels. If the number of anomaly pixels in a superpixel exceeds half of the number of pixels in the superpixel, the superpixel is marked as an abnormal superpixel, otherwise it is marked as a normal superpixel; the random walk algorithm is used to sample superpixels on the graph to form a graph dataset; S2. Design of Siamese Graph Attention Autoencoder: The graph attention autoencoder network is used as the backbone of the siamese network. The encoder takes two graph attention autoencoders with the same parameters as the main body, and the similarity of latent space features is measured between the latent spaces encoded in the middle of the two autoencoders; During the training process, the two graph attention encoders are randomly input with a class of graph data each time, and selective background reconstruction is performed according to the labels of the graph data; after each training, the two graph autoencoders synchronously update the same network parameters; S3. Network Training and Anomaly Detection: The network is trained using the reconstruction loss of node features between input and output, the reconstruction loss of the graph structure, and the contrast loss in the feature expression differentiation mechanism. After the training is completed, the output of the network is obtained. The output is subjected to a difference operation with the original data, and the Mahalanobis distance is used to perform anomaly detection on the difference image to obtain the final detection map.
2. The hyperspectral anomaly detection method based on twin graph attention encoding according to claim 1, wherein The specific content of the unsupervised background category search in S1 is as follows: The DBSCAN algorithm is used to obtain the initial classification of the original pixels, and the pixels are divided into two clusters of "background" and "anomaly", and the silhouette coefficient of each pixel is calculated respectively. The specific calculation formula is as follows: In the formula, SC(i) is the silhouette coefficient of the pixel. The closer SC(i) is to 1 ∈ [-1, 1], the more reasonable the clustering of sample i is. The closer SC(i) is to -1, the more sample i needs to be classified into another cluster. If SC(i) is approximately 0, it means that sample i is on the boundary of the two clusters; pixels with SC(i) ∈ [-0.1, 0.1] are selected to construct an "uncertain" cluster, and the pixels in the "uncertain" cluster do not participate in reconstruction and also do not participate in feature separation.
3. A hyperspectral anomaly detection method based on twin graph attention encoding according to claim 1, characterized in that, The specific content of the construction of the spatial-spectral feature map dataset in S1 is as follows: A1. The ERS is used to obtain tightly homogeneous and similarly sized uniform regions to learn the intrinsic low-dimensional features of different regions of hyperspectral data, so as to obtain superpixels representing the spatial-spectral information of the image. Each superpixel region is used as a node in the graph structure, and the adjacency relationship between superpixels is established; A2. Based on the representation method in A1, the hyperspectral data is converted into graph structure data to achieve more accurate feature embedding. The specific content is as follows: For the hyperspectral image X, perform entropy rate superpixel segmentation. Define the superpixel segmentation problem as an optimization problem in graphics. Map the image to a graph G=(V, E), where V represents the vertex set of the graph and E represents a set of edges e i between vertex v j and vertex v i,j that make up the edge set. The weight ω i,j of the edge represents the similarity between vertex v i and v j , which is calculated by the weight function ω:E→+{0}; then convert the graph segmentation problem into a subset selection problem, and use the objective function to find a set of edges such that the graph G'=(V, A) contains all connected subgraphs; finally, achieve the purpose of segmenting the image by removing the edges in the original edge set E through the objective function; In the ERS, the calculation method of the entropy rate H(A) of the random walk on the graph of the graph G'=(V, A) is shown in formula (2): In the formula, is the stationary distribution of the random walk, and ω i represents the sum of the weights of all edges passing through vertex v i . is the regularization constant, where |V| is the number of vertices of the graph; p i,j (A) is the probability of the random walk, and its calculation process is shown in formula (3): The objective function of the ERS, the random walk entropy rate term H(A), and the balance term B(A) satisfy the relationship shown in formula (4): In the formula, λ≥0 is a variable weight factor to coordinate the proportional relationship between the random walk entropy rate term H(A) and the balance term B(A), and the superpixel segmentation of the image is realized by maximizing the objective function; Taking the first principal component graph as the base image for superpixel segmentation, the first principal component F of the hyperspectral image X is extracted by PCA, and the ERS superpixel segmentation is performed on F through formula (5): where y is the number of superpixels, and X k (1 ≤ k ≤ y) is the k-th superpixel; For the superpixel regions that characterize the spatial-spectral information of the image, each superpixel region X i is regarded as a graph node, and the adjacency relationship between superpixels is established. The HSI is converted into an undirected graph G=(F, ε), where F is the vertex set, representing the set of obtained superpixels; ε represents the edges between superpixels, which is represented by an adjacency matrix; define the adjacency matrix of graph G as: A3. Combine the pseudo-labels obtained by unsupervised background class search with the graph structure data. In the normal pixel set and the abnormal pixel set, respectively use the random walk algorithm to construct several randomly generated graph data and send them into the subsequent reconstruction network; For each graph structure data, if the corresponding number of abnormal pixels exceeds 50% of the number of pixels contained in the superpixel, the superpixel is marked as an abnormal superpixel; otherwise, it is marked as a normal superpixel. The random walk algorithm is used to sample the graph structure data on graph G to construct a dataset. Among them, Nw represents the number of sampled random walk sequences; s represents the number of steps of the random walk. Each superpixel selected in each random walk is the same as the first sampled superpixel.
4. A hyperspectral anomaly detection method based on twin graph attention encoding according to claim 1, characterized in that, The graph attention autoencoder network described in S2 uses the self-attention mechanism to learn the relationships and representations between the nodes in the graph; the specific design of the graph attention encoder network includes the following: ① Determine the importance of each node by calculating the attention scores between each node and its neighbor nodes, and use the attention scores to perform weighted averaging on the nodes to generate the representation of the nodes; ② During the background estimation process, send the graph data set obtained by the spatio-spectral feature dynamic graph construction module into the encoder of the S2BAE to obtain the latent feature embedding vector. At the kth encoder layer, the correlation between the adjacent node j and the node i is shown in formula (7): where, σ(·) represents the sigmoid activation function, are the trainable parameters of the k-th encoder layer, and respectively represent the node features of node i and node j in the (k-1)-th encoder layer; for the first layer, the node features are the input samples, i.e., Subsequently, the normalized k-th attention coefficient is obtained: where, NBR i represents the neighbors of node i, that is, according to the adjacency matrix A m a set of nodes connected to node i, including node i itself; Generate the embedding of node i at the k-th encoder layer After passing through the L encoder layers, the output of the last layer is used as the final node representation, i.e., ③ Construct a latent feature comparator between the outputs of the two-path encoders to perform similarity measurement on the encoded latent feature embedding vectors; finally, send the latent feature embedding vectors into the decoder to obtain the reconstructed attribute information and connection information; ④ Use a decoder with the same number of layers as the encoder. Each decoder layer is the inversion of the structure of its corresponding encoder layer, that is, each decoder layer reconstructs the representation of the node using the representations of its neighbors according to the correlation of the nodes; in the kth decoder layer, the attention coefficient of the adjacent node j to the node i is defined as follows: where σ(·) represents the sigmoid activation function, are the trainable parameters of the k-th encoder layer, and represent the node features of nodes i and j in the k-th encoder layer, respectively; the output of the last encoder layer is the input of the decoder, that is where L is the number of encoder layers; The normalized attention coefficient between the adjacent node j and the node i in the kth layer decoder is calculated as follows: Among them, NBR i is the neighbor of node i, and the reconstructed node representation of the decoder at the (k - 1)-th layer is as follows: After passing through the L decoder layers, the output of the last decoder layer is used as the reconstructed node attributes, that is 5. A hyperspectral anomaly detection method based on twin graph attention encoding according to claim 1, characterized in that The siamese network described in S2 is a neural network structure used to calculate the similarity or distance between two inputs; the specific design of the siamese network includes the following: The Siamese network takes a pair of samples composed of two samples as input, and inputs them into the feature extraction network respectively. The inputs are mapped into two feature vectors G w (x1), G w (x2) in a new space. Among them, the two feature networks GATE1 and GATE2 share weights, and the specific formula is expressed as follows: Among them, x1 and x2 represent two inputs; f(·) represents the function of the sub-network; θ represents the weight parameters of the sub-network; g represents the distance metric function between features; D represents the output distance.
6. The hyperspectral anomaly detection method based on twin graph attention encoding according to claim 1, wherein The specific content of S3 is as follows: S3.
1. Minimize the reconstruction loss of the node features. Define the reconstruction loss of the node features as the l2 loss between the input of the encoder and the output of the decoder: Among them, ||·||2 represents the l2 norm; S3.
2. Measure whether the representations of the reconstructed adjacent nodes in S3.1 are similar, and define the reconstruction loss of the graph structure as follows: Among them, σ(·) represents the sigmoid activation function; S3.
3. Use the contrast loss in the feature expression differentiation mechanism to measure whether two similar samples in the feature space are still similar and whether two dissimilar samples are still dissimilar after feature extraction. The specific contrast loss is as follows: where d is the Euclidean distance between two embedding vectors; y represents whether the two samples have the same label, y = 1 means the two samples have the same label, and y = 0 means the labels are different; margin is the set threshold; S3.
4. Combining the content described in S3.1 to S3.3, the overall loss function of the network can be obtained: L = L att + L str + L CEDM (17) where, L att represents the reconstruction loss of node features; L str represents the reconstruction loss of the graph structure; L CEDM represents the contrastive loss; S3.
5. Use the overall loss function obtained in S3.4 to train the network. After training, the output of the network is obtained. Perform a difference operation on the output and the original highlight map image. The specific calculation formula is: Where X is the original hyperspectral image; is the reconstructed data obtained after the background distribution model tests the original hyperspectral image X; ΔX is the differential data; S3.
6. Based on the difference data obtained in S3.5, use the Mahalanobis distance to perform anomaly detection on the difference image, calculate the anomaly score of the tested pixel, and then detect the abnormal pixels in the difference image. The calculation process is shown in formula (19): M D (ΔX) = (Δx p -μ) T Γ -1 (Δx p -μ) (19) In the formula, represents the p-th vector of ΔX; represent the mean value and covariance matrix of ΔX respectively.
Citation Information
Patent Citations
Hyperspectral anomaly detection method for component projection optimization separation
CN110599466A
Hyperspectral remote sensing image change detection method
CN114359735A