Hyperspectral anomaly detection method based on sample similarity guidance and spatial spectrum mask auto-encoder

By using sample similarity guidance and spatial spectral mask autoencoder methods in hyperspectral anomaly detection, the problem of limited detection accuracy in complex backgrounds is solved, and more accurate background recovery and higher anomaly detection accuracy is achieved.

CN120219916APending Publication Date: 2025-06-27XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510290817.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing hyperspectral anomaly detection methods perform poorly in complex backgrounds, and the impact of abnormal pixels on background recovery is inevitable, resulting in limited detection accuracy.

Method used

Using a method based on sample similarity guidance and spatial spectral mask autoencoder, the influence of abnormal samples on background recovery is reduced through spatial-spectral masks, and multi-scale residual convolution blocks are designed on the latent layer features to mine the multi-scale features of the background.

Benefits of technology

Effectively alleviate the impact of abnormal pixels on background reconstruction, improve abnormal detection accuracy, and accurately restore the background through sample similarity guidance to improve detection effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219916A_ABST
    Figure CN120219916A_ABST
Patent Text Reader

Abstract

The invention discloses a hyperspectral anomaly detection method based on sample similarity guidance and a spatial spectrum mask auto-encoder, and the method comprises the steps: firstly, reducing abnormal pixels through a space-spectrum mask strategy in order to alleviate the influence of the abnormal pixels on background recovery, thereby avoiding the interference on subsequent feature extraction; besides, due to the complexity of the background, a multi-scale residual error convolution block is designed and is used for extracting the multi-scale features of the background. A space-spectrum mask strategy can lose a part of background information, so that the background is difficult to recover accurately in a decoding process. Therefore, the decoder is guided to decode by calculating the local similar features of the unmasked data blocks, so that the background is effectively reconstructed. And finally, performing anomaly detection on the residual error of the reconstructed background and the input image through the mahalanobis distance. According to the invention, the background can be effectively reconstructed, so that a good detection effect is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of hyperspectral anomaly detection, and particularly relates to a hyperspectral anomaly detection method based on sample similarity guidance and spatial-spectral masked autoencoder. Background Art

[0002] Hyperspectral anomaly detection is an important research direction in hyperspectral image analysis, and it is widely applied in fields such as geological exploration, battlefield reconnaissance, and maritime search and rescue. In the past few decades, researchers have developed a large number of anomaly detection algorithms. According to the basic principles of these algorithms, they can be roughly divided into statistical-based detection methods, representation-based detection methods, and deep learning-based detection methods.

[0003] Statistical-based anomaly detection algorithms usually assume that the background follows a specific distribution and perform anomaly detection through distance metrics. Among them, the most representative algorithm is the RX detector. The RX detector assumes that the background follows a multivariate normal distribution, and then performs anomaly detection by calculating the Mahalanobis distance between the pixel to be detected and the background. In addition, many improved algorithms improve the anomaly detection accuracy by making more reasonable distribution assumptions, estimating a purer background, and through kernel mapping, etc. Many studies have shown that due to the complex actual background distribution, a single distribution assumption is difficult to accurately model the actual background, so statistical-based algorithms perform poorly in complex backgrounds.

[0004] Different from statistical-based anomaly detection algorithms, representation-based anomaly detection algorithms represent the background and then use the reconstruction error for anomaly detection. Li et al. introduced collaborative representation into anomaly detection. Since there are obvious spectral differences between anomaly pixels and the surrounding background, the background pixels can be represented by the surrounding pixels, and the anomaly pixels will generate a large representation error. In addition, Cheng et al. performed anomaly detection through a graph structure and total variation constraints. This method can capture the local geometric structure, thereby improving the anomaly detection accuracy. However, representation-based methods usually require setting a large number of parameters. For different scenarios, their optimal parameters often vary greatly, which limits their practical applications.

[0005] In recent years, many deep learning-based algorithms have been successively proposed. Wang et al. designed the Auto-AD algorithm, which extracts features through multiple layers of convolution and performs downsampling, and then decodes. In addition, they constrained the loss function through the reconstruction error to suppress the reconstruction of anomaly pixels while restoring the background. Fan et al. made the AE able to capture the local features of the HSI through superpixel processing, and trained the network using the L (2,1) norm to avoid the influence of noise and anomalies on background restoration.

[0006] Although the above methods have achieved good detection accuracy, most algorithms process on the original images, making it difficult to avoid the influence of abnormal pixels on background restoration. In addition, due to the complex background and a large amount of noise in practical applications, existing methods still have deficiencies in background restoration, thus limiting the abnormal detection accuracy. Summary of the Invention

[0007] To overcome the above disadvantages of the prior art, the purpose of the present invention is to provide a hyperspectral anomaly detection method based on sample similarity guidance and spatial-spectral mask autoencoder. Through the spatial-spectral mask strategy, the influence of abnormal samples on subsequent background restoration is reduced. In addition, to solve the problem of inaccurate restoration caused by complex background, a multi-scale residual convolutional block is designed on the latent features, thereby effectively mining the multi-scale features of the background. And, during the decoding process, sample similarity is used to guide background reconstruction, restoring the background while effectively suppressing abnormal samples. Finally, the final anomaly detection result is obtained by using the Mahalanobis distance on the residual image.

[0008] To achieve the above purpose, the technical solution adopted by the present invention is:

[0009] A hyperspectral anomaly detection method based on sample similarity guidance and spatial-spectral mask autoencoder, characterized by including the following steps:

[0010] Step 1, perform spatial and spectral direction mask processing on the input hyperspectral image in sequence to obtain a mask image, where the height, width, and spectral channel number of the hyperspectral image are H, W, and C respectively;

[0011] Step 2, perform a sliding window process on the mask image to generate multiple pixel blocks of size p as mask samples, and use an encoder to extract features from each mask sample to obtain feature vectors;

[0012] Step 3, use the multi-scale residual convolutional block to extract the multi-scale features of the background with the feature vectors as the input;

[0013] Step 4, calculate the similarity between the unmasked data pixels corresponding to the mask samples to obtain a similarity vector;

[0014] Step 5, use the similarity vector as the guiding feature to guide the decoder to restore the background and achieve background reconstruction;

[0015] Step 6; perform anomaly detection on the reconstructed background and the residual image of the input hyperspectral image using the Mahalanobis distance.

[0016] Compared with the prior art, the beneficial effects of the present invention are:

[0017] 1. The present invention proposes an autoencoder based on a spatial-spectral masking strategy. The spatial-spectral mask can effectively alleviate the influence of abnormal pixels on background reconstruction, thereby improving the accuracy of anomaly detection.

[0018] 2. The present invention designs a multi-scale residual convolution block, which can effectively extract multi-scale features of the background, thereby enhancing the background recovery ability of the network.

[0019] 3. The present invention designs a background recovery module guided by sample similarity. The sample similarity in the input data block is used to guide the network to recover the background, thus avoiding the interference of abnormal pixels. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is the flow chart of the algorithm of the present invention.

[0021] Figure 2 is a schematic diagram of the multi-scale residual convolution block.

[0022] Figure 3 are the detection results of each algorithm on the first dataset. From left to right and from top to bottom, they are the labels corresponding to the dataset, the detection results of the RX algorithm, the detection results of the CRD algorithm, the detection results of the GTVLRR algorithm, the detection results of the PTA algorithm, the detection results of the RGAE algorithm, the detection results of the BockNet algorithm, the detection results of the Auto-AD algorithm, the detection results of the MSNet algorithm, and the detection results of the algorithm of the present invention.

[0023] Figure 4 are the detection results of each algorithm on the second dataset. From left to right and from top to bottom, they are the labels corresponding to the dataset, the detection results of the RX algorithm, the detection results of the CRD algorithm, the detection results of the GTVLRR algorithm, the detection results of the PTA algorithm, the detection results of the RGAE algorithm, the detection results of the BockNet algorithm, the detection results of the Auto-AD algorithm, the detection results of the MSNet algorithm, and the detection results of the algorithm of the present invention.

[0024] Figure 5 are the ROC curves of each algorithm on two datasets. (a) is the ROC curve of each algorithm on the first dataset, and (b) is the ROC curve of each algorithm on the second dataset.

[0025] Figure 6 are the background-anomaly diagrams of each algorithm on three datasets. (a) is the background-anomaly diagram of each algorithm on the first dataset, and (b) is the background-anomaly diagram of each algorithm on the second dataset. DETAILED DESCRIPTION OF THE INVENTION

[0026] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0027] As described above, in the existing hyperspectral anomaly detection methods, due to the influence of factors such as anomalous pixels and noise, the effect of background restoration is poor, and the accuracy of anomaly detection is limited.

[0028] To this end, the present invention provides a hyperspectral anomaly detection method based on sample similarity guidance and spatial-spectral mask autoencoder. First, in order to alleviate the influence of anomalous pixels on background restoration, a spatial-spectral mask strategy is used to reduce anomalous pixels, thereby avoiding interference with subsequent feature extraction. Secondly, due to the complexity of the background, a multi-scale residual convolution block is designed to extract multi-scale features of the background. And, since the spatial-spectral mask strategy will lose some background information, it is difficult to accurately restore the background during the decoding process. Therefore, the present invention calculates the local similarity features of the unmasked data blocks to guide the decoder to decode, thereby effectively reconstructing the background. Finally, anomaly detection is performed on the residual between the reconstructed background and the input image through the Mahalanobis distance. Experimental results prove that the present invention can effectively reconstruct the background, thus having a good detection effect.

[0029] Refer to as Figure 1 shown, the hyperspectral anomaly detection method of the present invention based on sample similarity guidance and spatial-spectral mask autoencoder mainly includes the following steps:

[0030] Step 1, taking the hyperspectral image as the input, performing mask processing on it in the spatial and spectral directions in sequence to obtain the masked image where H, W, and C respectively represent the height, width, and number of spectral channels of the hyperspectral image. Specifically:

[0031] The masks in the spatial and spectral directions are randomly generated according to a certain proportion, and the masked areas are marked as 0, and the unmasked areas remain the same as the original. In this embodiment, the mask processing on can be expressed by the following formula:

[0032]

[0033] where represents the input hyperspectral image, represents the hyperspectral image after spatial mask processing, represents the hyperspectral image after spatial-spectral mask processing, that is, the masked image, M spa ∈ H×W and respectively represent a spatially and spectrally random mask, and ⊙ represents the pixel multiplication operation.

[0034] Step 2, perform a sliding window process on the masked image to generate multiple pixel blocks of size p as mask samples

[0035] Specifically, perform a sliding window process on the image from left to right and top to bottom to generate H×W mask samples.

[0036] Step 3, use an encoder to extract features from the mask samples to obtain feature vectors

[0037] The encoder of the present invention has a two-layer structure. Each layer includes a Conv-BN-Activation block, and the convolution kernels are all 1×1. The activation function of the first layer is LeakyReLU, the number of output channels is C / 2, the activation function of the second layer is Sigmoid, and the number of output channels is k. In this embodiment, k = 50, and the rest are the same. Then the formula is expressed as:

[0038] P1 = LeakyReLU(BN(Conv 1×1 (P m )));

[0039] P2 = Sigmoid(BN(Conv 1×1 (P1)));

[0040] where and are the output results of the first and second layers of the encoder respectively. LeakyReLU represents the LeakyReLU activation function, Sigmoid represents the Sigmoid activation function, BN represents the batch normalization operation, and Conv 1×1 represents a 1×1 convolution.

[0041] Step 4, send the feature vectors into a multi-scale residual convolution block to extract the multi-scale features of the background.

[0042] Reference Figure 2, the multi-scale residual convolution block of the present invention includes three parallel branches, each branch is a Conv-BN-Activation block, and the convolution kernels of the three branches are different. Exemplarily, the first branch consists of a 1×1 convolution, BN, and LeakyReLU; the convolution kernel sizes of the second and third branches are 3×3 and 5×5 respectively, and the rest is the same as the first branch. The three branches respectively perform feature extraction on the feature vector, and then add and fuse the extracted features. The fusion result is then passed through a Conv-BN-Activation block. Immediately afterwards, its output is added to the input feature vector to obtain the output feature, which is the multi-scale feature of the background. Exemplarily, in the Conv-BN-Activation block where the fusion result is input, the convolution kernel is 3×3 and the activation function is LeakyReLU.

[0043] The above process can be expressed as:

[0044] P o = Conv 3×3 (Conv 1×1 (P2)+Conv 3×3 (P2)+Conv 5×5 (P2))+P2

[0045] where is the output feature after passing through the multi-scale residual convolution block, that is, the multi-scale feature of the background. It should be noted that BN and LeakyReLU are omitted in the above formula.

[0046] Step 5, calculate the similarity between the unmasked data m corresponding to the masked sample P pixel by pixel to obtain the similarity vector x.

[0047] Specifically, assume that the central pixel of P ori is First, calculate the spectral similarity between the central pixel and its surrounding pixels, and use this spectral similarity as the weight of this pixel. Among them, the weight of the i-th pixel P i of the unmasked data is expressed as:

[0048]

[0049] In the above formula, the denominator represents the sum of the similarities between all pixels in the current image block and the central pixel, the numerator is the similarity between the P i pixel and the central pixel, cos() represents the cosine similarity, P j is the j-th pixel of the unmasked data, and W i is the weight of P i .

[0050] When the weights of each pixel in the unmasked data P ori are obtained, the entire unmasked data P ori is encoded, and the encoded spectral vector is the similarity vector, expressed as:

[0051] x = P ori T W

[0052] where W is the weight of all pixels in the unmasked data.

[0053] Step 6, using the similarity vector as the guiding feature to guide the decoder to recover the background and achieve background reconstruction.

[0054] First, perform feature mapping on the similarity vector .

[0055] In this embodiment, two cascaded Conv-BN-Activation blocks are used to perform feature mapping on the similarity vector. The convolution kernel of each layer is 1×1. The activation function of the first layer is LeakyReLU, and the number of output channels is C / 2. The activation function of the second layer is Sigmoid, and the number of output channels is the same as that of the output channels of the second layer of the encoder, which is k. In this embodiment, k = 50, and the rest of the composition is the same as that of the first layer. The above process can be expressed as:

[0056] x1 = LeakyReLU(BN(Conv 1×1 (x)));

[0057] x2 = Sigmoid(BN(Conv 1×1 (x1)));

[0058] where and are the output results of the first layer and the second layer networks respectively.

[0059] Secondly, multiply the result of the feature mapping by the features of the corresponding stage decoder and then add them to achieve background reconstruction.

[0060] Corresponding to the encoder, the present invention uses a decoder with a two-layer structure for decoding and reconstructing the background. The first layer of the decoder is a Conv-BN-Activation block, and the second layer is a Conv-BN block. The convolution kernel of each layer is 1×1. The activation function of the first layer is LeakyReLU, and the number of output channels is C / 2. The number of output channels of the second layer is C. The decoding process is described as follows:

[0061] 1), the multi-scale features of the extracted background With the result of the second-layer feature mapping Perform a dot product, and then add it to P o and send it to the first layer of the decoder.

[0062] 2), Take the output of the first layer of the decoder and perform a dot product with the result of the first-layer feature mapping then add it to F1 and send it to the second layer of the decoder.

[0063] 3), The output of the second layer of the decoder

[0064] The above process is expressed as:

[0065] F1 = LeakyReLU(BN(Conv 1×1 (P o ⊙ x2 + P o ))) ;

[0066] F2 = BN(Conv 1×1 (F1 ⊙ x1 + F1)) ;

[0067] F2 represents the reconstruction result of an image patch sample. In the present invention, the background is reconstructed through two layers of decoders, and all the image patch samples are reconstructed together, which is the reconstructed background, expressed as

[0068] Step 7, Use the Mahalanobis distance to perform anomaly detection on the reconstructed background and the residual image of the input hyperspectral image X. The final detection result is expressed as:

[0069]

[0070] where Mahalanobis represents the Mahalanobis distance.

[0071] The effects of the present invention will be further described below in conjunction with simulation experiments.

[0072] 1. Simulation conditions:

[0073] The hardware environment for the simulation experiment of the present invention is an Intel Core i5-12400 CPU, 16-GB random access memory (RAM), and a Microsoft Windows 10 operating system; the simulation software is: MATLAB R2022a and Python 3.7.

[0074] 2. Experimental Content: To demonstrate the effectiveness of the hyperspectral anomaly detection method based on sample similarity guidance and spatial-spectral mask autoencoder, the present invention uses two publicly available datasets to verify the effectiveness of the algorithm and for simulation experiments. The first dataset is Cat Island, which consists of 188 bands, with a spatial size of 150×150 and a spectral resolution of 17.2 m / pixel. The airplane in the upper left corner is the anomaly target. The second dataset is Pavia, with a size of 100×100 and a total of 102 bands. Multiple vehicles on the bridge are used as anomaly targets.

[0075] Figure 3 And Figure 4 are the example diagrams of the detection results of each algorithm on the above two datasets, Figure 5 and Figure 6 show the ROC curves and background-anomaly block diagrams of each algorithm on the above two datasets. Table 1 is the AUC values of different algorithms on the two datasets.

[0076] Table 1 AUC Values of Different Algorithms on Two Datasets

[0077] Algorithm The first data set The second data set RX 0.9807 0.9887 CRD 0.9871 0.9653 GTVLRR 0.9752 0.9858 PTA 0.9786 0.9791 RGAE 0.9394 0.9914 BockNet 0.9855 0.9863 AuTo-AD 0.9806 0.9879 MSNet 0.9789 0.9967 The method of the present invention 0.9994 0.9977

[0078] It can be seen that the results of the simulation experiment confirm that the present invention can achieve good detection effects both subjectively and objectively.

[0079] The above is only the preferred embodiment of the present invention and is not used to limit the protection scope of the present invention.

Claims

1. A hyperspectral anomaly detection method based on sample similarity guidance and spatial spectral mask autoencoder, characterized in that: The following steps are involved: Step 1, performing spatial and spectral mask processing on the input hyperspectral image in sequence to obtain a mask image, wherein the height, width and number of spectral channels of the hyperspectral image are H, W and C respectively; Step 2, performing sliding window processing on the mask image to generate multiple pixel blocks of size p as mask samples, and using an encoder to extract features from each mask sample to obtain a feature vector; Step 3, taking the feature vector as input, and using a multi-scale residual convolution block to extract multi-scale features of the background; Step 4, calculating the similarity between the unmasked data pixels corresponding to the masked sample to obtain a similarity vector; Step 5, using the similarity vector as a guiding feature to guide the decoder to restore the background and achieve background reconstruction; Step 6: Perform anomaly detection on the reconstructed background and the residual image of the input hyperspectral image using Mahalanobis distance.

2. According to claim 1, the method for hyperspectral anomaly detection based on sample similarity guidance and spatial spectral mask autoencoder is characterized in that: In step 1, the input hyperspectral image is subjected to spatial and spectral masking in sequence to obtain a mask image, and the implementation formula is as follows: in, represents the input hyperspectral image, represents the hyperspectral image processed by spatial mask, represents the hyperspectral image processed by spatial-spectral mask, that is, the mask image; M spa ∈ H×W and denote random masks in space and spectrum respectively, ⊙ denotes pixel multiplication operation; In the step 2, sliding window processing is performed on the mask image to generate H×W mask samples.

3. According to claim 2, the method for hyperspectral anomaly detection based on sample similarity guidance and spatial spectral mask autoencoder is characterized in that: The encoder has a two-layer structure, each layer includes a Conv-BN-Activation block, and the convolution kernel is 1×1. The activation function of the first layer is LeakyReLU, the number of output channels is C / 2, the activation function of the second layer is Sigmoid, the number of output channels is k, and the rest of the components are the same; In step 5, the decoder has a two-layer structure, the first layer is a Conv-BN-Activation block, the second layer is a Conv-BN block, the convolution kernel of each layer is 1×1, the activation function of the first layer is LeakyReLU, the number of output channels is C / 2, and the number of output channels of the second layer is C.

4. The method for hyperspectral anomaly detection based on sample similarity guidance and spatial spectral mask autoencoder according to claim 1, characterized in that: The multi-scale residual convolution block includes three parallel branches, each of which is a Conv-BN-Activation block, and the convolution kernels of the three branches are different; the three branches respectively extract features from the feature vector, and then add and fuse the extracted features. The fusion result passes through a Conv-BN-Activation block, and its output is added to the feature vector. The obtained output feature is the multi-scale feature of the background.

5. According to claim 4, the method for hyperspectral anomaly detection based on sample similarity guidance and spatial spectral mask autoencoder is characterized in that: The convolution kernels of the three branches are 1×1, 3×3 and 5×5 respectively. In the Conv-BN-Activation block input with the fusion result, the convolution kernel is 3×3, and the activation function of each Conv-BN-Activation block is LeakyReLU.

6. The method for hyperspectral anomaly detection based on sample similarity guidance and spatial spectral mask autoencoder according to claim 1, characterized in that: The step 4 is to calculate the similarity between the unmasked pixels corresponding to the masked sample to obtain a similarity vector, and the implementation method is as follows: Step 4.1, calculate the center pixel of the unmasked data The spectral similarity between the pixel and the surrounding pixels is used as the weight of the pixel; Step 4.2, after obtaining the weight of each pixel in the unmasked data, encode the entire unmasked data, and the encoded spectral vector is the similarity vector.

7. The method for hyperspectral anomaly detection based on sample similarity guidance and spatial spectral mask autoencoder according to claim 6, characterized in that: In step 4.1, the i-th pixel P of the unmasked data i The weight is expressed as: Among them, cos() represents cosine similarity, P j is the jth pixel of the unmasked data; In step 4.2, the encoded spectral vector x is expressed as: x=P ori T W In the formula, W is the weight of all pixels in the unmasked data.

8. The method for hyperspectral anomaly detection based on sample similarity guidance and spatial spectral mask autoencoder according to claim 1, characterized in that: The step 5 uses the similarity vector as a guiding feature to guide the decoder to restore the background, and the implementation method is as follows: Step 5.1, performing feature mapping on the similarity vector; Step 5.2, multiplying the result of the feature mapping with the feature of the decoder of the corresponding stage and then adding them to achieve background reconstruction.

9. The method for hyperspectral anomaly detection based on sample similarity guidance and spatial spectral mask autoencoder according to claim 8, characterized in that: In the step 5.1, a two-layer cascaded Conv-BN-Activation block is used to perform feature mapping on the similarity vector. The convolution kernel of each layer is 1×1, the activation function of the first layer is LeakyReLU, the number of output channels is C / 2, the activation function of the second layer is Sigmoid, the number of output channels is k, and the rest of the components are the same as the first layer.

10. The method for hyperspectral anomaly detection based on sample similarity guidance and spatial spectral mask autoencoder according to claim 9, characterized in that: The result of the feature mapping is multiplied by the feature of the decoder at the corresponding stage and then added. The implementation method is as follows: First, the multi-scale features of the extracted background The result of the second layer feature map Perform a dot product and then add it to P o Add them together and send them to the first layer of the decoder; Then, the output of the first layer of the decoder The result of the first layer feature map Perform point multiplication, then add it to F1 and send it to the second layer of the decoder; Finally, the decoder second layer output The background is reconstructed by a two-layer decoder and the result is expressed as