Hyperspectral anomaly detection method based on joint feature and trilateral filtering guidance

By introducing a trilateral filtering algorithm and an empty spectrum joint attention mechanism in hyperspectral anomaly detection, combined with an adaptive loss function, the problems of incomplete anomaly target suppression and insufficient background feature weights in the prior art are solved, and the accuracy and robustness of the detection are significantly improved.

CN120032245APending Publication Date: 2025-05-23XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510043019.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing hyperspectral anomaly detection algorithms have shortcomings in utilizing the spatial and spectral information of hyperspectral images, resulting in incomplete suppression of abnormal targets, insufficient background feature weights, and abnormal reconstruction phenomena are more common.

Method used

The hyperspectral anomaly detection method based on joint features and trilateral filtering guidance is adopted to build an AE network, combine trilateral filtering algorithm to background purification of the image, and use the null spectrum joint attention mechanism to improve the feature expression ability, and finally suppress abnormal reconstruction through an adaptive loss function.

Benefits of technology

It significantly improves the accuracy and robustness of abnormal detection, can more efficiently utilize the null spectral characteristics of hyperspectral images, improves detection performance, and has high practical application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032245A_ABST
    Figure CN120032245A_ABST
Patent Text Reader

Abstract

The invention discloses a hyperspectral anomaly detection method based on joint features and trilateral filtering guidance. The hyperspectral anomaly detection method comprises the following steps: constructing an AE network; carrying out background purification on a to-be-detected image through the guide module based on trilateral filtering to obtain a guide feature map; performing element-by-element multiplication on the guide feature map and the learned channel weight through an attention mechanism module based on spatial-spectral combination to obtain a feature map after channel attention weighting; multiplying the guide feature map by the learned spatial weight through an attention mechanism module based on spatial-spectral combination to obtain a feature map after spatial attention weighting; multiplying the feature map subjected to the channel attention weighting by the feature map subjected to the space attention weighting to obtain final feature representation; and performing differential detection on the final feature representation to obtain a detection result image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of hyperspectral image processing, and in particular relates to a hyperspectral anomaly detection method based on joint features and trilateral filtering guidance. Background Art

[0002] With the widespread application of hyperspectral imaging technology, hyperspectral images have shown great potential in target detection, environmental monitoring and other fields. However, existing hyperspectral anomaly detection algorithms are insufficient in utilizing the spatial and spectral information of hyperspectral images. When processing hyperspectral images, traditional fully convolutional autoencoders often fail to fully utilize their spatial-spectral characteristics, resulting in incomplete suppression of abnormal targets.

[0003] In addition, insufficient background feature weights and anomaly reconstruction are also common problems in existing technologies. For example, some anomaly detection methods based on statistical models do not work well in complex backgrounds, while deep learning models may be difficult to apply effectively due to insufficient training data or high model complexity. Therefore, a new method is urgently needed that can more efficiently utilize the spatial-spectral characteristics of hyperspectral images and improve the accuracy of anomaly detection. Summary of the invention

[0004] In view of this, the main purpose of the present invention is to provide a hyperspectral anomaly detection method based on joint features and trilateral filtering guidance.

[0005] To achieve the above object, the technical solution of the present invention is achieved as follows:

[0006] The embodiment of the present invention provides a hyperspectral anomaly detection method based on joint features and trilateral filtering guidance, the method comprising:

[0007] Constructing an AE network, wherein the AE network includes a guidance module based on trilateral filtering, an attention mechanism module based on spatial-spectral joint, and an adaptive loss function based on local features;

[0008] The guidance module based on trilateral filtering is used to purify the background of the image to be detected to obtain a guidance feature map;

[0009] The guided feature map is element-wise multiplied by the learned channel weights through an attention mechanism module based on spatial-spectral joint to obtain a feature map weighted by channel attention;

[0010] Multiplying the guided feature map with the learned spatial weights through an attention mechanism module based on spatial-spectral joint to obtain a feature map weighted by spatial attention;

[0011] Multiplying the feature map weighted by channel attention and the feature map weighted by spatial attention to obtain a final feature representation;

[0012] Perform differential detection on the final feature representation to obtain a detection result image.

[0013] In the above scheme, the guiding module based on trilateral filtering is constructed by a spatial correlation weight factor, a structural feature weight factor, and a spectral distance weight factor, specifically:

[0014] The spatial association weight factor In the formula, x is the pixel to be tested, ε is the neighboring pixel within a certain relevant distance of x, and w s (x,ε) is the weight factor based on the spatial position distance, d s (x,ε) is the distance between the neighboring pixel and the pixel to be tested, σ s Contribution parameters for distance;

[0015] The gray structure feature weight factor In the formula, x is the pixel to be tested, ε is the neighboring pixel within a certain relevant distance of x, and w g (x,ε) is the weight factor based on the grayscale structure feature, d g (x,ε) is the grayscale structure eigenvalue between the neighboring pixel and the pixel to be tested, σ g Contribute parameters to structural features;

[0016] The spectral distance weight factor In the formula, x is the pixel to be tested, ε is the neighboring pixel within a certain relevant distance of x, and w o (x,ε) is the weight factor based on the spectral vector, d o (x,ε) is the spectral vector difference between the neighboring pixel and the pixel to be tested, σ o Contribute parameters to the spectrum;

[0017] Determine the guiding module of the three-side filter according to the three weight factors In the formula, x is the center pixel, ε i represents the neighborhood point, n is the number of neighborhood points of x, O represents the image before reconstruction, and R represents the image after filtering.

[0018] In the above scheme, the distance d s (x,ε) uses the Euclidean distance Where X is the pixel to be tested, Y is the neighboring pixel, the coordinate value of the pixel to be tested is fixed (a, b), and the coordinate value of the adjacent pixel is (i, j).

[0019] In the above scheme, the gray structure characteristic value d g (x, ε) is the gradient difference d between the current pixel to be tested and the related pixels around the pixel to be tested g(x,ε)=G(x)-G(ε), where G(x) is the gradient value of the pixel to be tested, and G(ε) is the gradient value of the adjacent pixels of the pixel to be tested.

[0020] In the above scheme, the gradient is calculated from the vertical, horizontal and diagonal directions Where (i, j) represents the current position coordinates, F(i, j) represents the current pixel gray value, G x , G y and G xy Represents the grayscale difference in the horizontal, vertical and diagonal directions respectively.

[0021] In the above scheme, the spectral vector difference value between the neighborhood pixel and the pixel to be tested is In the formula, N is the total number of bands, i is the number of the current band, and x is i and ε i are the grayscale values ​​of the spectral vector of the pixel to be tested and the spectral vector of the adjacent pixels in the i-th band after normalization.

[0022] In the above scheme, the guided feature map is element-wise multiplied by the learned channel weight through the attention mechanism module based on spatial-spectral union to obtain the feature map weighted by channel attention, specifically including: obtaining the feature map weighted by channel attention according to CA(F)=σ(MLP([AvgPool(F);MaxPool(F)]))·F, wherein F is the input feature map, [;] represents connecting the results of channel maximum pooling and channel average pooling in the feature dimension, MLP represents multilayer perceptron, σ represents Sigmoid activation function, CA(F) represents the feature map weighted by channel attention, which is obtained by multiplying the channel weight with the original image.

[0023] In the above scheme, the guided feature map is multiplied by the learned spatial weight through the attention mechanism module based on spatial-spectral joint to obtain the feature map after spatial attention weighting, specifically including: according to SA(F)=σ(f n×n (CA(F); AvgPool(CA(F)); MaxPool(CA(F))))·CA(F) obtains the feature map after spatial attention weighting, where F is the input feature map, [;] means connecting the results of spatial maximum pooling and spatial average pooling in the channel dimension, σ represents the Sigmoid activation function, n represents the dimension size of the joint feature fusion convolution kernel, and SA(F) represents the feature map after spatial attention weighting, which is obtained by multiplying the spatial weight with the feature map.

[0024] In the above scheme, the method further comprises: constructing a global error map of the detection result image based on the local feature adaptive loss function to obtain a reconstruction error G map;

[0025] The reconstruction error G map is converted into a weight map, and the training of the AE network is completed through the weight map.

[0026] In the above scheme, the global error map of the detection result image is constructed based on the local feature adaptive loss function to obtain the reconstruction error G map, which specifically includes:

[0027] The reconstruction error calculation for each pixel is expressed as where x i,j represents the pixel grayscale value of the input hyperspectral image, Represents the pixel grayscale value reconstructed by the network, and then the reconstruction error G can be obtained by combining the reconstruction errors of all pixels as G = [d 1,1 ,...,d 1,W ;...;d H,1 ,...,d H,W ], where d i,j represents the reconstruction error vector at the point (i, j), H and W represent the height and width of the hyperspectral image respectively;

[0028] The reconstruction error G map is converted into a weight map, and the AE network training is completed through the weight map, specifically including: dividing the image into 10 parts based on the global image dimension, re-defining the flag position in each part, if the difference between the maximum reconstruction error vector in the area and the global maximum reconstruction error does not exceed 10%, the maximum reconstruction error in the area is used as the standard for weight allocation in the area, if it exceeds 10%, the global maximum reconstruction error vector is still used as the standard for weight allocation, and the specific mathematical process is as follows:

[0029]

[0030] W=[ω 1,1 ,...,ω 1,W ;...;ω H,1 ,...,ω H,W ];

[0031] Where n = 1, 2, ..., 100 represents the divided area mark, represents the global maximum reconstruction error vector, represents the maximum reconstruction error vector in the current local area, represents the final selection criteria within the region, ω i,j represents the weight at the point (i, j), W represents the weight map based on the weight set of each point, and the weight map is updated every 100 iterations. In the first 100 iterations, all elements of the weight map are initialized to 1;

[0032] Based on the weight map, the global adaptive loss function

[0033] In the formula Represents the vector inner product. The first term is a loss term based on the mean square error, which is used to quantify the reconstruction error. The second term is a loss term based on the spectral angle, which is used to limit the similarity of the spectral dimension between the two images, thereby completing the network parameter update and the final background reconstruction.

[0034] Compared with the prior art, the present invention has the following beneficial effects:

[0035] The present invention introduces a trilateral filtering algorithm to purify the background of hyperspectral images, combines the spatial-spectral joint attention mechanism to improve the network feature expression ability, and uses an adaptive local weighted loss function to further suppress abnormal reconstruction, thereby significantly improving the accuracy and robustness of anomaly detection. Experimental results show that the method of the present invention has superior detection performance to traditional methods on different data sets and has high practical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The drawings described herein are used to disclose a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0037] Figure 1 is a schematic diagram of a joint feature filtering template in an embodiment of the present invention;

[0038] Figure 2 This is a schematic diagram of a joint representation filtering template based on a trilateral filtering algorithm;

[0039] Figure 3 This is a schematic diagram of the AutoEncoder network structure;

[0040] Figure 4 This is the anomaly detection result graph of each algorithm for the Cat Island dataset;

[0041] Figure 5 This is the logarithmic ROC curve of the Cat Island data detection results. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0043] The embodiment of the present invention provides a hyperspectral anomaly detection method based on joint features and trilateral filtering guidance, such as Figure 1 As shown, the method includes:

[0044] Step 101: constructing an AE network, wherein the AE network includes a guidance module based on trilateral filtering, an attention mechanism module based on spatial-spectral joint, and an adaptive loss function based on local features;

[0045] Specifically, the AE network consists of two modules: the encoding module and the decoding module. The encoder maps the input sample to the feature space, which is the encoding process. Then the decoder maps the abstract features back to the original space to obtain the reconstructed sample, which is the decoding process. The optimization goal is to optimize the encoder and decoder parameters simultaneously by minimizing the reconstruction error, so as to learn the abstract feature representation image for the sample input image. AE does not need to use the sample label during the optimization process. In essence, it uses the sample input as both the input and output of the neural network, and hopes to learn the abstract feature representation image of the sample by minimizing the reconstruction error. This unsupervised optimization method greatly improves the versatility of the model.

[0046] For the AE model based on neural network, the encoder part compresses the data by reducing the number of neurons layer by layer; the decoder part increases the number of neurons layer by layer based on the abstract representation of the data, and finally reconstructs the input samples. The specific structure is as follows: Figure 2 shown.

[0047] Figure 2 The encoder in contains 15 convolutional layers, each followed by batch normalization and leaky rectified linear unit (LeakyReLU) activation functions. Modules 1 and 4 each contain a 1×1 convolutional layer with a stride of 1. The feature maps generated by modules 1 and 4 are not input to the next convolutional layer, but are connected to the feature maps of the corresponding layers of the decoder through jump connections. Jump connections supplement the features of spatial details in the early layers of the network, thereby improving the spatial accuracy of the reconstructed background. At the same time, except for the convolutional layer in module 2 that reduces the dimension of the hyperspectral image and generates a d-dimensional feature map, other convolutional layers do not reduce the dimension of the feature map, which means that the dimension of the feature map remains unchanged at the feature dimension d during the encoding process, and the spectral features are preserved in this way. Because the purpose of the network is to reconstruct the background to suppress abnormal features, and the anomalies in hyperspectral images are usually small in size, the number of abnormal pixels is very limited. Therefore, the convolutional layers in modules 2 and 5 use a 3×3 convolutional layer with a stride of 2 to perform spatial downsampling, while the module 3 located after each module 2 and each module 5 contains a 3×3 convolutional layer with a stride of 1, thereby weakening abnormal features in the feature map during the encoding process.

[0048] The decoder contains 11 convolutional layers. Unlike the encoder, the decoder uses nearest neighbor interpolation with a scale of 2 to perform upsampling. The input of each module 6 is a 2D-dimensional feature map, that is, two d-dimensional feature maps are superimposed in the spectral dimension through skip connections. By including a 3×3 convolutional layer following batch normalization in module 6, the input 2D-dimensional feature map is reduced to a d-dimensional feature map after convolution. Module 4 contains a 1×1 convolutional layer with a stride of 1, which is located after module 6. Finally, module 7 contains a 1×1 convolutional layer with a stride of 1, through which the 128-dimensional feature map is increased to the same dimension as the original hyperspectral image. Unlike other modules, the convolutional layer in module 7 is followed by a sigmoid activation function. The input of the network is the original hyperspectral image, and the final output is a reconstructed background image with the same shape as the input hyperspectral image.

[0049] The guiding module based on trilateral filtering is constructed by a spatial correlation weight factor, a structural feature weight factor, and a spectral distance weight factor, specifically:

[0050] The spatial association weight factor In the formula, x is the pixel to be tested, ε is the neighboring pixel within a certain relevant distance of x, and w s (x,ε) is the weight factor based on the spatial position distance, d s (x,ε) is the distance between the neighboring pixel and the pixel to be tested, σ s Contribution parameters for distance;

[0051] The gray structure feature weight factor In the formula, x is the pixel to be tested, ε is the neighboring pixel within a certain relevant distance of x, and w g (x,ε) is the weight factor based on the grayscale structure feature, d g (x,ε) is the grayscale structure eigenvalue between the neighboring pixel and the pixel to be tested, σ g Contribute parameters to structural features;

[0052] The spectral distance weight factor In the formula, x is the pixel to be tested, ε is the neighboring pixel within a certain relevant distance of x, and w o (x,ε) is the weight factor based on the spectral vector, d o (x,ε) is the spectral vector difference between the neighboring pixel and the pixel to be tested, σ o Contribute parameters to the spectrum;

[0053] Since hyperspectral images have the property of being a unified image and spectrum, each pixel has different mathematical characteristics in its reflection spectrum under different optical properties. The reflection spectrum of each pixel can be extracted separately and regarded as an independent spectral vector for similarity comparison. Therefore, in addition to the weighting of the conventional spatial grayscale structure, adding a spectral distance weight factor helps to give a more accurate weighting to the correlation between the pixel to be tested and the surrounding pixels for precise filtering.

[0054] Determine the guiding module of the three-side filter according to the three weight factors In the formula, x is the center pixel, ε i represents the neighborhood point, n is the number of neighborhood points of x, O represents the image before reconstruction, and R represents the image after filtering.

[0055] o(x) is the regularization parameter, calculated as

[0056] The distance d s (x,ε) uses the Euclidean distance In the formula, X is the pixel to be tested, Y is the neighboring pixel, the coordinate value of the pixel to be tested is fixed (a, b), and the coordinate value of the adjacent pixel is (i, j). It can be known that the closer the pixel is to the pixel adjacent to the coordinate to be tested, the higher the weight value will be, and the farther the pixel is from the pixel to be tested, the lower the weight of the pixel to be tested.

[0057] d s (x, ε) is calculated using Euclidean distance. On the one hand, the Euclidean distance itself is measured by the physical distance between pixels in space, which has excellent intuitiveness and explainability. On the other hand, the Euclidean distance only needs to sum the squares of the coordinate differences and then take the square root, which is simpler and more efficient than distance calculation methods such as Mahalanobis distance and cosine distance.

[0058] The grayscale structure eigenvalue d g (x, ε) is the gradient difference d between the current pixel to be tested and the related pixels around the pixel to be tested g (x,ε)=G(x)-G(ε), where G(x) is the gradient value of the pixel to be tested, and G(ε) is the gradient value of the adjacent pixels of the pixel to be tested.

[0059] In the process of digital image processing, the gradient value can well describe the change of gray value in a certain space. For the task of anomaly detection, the local gradient change reflects the smoothness of the image in the current area. When there is an abnormal target, the gradient value around the pixel to be tested will show obvious changes, which is significantly different from the smooth background pixels. Therefore, the gradient value is selected as the structural feature with good robustness and information expression ability.g (x, ε) is used as the gradient change difference between the current pixel to be tested and the related pixels around the pixel to be tested.

[0060] Considering the accuracy of the hyperspectral anomaly detection task, abnormal pixels are often composed of only a small number of pixels. Therefore, in addition to considering the traditional vertical and horizontal directions, the diagonal direction is added to calculate the gradient change factors in the three directions. The gradient is calculated from the vertical, horizontal and diagonal directions.

[0061] Where (i, j) represents the current position coordinates, F(i, j) represents the current pixel gray value, G x , G y and G xy Represents the grayscale difference in the horizontal, vertical and diagonal directions respectively.

[0062] It can be seen from the above formula that if the pixel to be tested is in a smooth background position, the adjacent pixels will be assigned equal weights. If there are changes in the surroundings, the weight value of the change point will be reduced, thereby achieving the purpose of suppressing the weight of the change point caused by abnormal targets.

[0063] The spectral vector difference value between the neighborhood pixel and the pixel to be tested In the formula, N is the total number of bands, i is the number of the current band, and x is i and ε i are the grayscale values ​​of the spectral vector of the pixel to be tested and the spectral vector of the adjacent pixel after normalization in the i-th band. It can be seen that when the pixel to be tested and the adjacent pixel belong to the same background, the spectral vectors are similar, the smaller the corresponding spectral vector difference value is, the greater the weight obtained in the end is, and if there is anomaly between adjacent pixels, the spectral vector difference value increases accordingly, and the weight obtained in the background reconstruction process is smaller, so as to achieve the purpose of suppressing abnormal points.

[0064] From the principle of the three-side filtering algorithm, it can be found that the algorithm takes neighboring pixels as the benchmark, and mainly reconstructs the background for the sudden change of spatial pixel values ​​and spectral difference caused by target anomalies. If the current pixel is a background point, the surrounding abnormal points will be reconstructed due to the gradient change and spatial distance, resulting in a reduced weight; if the current point is an abnormal point and there are mostly background points around, the background points can also obtain a larger weight through spatial distance to suppress the reconstruction of abnormal points; but if the current point is an abnormal point and the abnormal target is clustered in a small range, the spatial association weight, structural feature weight and spectral distance weight will all be higher weight values, and the abnormality will not be suppressed. For this situation, a joint representation filter template is designed to suppress it.

[0065] The joint representation filter template is based on the inner and outer intervals. The inner interval refers to a rectangular range with a smaller spatial range, and the outer interval refers to a rectangular range with a larger spatial range. The specific schematic diagram is as follows Figure 3 As shown. The inner interval is mainly responsible for the pixel isolation function. In view of the phenomenon that abnormal target pixel points are clustered in a small range, the associated pixels within the inner interval are isolated through rectangular space so that they do not participate in the background purification task. The outer interval is mainly responsible for performing weight assignment based on the trilateral filtering algorithm and subsequent background purification tasks. It can be found that through the joint representation of the inner and outer intervals, when abnormal point aggregation occurs, if the target is an abnormal point, then since the abnormal point is clustered in the inner interval and does not participate in the subsequent weight assignment and background purification tasks, the abnormality will be suppressed and reconstructed as the background. If the target is a background point, the trilateral filtering weight of the background in the outer interval is higher, so the abnormal point is also reconstructed as the background.

[0066] Step 102: performing background purification on the image to be detected by the guiding module based on trilateral filtering to obtain a guiding feature map;

[0067] Step 103: multiplying the guide feature map by the learned channel weight element by element through the spatial-spectral joint-based attention mechanism module to obtain a feature map weighted by channel attention;

[0068] Specifically, the input feature map is first subjected to channel maximum pooling and channel average pooling operations to obtain the maximum feature response and average feature response in each channel. Then, the outputs of channel maximum pooling and channel average pooling are connected to form a feature description vector and input into the MLP part, and the channel weights are learned through the multi-layer perceptron structure. The learned channel weights are normalized by the Sigmoid activation function to represent the importance weight of each channel. Finally, the original feature map is multiplied by the learned channel weights to emphasize the channels that are helpful for anomaly detection and suppress irrelevant channels to obtain the weighted feature map.

[0069] According to CA(F)=σ(MLP([AvgPool(F);MaxPool(F)]))·F, the feature map after channel attention weighting is obtained, where F is the input feature map, [;] represents connecting the results of channel maximum pooling and channel average pooling in the feature dimension, MLP represents multilayer perceptron, σ represents Sigmoid activation function, CA(F) represents the feature map after channel attention weighting, which is obtained by multiplying the channel weight with the original image.

[0070] Step 104: multiplying the guided feature map by the learned spatial weight through the spatial-spectral joint attention mechanism module to obtain a feature map weighted by spatial attention;

[0071] Specifically, the input feature map is subjected to spatial maximum pooling and spatial average pooling operations respectively, thereby obtaining the maximum feature response and average feature response in each spatial position. Then, the outputs of spatial maximum pooling and spatial average pooling are connected in the spectral dimension and residually connected with the original feature map after channel attention weighting to form a feature description vector input to the convolution layer to obtain the spatial weight. The obtained spatial weight is normalized by the Sigmoid activation function to represent the importance weight of each spatial position. Finally, the original feature map is multiplied by the learned spatial weight to obtain the weighted feature map, which highlights the important spatial positions.

[0072] According to SA(F)=σ(f n×n (CA(F); AvgPool(CA(F)); MaxPool(CA(F))))·CA(F) obtains the feature map after spatial attention weighting, where F is the input feature map, [;] means connecting the results of spatial maximum pooling and spatial average pooling in the channel dimension, σ represents the Sigmoid activation function, n represents the dimension size of the joint feature fusion convolution kernel, and SA(F) represents the feature map after spatial attention weighting, which is obtained by multiplying the spatial weight with the feature map.

[0073] Step 105: multiplying the feature map weighted by channel attention and the feature map weighted by spatial attention to obtain a final feature representation;

[0074] Specifically, the channel attention module is used to adjust the channel weights of the feature map to highlight important feature channels, and the spatial attention module is used to adjust the spatial weights of the feature map to highlight important spatial locations. These two modules complement each other and work together on the input feature map, allowing the network to focus on important background features in a targeted manner. Finally, the outputs of the two modules are multiplied to obtain the final feature representation. In this way, the importance of features in both channels and space is taken into account at the same time, making the network more accurate and effective in extracting and utilizing features.

[0075] Step 106: Perform differential detection on the final feature representation to obtain a detection result image.

[0076] The method further comprises: constructing a global error map of the detection result image based on a local feature adaptive loss function to obtain a reconstruction error G map;

[0077] The reconstruction error G map is converted into a weight map, and the training of the AE network is completed through the weight map.

[0078] The global error map of the detection result image is constructed based on the local feature adaptive loss function to obtain the reconstruction error G map, which specifically includes:

[0079] The reconstruction error calculation for each pixel is expressed as where x i,j represents the pixel grayscale value of the input hyperspectral image, Represents the pixel grayscale value reconstructed by the network, and then the reconstruction error G can be obtained by combining the reconstruction errors of all pixels as G = [d 1,1 ,...,d 1,W ;...;d H,1 ,...,d H,W ], where d i,j represents the reconstruction error vector at the point (i, j), H and W represent the height and width of the hyperspectral image respectively;

[0080] By constructing the reconstruction error map, two characteristics can be found: first, since the number of pixels occupied by abnormal targets is relatively small, it is often difficult for the network to reconstruct abnormal target pixels during the initial training process, while background pixel reconstruction is relatively easy; second, since the grayscale value of the abnormal point itself is different from the surrounding background pixels, and the abnormal targets are often distributed in a discrete manner, they have local highlighting characteristics.

[0081] By reconstructing the abnormal image and combining the characteristics, a weight calculation method can be proposed. First, the global maximum reconstruction error vector max(E) can be used as a flag for weight allocation and the difference calculation with the other points, which represents how much grayscale difference should be given to the abnormal point by the background reconstruction performed by the current network. Using this for global evaluation, the improvement direction of other points in the reconstructed image in the background reconstruction process can be determined. If the target point is an abnormal point, there will also be a large reconstruction error, and the final weight value will be reduced accordingly; if the target point is a background point, it will obtain a larger weight in the subsequent network training process due to the lower reconstruction error.

[0082] The reconstruction error G map is converted into a weight map, and the AE network training is completed through the weight map, specifically including: dividing the global image dimension into 10 parts, re-defining the flag position in each part, if the difference between the maximum reconstruction error vector in the area and the global maximum reconstruction error does not exceed 10%, the maximum reconstruction error in the area is used as the standard for weight allocation in the area, if it exceeds 10%, the global maximum reconstruction error vector is still used as the standard for weight allocation, and the specific mathematical process is as follows:

[0083]

[0084]

[0085] W=[ω 1,1 ,...,ω 1,W ;...;ω H,1 ,...,ωH,W ];

[0086] Where n = 1, 2, ..., 100 represents the divided area mark, represents the global maximum reconstruction error vector, represents the maximum reconstruction error vector in the current local area, represents the final selection criteria within the region, ω i,j represents the weight at the point (i, j), W represents the weight map based on the weight set of each point, and the weight map is updated every 100 iterations. In the first 100 iterations, all elements of the weight map are initialized to 1;

[0087] Based on the weight map, the global adaptive loss function

[0088] In the formula Represents the vector inner product. The first term is a loss term based on the mean square error, which is used to quantify the reconstruction error. The second term is a loss term based on the spectral angle, which is used to limit the similarity of the spectral dimension between the two images, thereby completing the network parameter update and the final background reconstruction.

[0089] The present invention makes improvements on the original fully convolutional AE network. On the one hand, a guiding layer is added, and the original hyperspectral image is processed by the trilateral filtering algorithm and the result is connected to the decoder to guide the reconstruction and improve the robustness of the network. On the other hand, a spatial-spectral joint attention module composed of spatial attention and channel attention is inserted in the process of convolutional feature extraction. The receptive field between layers is enhanced by jump connections based on the spatial-spectral joint attention module, and the spatial-spectral joint information advantage brought by the hyperspectral image is fully utilized. Finally, an adaptive loss function based on local features is added to enhance the network's ability to suppress abnormal reconstruction.

[0090] Three groups of real hyperspectral remote sensing images collected in different scenes are selected as experimental data sets to evaluate the detection performance of the proposed anomaly detection algorithm and the comparison algorithm. The data sets are introduced in detail below.

[0091] (1) Los Angeles Dataset

[0092] The Los Angeles dataset was collected by the AVIRIS sensor in the airspace near Los Angeles Airport in the United States. Its spectral resolution is 10nm, the spectral range is 370-2510nm, and the image space size is 100×100 pixels. After removing the bands with large noise, 205 main bands remain. The main scenes cover the airport and part of the urban landscape, including buildings, highway runways, and aprons. Two aircraft parked on the apron are distinguished as abnormal targets in the hyperspectral image.

[0093] (2) Cat Island Dataset

[0094] The Cat Island dataset was collected by the AVIRIS sensor in the airspace near Cat Island, Japan. Its spectral resolution is 10nm, the spectral range is 370-2510nm, the spatial resolution is 17.2m, and the image space size is 150×150 pixels. After removing the bands with large noise, 188 main bands remain. The main scenes cover the sea surface and part of the Cat Island landscape, including targets such as island reefs and surrounding sea areas. The main ships operating in the sea area are distinguished as abnormal targets in the hyperspectral image.

[0095] (3) Abu-urban-4 dataset

[0096] The Abu-urban-4 dataset was collected by the AVIRIS sensor in the airspace of Los Angeles, USA. Its spectral resolution is 10nm, the spectral range is 370-2510nm, the spatial resolution is 17.2m, and the image space size is 100×100 pixels. After removing the bands with large noise, 205 main bands remain. The main scenes cover the urban area and some mountainous areas, including roads, buildings and other targets. The houses in the scenes are distinguished as abnormal targets in the hyperspectral image.

[0097] There are two experimental environments. Experimental environment one is mainly used to train and test the constructed hyperspectral anomaly detection algorithm based on joint features and trilateral filtering guided autoencoder and some contrast algorithms based on neural networks. Experimental environment two is mainly used to test some traditional algorithms, and relies on Matlab's powerful linear computing and drawing capabilities to integrate the final detection results and visualize related detection indicators.

[0098] The experimental environment configuration information is as follows:

[0099] (1) Experimental environment 1: System Windows 11, code tool library and version Pytorch 1.10.2 and Python 3.9, GPU: GeForce RTX4060 video memory: 16GB

[0100] (2) Experimental environment 2: System: Windows 11, Code tool library and version: MatlabR 2021a, CPU: Intel (R Core (TM) i7-13700H Main frequency: 2.4GHz

[0101] The present invention proposes that the training parameters of the joint feature and trilateral filtering guided autoencoder network (hereinafter referred to as JFTBF-AE algorithm) include training round epoch, calibration loss th, learning rate lr, hidden layer feature dimension d, and the network training end standard is based on the training round and calibration loss. When the round reaches the epoch specified value or the network loss is less than the calibration loss th value, the network training ends. Set the initial learning rate lr to 0.01, the hidden layer feature dimension d to 128, and plan to start training based on 800.

[0102] This comparative experiment selects RX, KRX, PTA, AHMID, MsRFQFT, and PCA-TLRSR as the comparative algorithms to compare the superiority of the JFTBF-AE algorithm in this chapter. Among them, the RX algorithm does not need to set additional parameters because it is an adaptive type algorithm. The KRX algorithm usually has robust performance on various data sets, so the parameters are set according to the suggestions given in the original text, and the clustering category l is proposed. K is 6, the kernel function parameter S K is 300.

[0103] The following is a comparative experimental analysis.

[0104] The experimental data includes targets such as Carter Island reefs, surrounding waters, and operating vessels, among which the operating vessels are identified as abnormal targets and the rest of the scenes are identified as background parts.

[0105] In order to verify the advantages of this method in hyperspectral anomaly detection, six additional algorithms were selected for comparative experiments, namely: RX, KRX, PTA, AHMID, MsRFQFT, PCA-TLRSR. The anomaly detection results of each algorithm on the Cat Island dataset are shown in Figure 2. Figure 4 As shown in the figure, (a) is the original color image of the dataset, (b) is the abnormal target distribution map, and the rest are pseudo-color images of the abnormality detection results of each algorithm.

[0106] From the figure, we can see that although RX can suppress most of the background and detect the position of the target, the target response is low, so the performance is average. The AHMID algorithm reconstructs the background more fully, but the suppression of abnormal targets is weak, and the abnormal targets in the detection results are not obvious. The KRX algorithm, PTA algorithm, MsRFQFT algorithm and PCA-TLRSR algorithm do not reconstruct the background sufficiently, resulting in a large area of ​​obvious background in the final differential result. The algorithm in this chapter fully reconstructs the background in the process, so the false alarm rate is reduced, and the abnormal targets are effectively highlighted through the attention mechanism. The abnormal targets in the detection results have a certain brightness while maintaining a good degree of integrity. In comparison, better detection results are achieved.

[0107] In order to objectively evaluate the algorithm detection results, the logarithmic ROC curves of the algorithm detection results were drawn for analysis and comparison. Figure 5 shown.

[0108] It can be seen from the figure that when the false alarm probability is less than 10 -4 In this case, the detection probability of each algorithm is low, and as the false alarm probability gradually increases, the algorithm in this chapter and the RX algorithm are closer to each other in the process. Both have good anomaly detection effects for the CatIsland algorithm, while the anomaly detection effects of other algorithms are relatively poor.

[0109] In order to verify the advantages of the algorithm in this chapter in hyperspectral anomaly detection, six additional algorithms were selected for comparative experiments, namely: RX, KRX, PTA, AHMID, MsRFQFT, PCA-TLRSR. Visual subjective analysis showed that the RX algorithm, PTA algorithm, MsRFQFT algorithm, PCA-TLRSR algorithm and the algorithm in this chapter can suppress the background well and detect abnormal targets. Among them, the reconstruction effect of the background part in the process of the PCA-TLRSR algorithm and the algorithm in this chapter is better than that of other comparative algorithms. Although the KRX algorithm and the AHMID algorithm can detect more abnormal targets, the background is not well reconstructed in the process, resulting in a large number of false alarms, showing a large area of ​​non-abnormal bright points.

[0110] Aiming at the problem that the traditional fully convolutional autoencoder network does not fully utilize the spatial and spectral information of hyperspectral images, resulting in incomplete suppression of abnormal targets, this method studies and designs a hyperspectral anomaly detection algorithm based on joint features and trilateral filtering guided autoencoders. The main body is: first, in order to improve the network's ability to express spatial and spectral features, a spatial and spectral joint attention module is designed to weight the feature information in the process from the spatial and spectral dimensions. Secondly, in order to improve the utilization of the spatial and spectral information of hyperspectral images to suppress abnormal reconstruction, the reconstruction weights are assigned in the dual interval by combining the features of the three dimensions of spatial distance, grayscale structure and spectral vector, so as to achieve background purification to guide the network to perform background reconstruction more accurately. Finally, the reconstruction error in the local area of ​​the comprehensive differential image is redistributed to adaptively further suppress abnormal reconstruction. The algorithm jointly utilizes the spatial features, spectral features and abnormal distribution characteristics of hyperspectral images, effectively increases the network's ability to express background features, suppresses the reconstruction probability of abnormal targets, and improves the network's ability to detect abnormal targets. By comparing with six algorithms on three hyperspectral data, the experimental results prove the effectiveness of the proposed algorithm. Overall, this method has good detection performance.

[0111] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention.

Claims

1. A hyperspectral anomaly detection method based on joint features and trilateral filtering guidance, characterized in that: The method includes: Constructing an AE network, wherein the AE network includes a guidance module based on trilateral filtering, an attention mechanism module based on spatial-spectral joint, and an adaptive loss function based on local features; The guidance module based on trilateral filtering is used to purify the background of the image to be detected to obtain a guidance feature map; The guided feature map is element-wise multiplied by the learned channel weights through an attention mechanism module based on spatial-spectral joint to obtain a feature map weighted by channel attention; Multiplying the guided feature map with the learned spatial weights through an attention mechanism module based on spatial-spectral joint to obtain a feature map weighted by spatial attention; Multiplying the feature map weighted by channel attention and the feature map weighted by spatial attention to obtain a final feature representation; Perform differential detection on the final feature representation to obtain a detection result image.

2. The hyperspectral anomaly detection method based on joint features and trilateral filtering guidance according to claim 1 is characterized in that: The guiding module based on trilateral filtering is constructed by a spatial correlation weight factor, a structural feature weight factor, and a spectral distance weight factor, specifically: The spatial association weight factor In the formula, x is the pixel to be tested, ε is the neighboring pixel within a certain relevant distance of x, and w s (x,ε) is the weight factor based on the spatial position distance, d s (x,ε) is the distance between the neighboring pixel and the pixel to be tested, σ s Contribution parameter for distance; The gray structure feature weight factor In the formula, x is the pixel to be tested, ε is the neighboring pixel within a certain relevant distance of x, and w g (x,ε) is the weight factor based on the grayscale structure feature, d g (x,ε) is the grayscale structure eigenvalue between the neighboring pixel and the pixel to be tested, σ g Contribute parameters to structural features; The spectral distance weight factor In the formula, x is the pixel to be tested, ε is the neighboring pixel within a certain relevant distance of x, and w o (x,ε) is the weight factor based on the spectral vector, d o (x,ε) is the spectral vector difference between the neighboring pixel and the pixel to be tested, σ o Contribute parameters to the spectrum; Determine the guiding module of the three-side filter according to the three weight factors In the formula, x is the center pixel, ε i represents the neighborhood point, n is the number of neighborhood points of x, O represents the image before reconstruction, and R represents the image after filtering.

3. The hyperspectral anomaly detection method based on joint features and trilateral filtering guidance according to claim 2 is characterized in that: The distance d s (x,ε) uses the Euclidean distance Where X is the pixel to be tested, Y is the neighboring pixel, the coordinate value of the pixel to be tested is fixed (a, b), and the coordinate value of the adjacent pixel is (i, j).

4. The hyperspectral anomaly detection method based on joint features and trilateral filtering guidance according to claim 3 is characterized in that: The grayscale structure eigenvalue d g (x, ε) is the gradient change difference d between the current pixel to be tested and the related pixels around the pixel to be tested g (x,ε)=G(x)-G(ε), where G(x) is the gradient value of the pixel to be tested, and G(ε) is the gradient value of the adjacent pixels of the pixel to be tested.

5. The hyperspectral anomaly detection method based on joint features and trilateral filtering guidance according to claim 4 is characterized in that: Calculate gradients in vertical, horizontal and diagonal directions Where (i, j) represents the current position coordinates, F(i, j) represents the current pixel gray value, G x , G y and G xy Represents the grayscale difference in the horizontal, vertical and diagonal directions respectively.

6. The hyperspectral anomaly detection method based on joint features and trilateral filtering guidance according to claim 5 is characterized in that: The spectral vector difference value between the neighborhood pixel and the pixel to be tested In the formula, N is the total number of bands, i is the number of the current band, and x is i and ε i are the grayscale values ​​of the spectral vector of the pixel to be tested and the spectral vector of the adjacent pixels in the i-th band after normalization.

7. The hyperspectral anomaly detection method based on joint features and trilateral filtering guidance according to claim 6 is characterized in that: The described attention mechanism module based on spatial-spectral union multiplies the guided feature map by the learned channel weight element by element to obtain the feature map weighted by channel attention, specifically including: obtaining the feature map weighted by channel attention according to CA(F)=σ(MLP([AvgPool(F);MaxPool(F)]))·F, wherein F is the input feature map, [;] represents connecting the results of channel maximum pooling and channel average pooling in the feature dimension, MLP represents multilayer perceptron, σ represents Sigmoid activation function, CA(F) represents the feature map weighted by channel attention, which is obtained by multiplying the channel weight with the original image.

8. The hyperspectral anomaly detection method based on joint features and trilateral filtering guidance according to claim 7 is characterized in that: The guide feature map is multiplied by the learned spatial weight through the attention mechanism module based on spatial-spectral joint to obtain the feature map after spatial attention weighting, specifically including: according to SA(F)=σ(f n×n (CA(F); AvgPool(CA(F)); MaxPool(CA(F))))·CA(F) obtains the feature map after spatial attention weighting, where F is the input feature map, [;] means connecting the results of spatial maximum pooling and spatial average pooling in the channel dimension, σ represents the Sigmoid activation function, n represents the dimension size of the joint feature fusion convolution kernel, and SA(F) represents the feature map after spatial attention weighting, which is obtained by multiplying the spatial weight with the feature map.

9. The hyperspectral anomaly detection method based on joint features and trilateral filtering guidance according to claim 1 is characterized in that: The method further includes: constructing a global error map of the detection result image based on a local feature adaptive loss function to obtain a reconstruction error G map; The reconstruction error G map is converted into a weight map, and the training of the AE network is completed through the weight map.

10. The hyperspectral anomaly detection method based on joint features and trilateral filtering guidance according to claim 9 is characterized in that: The global error map of the detection result image is constructed based on the local feature adaptive loss function to obtain the reconstruction error G map, which specifically includes: The reconstruction error calculation for each pixel is expressed as where x i,j represents the pixel grayscale value of the input hyperspectral image, Represents the pixel grayscale value reconstructed by the network, and then the reconstruction error G can be obtained by combining the reconstruction errors of all pixels as G = [d 1,1 ,...,d 1,W ;...;d H,1 ,...,d H,W ], where d i,j represents the reconstruction error vector at the point (i, j), H and W represent the height and width of the hyperspectral image respectively; The reconstruction error G map is converted into a weight map, and the AE network training is completed through the weight map, specifically including: dividing the global image dimension into 10 parts, re-defining the flag position in each part, if the difference between the maximum reconstruction error vector in the area and the global maximum reconstruction error does not exceed 10%, the maximum reconstruction error in the area is used as the standard for weight allocation in the area, if it exceeds 10%, the global maximum reconstruction error vector is still used as the standard for weight allocation, and the specific mathematical process is as follows: W=[ω 1,1 ,...,oh 1,W ;...;oh H,1 ,...,oh H,W ]; Where n = 1, 2, ..., 100 represents the divided area mark, represents the global maximum reconstruction error vector, represents the maximum reconstruction error vector in the current local area, represents the final selection criteria within the region, ω i,j represents the weight at the point (i, j), W represents the weight map based on the weight set of each point, and the weight map is updated every 100 iterations. In the first 100 iterations, all elements of the weight map are initialized to 1; Based on the weight map, the global adaptive loss function In the formula Represents the vector inner product. The first term is a loss term based on the mean square error, which is used to quantify the reconstruction error. The second term is a loss term based on the spectral angle, which is used to limit the similarity of the spectral dimension between the two images, thereby completing the network parameter update and the final background reconstruction.