Implicit self-supervision hyperspectral characterization anomaly detection method based on background credibility enhancement
By employing an implicit self-supervised hyperspectral characterization method, combined with color space mapping and a spectral feature cube builder, the problems of robustness and insufficient semantic information in hyperspectral anomaly detection are solved, achieving high-precision anomaly detection in complex backgrounds.
Patent Information
- Application Number
- CN202511057133.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-14
AI Technical Summary
Existing hyperspectral anomaly detection technologies have poor robustness and limited discrimination ability in complex background environments, and lack high-level semantic information, resulting in insufficient detection accuracy.
An implicit self-supervised hyperspectral characterization method based on background credibility enhancement is adopted. By combining implicit neural representation network and self-supervised learning with color space mapping and spectral feature cube builder, high-precision reconstruction of background spectrum and differentiation of abnormal targets are achieved. A background credibility enhancement mechanism is introduced to suppress abnormal interference.
It significantly improves the accuracy and robustness of hyperspectral anomaly detection, effectively models spectral-spatial dependencies in complex backgrounds, and enhances the ability to detect anomalous targets.
Smart Images

Figure CN120953804A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent image analysis technology, specifically to an anomaly detection method and system based on implicit self-supervised hyperspectral characterization with enhanced background credibility. Background Technology
[0002] Currently, HAD technology can be mainly divided into two categories: traditional learning-based and deep learning-based. Traditional hyperspectral image detection methods refer to the construction of detection algorithms based on traditional machine learning, statistical learning, and pattern recognition technologies. These include statistical methods that accurately model the background distribution, such as the Reed-Xiaoli (RX) algorithm (Reed IS, Yu X. Adaptive multiple-band CFAR detection of an optical pattern with unknown spectral distribution[J]. IEEE Transactions on Acoustics Speech & Signal Processing, 1990, 38(10): 1760-1770.), and methods based on representation learning and dictionary learning, such as low-rank representation and sparse representation. The RX algorithm assumes that background pixels follow a multivariate Gaussian distribution, while anomalous targets deviate from the background distribution. It detects anomalous targets by calculating the Mahalanobis distance between the two distributions. However, in real hyperspectral images, the zero-mean multivariate normal distribution assumption is often difficult to accurately model background features in the scene, and the global background estimation for the entire image is not accurate. Therefore, researchers have proposed many improvement strategies for the original RX method. For example, by using a dual-window sliding strategy, the hyperspectral image is set to follow a Gaussian model in a local region to improve the inaccuracy of parameter estimation in the RX algorithm and improve the detection effect. In addition, the RX improvement methods also include: kernel RX (Kwon H, Nasrabadi N. Kernel RX-algorithm: anonlinear anomaly detector for hyperspectral imagery[J]. IEEE Transactions on Geoscience and Remote Sensing, 2005, 43(2):388–397.) and weighted RX (Guo Q, Zhang B, Ran Q, et al. Weighted-rxd and linear filter-based rxd: Improving backgroundstatistics estimation for anomaly detection in hyperspectral imagery[J]. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2014, 7(6):2351–2366.).Regularized KRX (Theiler J, G. Grosklos. Problemmatic projection to the in-sample subspace for a kernelized anomaly detector[J].IEEE Geoscience and Remote Sensing Letters, 2016, 13(4): 485–489.), Cluster kernel RX algorithm (Zhou J, Kwan C, Ayhan B, et al. A novel cluster kernel rx algorithm for anomaly and change detection using hyperspectral images[J].IEEE Transactions on Geoscience and Remote Sensing, 2016, 54(11): 6497–6504.), Fractional Fourier entropy-based algorithm (Tao R, Zhao X, Li W, et al. Hyperspectral Anomaly Detection by Fractional Fourier Entropy[J].IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2019, PP(99): 1-10.), and superpixel-based RX (Ren L, Zhao L, Wang YA superpixel-based dual window rx for hyperspectral anomaly detection[J].IEEE Geoscience and Remote Sensing Letters,2020,17(7):1233–1237.). Hyperspectral image anomaly detection algorithms based on the RX operator have strong ease of use and execution efficiency, but they still cannot accurately express the complex distribution characteristics of hyperspectral images, and it is difficult to effectively filter out the interference of anomaly and noise data, which will seriously affect the results of anomaly detection. In order to avoid making distribution assumptions about the data, many scholars have proposed representation learning-based methods, such as sparse representation, low-rank representation, etc., combined with dictionary learning methods, to model the background or anomalies through the prior constraints of the data itself. Among them, Wang et al. (Wang X, Wang L, Wu H, et al.)A double-dictionary-based nonlinear representation model for hyperspectral subpixel target detection[J].IEEE Transactions on Geoscience and RemoteSensing,2022,60:1–16.) A nonlinear model based on background and target dictionaries is used to represent hyperspectral images. An overcomplete background dictionary construction strategy is designed to combine spectral angular distance with sparse representation to more effectively represent the background part, and can reliably separate the background and the target. Guo et al. (Guo T, He L, Luo F, et al. Anomaly detection of hyperspectral image with hierarchical antinoise mutual-incoherence induced low-rank representation[J].IEEE Transactions on Geoscience and RemoteSensing,2023,61:1–13.) proposed a noise-resistant hierarchical mutual-incoherence induced discriminative learning method. By designing structural inconsistency constraints and a hierarchical alternation strategy, anomaly detection under complex mixed noise interference is achieved. To effectively utilize the spatial dimension information of hyperspectral images, He et al. (He X, Wu J, Ling Q, et al. Anomaly detection for hyperspectral imagery via tensor low-rank approximation with multiple subspace learning[J]. IEEE Transactions on Geoscience and Remote Sensing, 2023, 61: 1–17.) developed a novel tensor low-rank approximation detection algorithm. In addition, anomaly detection methods based on collaborative representation have also been used for anomaly detection in hyperspectral images. Lin et al. (Lin S, Zhang M, Cheng X, et al. Hyperspectral anomaly detection via sparse representation and collaborative representation[J].)The IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2023, 16: 946–961, proposes a novel HAD method that combines sparse and cooperative representations.
[0003] With the development of deep learning technology, deep learning methods for remote sensing image processing have been proposed one after another. Due to the powerful feature representation capabilities of deep learning, deep learning methods have achieved significant advantages in the field of remote sensing image processing. In the field of hyperspectral anomaly detection, Zhang et al. (Zhang L, Cheng B. Transferred CNN based on tensor for hyperspectral anomaly detection[J].IEEE Geoscience and Remote Sensing Letters,2020,17(12):2115–2119.) proposed a tensor-based transport convolutional neural network for hyperspectral anomaly detection. Song et al. (Song X, Zhang T, Xu Z, et al. Anomaly detection of hyperspectral images based on transformer with spatialcspectral dual-window mask[J].IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing,2023,16:1414–1426.) used a transform neural network to fully extract features from both global and local perspectives to achieve anomaly detection. Lian et al. (Lian J, Wang L, Sun H, et al. Gt-had: Gated transformer for hyperspectral anomaly detection[J]. IEEE Transactions on Neural Networks and Learning Systems, 2025, 36(2):3631–3645.) proposed a gated Transformer network, which effectively improves the ability to distinguish between background and anomaly in hyperspectral anomaly detection by introducing spatial-spectral similarity modeling and content matching mechanisms. Currently, researchers have proposed using auxiliary tasks to solve the hyperspectral image anomaly detection task under the deep learning framework. Among them, the most common auxiliary task is deep reconstruction learning.Jiang et al. (Jiang T, Li Y, Xie W, et al. Discriminative reconstruction constrained generative adversarial network for hyperspectral anomaly detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2020, 58(7):4666–4679.) proposed an anomaly detection algorithm based on discriminative reconstruction constrained generative adversarial network, which for the first time used an unsupervised GAN model to solve the problem of high-dimensional data background modeling. Fan et al. (Fan G, Ma Y, Mei X, et al. Hyperspectral anomaly detection with robust graph autoencoders[J]. IEEE Transactions on Geoscience and Remote Sensing, 2022, 60:1–14.) proposed a robust graph autoencoder detector, which embeds a graph regularization term based on superpixel segmentation into a norm-based autoencoder. Wang et al. (Wang D, Zhuang L, Gao L, et al. Bocknet: Blind-block reconstruction network with a guard window for hyperspectral anomaly detection[J]. IEEE Transactions on Geoscience and Remote Sensing, 2023, 61: 1–16.) proposed a blind-block network that reduces the contribution of anomalies to the reconstruction results by creating a blind block at the center of the network's receptive field. Li et al. (Li K, Ling Q, Wang Y, et al. Spectral difference guided graph attention autoencoder for hyperspectral anomaly detection[J]. IEEE Transactions on Instrumentation and Measurement, 2023, 72: 1–17.) introduced spectral sharpening constraints into a graph attention autoencoder network to guide attention coefficient learning, thereby achieving better background modeling.Cheng et al. (Cheng X, Huo Y, Lin S, et al. Deep feature aggregation network for hyperspectral anomaly detection[J].IEEE Transactions on Instrumentation and Measurement, 2024, 73: 1–16.) proposed a novel deep feature aggregation network that effectively improves the accuracy of background reconstruction and the ability to distinguish anomalous targets in hyperspectral anomaly detection by introducing multi-background pattern modeling and a joint attention mechanism. Besides the above algorithm, there are many other excellent hyperspectral anomaly detection algorithms based on deep reconstruction learning. These algorithms improve their detection performance by designing more reasonable constraints and network structures.
[0004] Existing HAD (High-Intensity Detection) technologies suffer from poor robustness, limited discriminative ability, and high false alarm rates when handling complex background environments. Many methods rely on manually defined background priors and lack a unified mechanism to simultaneously model global spectral-spatial correlations, resulting in insufficient detection capabilities in heterogeneous backgrounds, cluttered noise, and blurred target regions. The reasons for this are as follows:
[0005] (1) Traditional methods rely on Gaussian distribution assumptions or shallow feature modeling, which cannot effectively express nonlinear distributions and complex structural backgrounds;
[0006] (2) Some deep learning methods only use pixel-level reconstruction or local attention mechanisms, which have limited receptive fields and fail to fully capture long-distance background dependencies.
[0007] (3) Many models lack guidance from high-level semantic information, especially in training, they fail to effectively shield abnormal interference, resulting in insufficient learning of normal background and reduced separation ability.
[0008] (4) Most existing methods use pixel-by-pixel processing, which ignores the consistency of spatial structure and background clustering characteristics, resulting in blurred boundaries of abnormal areas or local misjudgment.
[0009] To address the aforementioned issues, this invention proposes a method and system for anomaly detection based on implicit self-supervised hyperspectral characterization with enhanced background credibility. Summary of the Invention
[0010] 1. The technical problem to be solved by the present invention
[0011] The purpose of this invention is to propose an implicit self-supervised hyperspectral representation anomaly detection method and system based on background credibility enhancement to address the problems of insufficient background distribution modeling, lack of high-level semantic information extraction capability, and the impact of anomaly interference on learning performance in existing hyperspectral anomaly detection technologies in complex scenes. This invention proposes a background-aware implicit representation learning method for hyperspectral anomaly detection, effectively modeling the spectral-spatial dependencies between background pixels and improving the reconstruction capability of normal targets and the differentiation capability of anomaly targets through mapping from the visible light domain to the hyperspectral domain. Simultaneously, a data-driven background credibility enhancement mechanism is introduced to mine the clustering structure of the background spectrum, effectively suppressing the interference of anomaly pixels on the model training process.
[0012] 2. Technical Solution
[0013] To achieve the above objectives, the present invention provides the following technical solution:
[0014] A hyperspectral anomaly detection method based on background confidence enhancement implicit self-supervised hyperspectral characterization includes the following steps:
[0015] S1. Design a color space mapping generator to construct color space mapping data for network training;
[0016] S2. Design a spectral feature cube builder to provide real hyperspectral images (HSI) data required for guiding network training;
[0017] S3. Design an implicit self-supervised hyperspectral characterization anomaly detection module, specifically including:
[0018] S3.1 Construct an implicit self-supervised hyperspectral representation network, introducing implicit neural representation (INR) into the field of hyperspectral anomaly detection. After inputting training image samples, the background spectrum is reconstructed with high precision by combining a color space mapping generator and a spectral feature cube builder. The background spectral information is recovered from the training samples, and the corresponding real hyperspectral image (HSI) cube is reconstructed. This realizes the mapping of samples from color space to spectral space, so as to reveal the deep connection between the appearance and spectrum of the samples and capture the high-level semantic information of normal targets.
[0019] S3.2 Design a background credibility enhancement learning mechanism to obtain pure and high-quality normal samples when processing hyperspectral anomaly detection datasets, ensuring that the network is not affected by abnormal targets during the training phase;
[0020] S3.3 Design training loss and anomaly detection. During the training phase, the accuracy of background region reconstruction is used as the metric for the loss function, and the root mean square error is used as the indicator to evaluate the quality difference between the reconstructed image and the real image. During the testing phase, the image is divided into non-intersecting spatial cubes, and anomaly detection is performed using self-supervised differential recovery features and the spatial distribution characteristics of abnormal targets.
[0021] Preferably, S1 specifically includes the following:
[0022] For the RGB color image corresponding to hyperspectral image Color space mapping data was obtained using a data augmentation method employing a k×k sliding window. Enhance the network's ability to extract spectral and spatial information; among which, hsi m and hsi n The horizontal and vertical pixel counts of the image are respectively represented; k represents the size of the color space mapping data block; i∈1,…,n represents the index number of the color space mapping data, and n is the number of data blocks. The size of the block must satisfy k<hsi m And k < hsi n ;
[0023] The sample set is expanded using a sliding window method to provide image samples for training an implicit self-supervised hyperspectral representation network; the specific process of the sliding window is shown in equation (1):
[0024]
[0025] Among them, w i and h i represents the initial position index of the i-th color space mapping data; p represents the step size of the sliding window.
[0026] Preferably, S2 specifically includes the following:
[0027] Dimensionality reduction techniques are employed to process real hyperspectral images (HSI) to reduce the number of spectral bands in the images; this is done when processing the original real hyperspectral image (HSI) data. During the process, a method based on an optimized clustering framework is used to select the top n spectral bands that best represent the data characteristics. s One spectral channel;
[0028] Generating data X using the sliding window technique, compared to the original data X. i Corresponding real hyperspectral image (HSI) cube
[0029] The resulting real hyperspectral image (HSI) cube S iIt is used to guide the network to complete the self-supervised learning process, avoid the interference of redundant information on the model's learning effect, and enable the model to master the mapping relationship between the appearance space and spectral space of normal samples.
[0030] Preferably, S3.1 specifically includes the following:
[0031] Deep features of the input samples are extracted using a feature encoder.
[0032] Using the obtained deep features M in and the number of target spectral bands n s As input, a spectral image that meets the target band quantity requirement is reconstructed, and the process is shown in equation (2):
[0033]
[0034] in, Represents the reconstructed true hyperspectral image (HSI) cube;
[0035] Deep features are enhanced by using Spectral Feature Refinement Technique (SFRT) and Spatial-Spectral Synergistic Attention Mechanism (SSSAM). The number of target spectral bands is regarded as the coordinates of a continuous implicit function, thereby mapping deep features onto spectral intensity and achieving accurate projection from deep features to spectral intensity.
[0036] Preferably, the Spectral Feature Refinement Technique (SFRT) is used to analyze the spatial-spectral correlation, which utilizes spectral interpolation to characterize the spatial-spectral interaction of features, specifically including:
[0037] For deep features The analysis is conducted along both horizontal and vertical directions to uncover the intrinsic relationship between spatial and spectral dimensions, which are represented as follows: and
[0038] Using bilinear interpolation, the horizontal spectral features of the upsampled data are constructed respectively. and vertical spectral characteristics The process, represented by rich features, is shown in equation (3):
[0039]
[0040] Where BI(·) represents bilinear interpolation;
[0041] Convert the upsampled spectral features into a feature map. By fusing features through splicing operations, a three-dimensional deep feature suitable for implicit neural representation (INR) is constructed, as shown in Equation (4):
[0042]
[0043] Among them, M out This represents the output characteristics after processing with Spectral Feature Refinement Technology (SFRT); Conv 3D (·) and [·,·] represent a 3D convolutional layer with the LeakyReLU activation function and a splicing operation, respectively.
[0044] Preferably, the spatial-spectral collaborative attention mechanism (SSSAM) focuses on achieving global attention in both spatial and spectral dimensions. It utilizes a spatial-spectral attention mechanism to analyze the similarities between various channels, thereby enhancing the comprehensiveness of spatial-spectral information. Specifically, it includes:
[0045] Using a three-dimensional tensor as input, the three-dimensional features are converted into two-dimensional unit embeddings and mapped through a multilayer perceptron, as shown in equation (5):
[0046] v = MLP(Reshape(M out (5)
[0047] Where MLP(·) and Reshape(·) represent the multilayer perceptron and reshaping operation, respectively; This indicates the reshaped unit embedding;
[0048] During the attention computation process, unit embeddings are transformed into key embeddings and memory embeddings to depict the correspondence between deep features. Attention maps for each channel are generated to reveal the interaction between channels, thereby enhancing the richness of feature information.
[0049] By compressing spatial-spectral information into a one-dimensional form through unit embedding, and using attention maps to deeply mine the connections between spectral information, efficient calculation of channel similarity is achieved. The process is shown in Equation (6):
[0050]
[0051] in, These represent the query, key, and value matrices, respectively. For relevant weights; The superscript T and represent batch matrix multiplication and batch transpose operations, respectively;
[0052] Using a multilayer perceptron to process the output of the unit embedding v * The embedding is reshaped into a three-dimensional feature form, as shown in equation (7):
[0053] M out * =Reshape(MLP(v * (7)
[0054] Among them, M out * This represents the reshaped 3D features.
[0055] Preferably, S3.2 specifically includes the following:
[0056] Clustering is performed on the query matrix to leverage the low-rank properties of the normal background to delve deeper into background features, obtaining n Cs Each category and its corresponding cluster center vector Then, key features of the background are extracted;
[0057] After determining the cluster centers, the augmented query matrix and augmented key matrix are constructed using the row concatenation method. Where m = k × k;
[0058] Calculate the attention distribution according to equation (8):
[0059]
[0060] Among them, A P The value of K′ is in the range [0,1]; T Represents the transpose of the augmented bond matrix; Scaling factor used to adjust the scale;
[0061] Attention Distribution A P pass Represents the global attention relationship between features, where A P_t =A P [1:m; 1:m]; through This represents the attention relationship between features and cluster centers, where A P_c =A P [1:m;m+1:m+n c ];
[0062] Using clustering algorithms in A P_c In the matrix, the probability p that each pixel belongs to the background is calculated. i And based on the p of all pixels i Values, select the top γ percentile pixels to construct the background weight map M. bw ; in M bw In this context, pixels with a value of "0" represent the foreground area, while other pixels are considered the background. The closer a pixel value is to "1", the higher its likelihood of being the background.
[0063] By calculating the local average of attention scores, the local energy of the attention distribution is evaluated, thereby creating a local attention energy map E. i,j The process is shown in equation (9):
[0064]
[0065] Where κ∈[i±τ,j±τ] represents the spatial local neighborhood at position (i,j), and the neighborhood contains N... a 1 pixel This represents the correlation between the pixel at position (i,j) and its neighboring pixels in the image;
[0066] Based on this, the weight ω of a pixel belonging to the foreground region is calculated. i,j The calculation formula is shown in equation (10):
[0067]
[0068] in, Represents the background weight map M bw The larger the pixel value, the greater the probability that the pixel belongs to the foreground.
[0069] Preferably, the calculation process of the training loss in S3.3 is as shown in equation (11):
[0070]
[0071] Among them, M bw Represents the confidence background weight map; ⊙ represents the dot product of the corresponding spatial locations; represents the normalization factor; S represents the true hyperspectral image cube. Represents the reconstructed hyperspectral image cube; the subscript F indicates the Frobenius norm.
[0072] Preferably, the judgment principle for anomaly detection in S3.3 is as follows:
[0073] If a pixel has both a high confidence foreground weight and a large recovery error, then the probability of that pixel being judged as an anomalous target also increases; the formula for calculating the anomalous score is shown in equation (12):
[0074]
[0075] Where ADs represent outlier scores; S i,j This represents the hyperspectral data to be detected; This represents reconstructed hyperspectral data; χ is a weighting coefficient used to adjust the relationship between the two; ω i,j This indicates the weight of a pixel belonging to the foreground region.
[0076] A hyperspectral anomaly detection system based on background confidence enhancement implicit self-supervised hyperspectral characterization includes:
[0077] The color space mapping generator module constructs color space mapping data for network training.
[0078] The spectral feature cube builder module provides real hyperspectral image (HSI) data required to guide network training;
[0079] Design an implicit self-supervised hyperspectral characterization anomaly detection module, specifically including:
[0080] Implicit self-supervised hyperspectral representation network unit is used to introduce implicit neural representation (INR) into the field of hyperspectral anomaly detection. After inputting training image samples, it combines a color space mapping generator and a spectral feature cube builder to reconstruct the background spectrum with high accuracy, recover the background spectral information from the training samples, and reconstruct the corresponding real hyperspectral image (HSI) cube. This realizes the mapping of samples from color space to spectral space, so as to reveal the deep connection between the appearance and spectrum of the sample and capture the high-level semantic information of normal targets.
[0081] The background credibility enhancement learning mechanism unit is used to obtain clean and high-quality normal samples when processing hyperspectral anomaly detection datasets, ensuring that the network is not affected by anomalous targets during the training phase.
[0082] The training loss and anomaly detection unit are designed to use the accuracy of background region reconstruction as the metric for the loss function during the training phase, and the root mean square error is used as the indicator to evaluate the quality difference between the reconstructed image and the real image. During the testing phase, the image is divided into non-intersecting spatial cubes, and anomaly detection is performed using self-supervised differential recovery features and the spatial distribution characteristics of anomalous targets.
[0083] 3. Beneficial effects
[0084] (1) Implicit mapping mechanism from visible light to hyperspectral light: This invention proposes a self-supervised mapping method from the visible light domain to the hyperspectral domain. By establishing a cross-domain self-representation model, the ability to extract high-order semantic features is significantly enhanced. This mechanism achieves the separation of background and abnormal targets in unlabeled scenarios, effectively alleviating the problems of strong dependence on explicit labels and lack of semantic information in existing methods.
[0085] (2) Construction of Implicit Neural Representation Network (IHSRN): This invention introduces an implicit neural network (INR) as the core modeling framework, combines spatial-spectral features for high-precision reconstruction, and improves feature representation ability through continuous function modeling. Compared with existing explicit reconstruction or pixel-level regression methods, this invention can effectively model long-distance background dependencies and improve the robustness of anomaly detection.
[0086] (3) Background credibility enhancement mechanism: This invention designs a cluster-driven background credibility weight estimation method, which automatically identifies high-credibility background regions during the training phase and dynamically adjusts them by constructing an attention distribution, thereby suppressing anomalous interference. Compared with existing methods that rely on fixed background priors or manual rules, this mechanism is more adaptable and data-driven, and is particularly suitable for anomaly detection in complex backgrounds.
[0087] In summary, compared with existing methods, this invention has stronger detection accuracy and higher robustness. It demonstrates innovation and technical superiority in modeling methods, information fusion mechanisms, and training strategies, and has broad practical value and promotion potential. Attached Figure Description
[0088] Figure 1 This is a flowchart of the anomaly detection method based on implicit self-supervised hyperspectral characterization with enhanced background credibility proposed in this invention;
[0089] Figure 2 This is a diagram of the detection results proposed in Embodiment 2 of the present invention. Detailed Implementation
[0090] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0091] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0092] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0093] This invention aims to solve two key technical challenges in hyperspectral anomaly detection (HAD):
[0094] First, existing methods generally lack the ability to model high-level semantic information, resulting in limited ability to identify abnormal targets in complex scenes;
[0095] Secondly, traditional models struggle to effectively characterize long-distance spectral-spatial correlations between background pixels, resulting in insufficient anomaly detection accuracy.
[0096] This invention proposes an implicit representation learning method that integrates a background credibility enhancement mechanism. By constructing an implicit mapping network from visible light images to hyperspectral images, it achieves fine distinction between background and anomalous targets without relying on manual annotation, thereby improving the accuracy and robustness of anomalous target detection in hyperspectral images. It has broad engineering application and research value.
[0097] The following description, in conjunction with relevant accompanying drawings and specific examples, illustrates the proposed method and system for anomaly detection based on implicit self-supervised hyperspectral characterization with enhanced background credibility.
[0098] Example 1:
[0099] Please see Figure 1 This invention proposes a method and system for detecting anomalies in hyperspectral characterization based on background credibility enhancement, which mainly consists of three parts: a color space mapping generator, a spectral feature cube builder, and an implicit self-supervised hyperspectral characterization anomaly detection module.
[0100] Part 1: Color Space Mapping Generator
[0101] The color space mapping generator aims to construct color space mapping data for network training. This is for RGB color images corresponding to hyperspectral images. Color space mapping data was obtained using a data augmentation method employing a k×k sliding window. This enhances the network's ability to extract spectral and spatial information. Among them, hsi m and hsi n Let represent the horizontal and vertical pixel counts of the image, respectively. k is the size of the color space mapping data block, i∈1,…,n are the index numbers of the color space mapping data, and n is the number of data blocks. The size of each block must satisfy k. <hsi m And k <hsi nThe generated color space mapping data lays the foundation for the network to extract spectral and spatial information. Simultaneously, the sliding window method can expand the sample set, providing necessary image samples for training the implicit self-supervised hyperspectral representation network. The specific process of the sliding window is shown in Equation 1.
[0102]
[0103] Among them, w i and h i Let represent the initial position index of the i-th color space mapping data, and p represent the step size of the sliding window.
[0104] Part Two: Spectral Feature Cube Constructor
[0105] The spectral feature cube builder is designed to provide the real hyperspectral image (HSI) data necessary for guiding network training. The implicitly self-supervised hyperspectral representation network maps the input color space data X... i Convert to the corresponding HSI cube During the process, the real HSI S i It plays a crucial role in guiding the network to optimize its path. However, HSI generally has a high spectral dimension. Too many spectral bands not only lead to a significant increase in computational load, but also the information redundancy between spectral bands may weaken the model's accuracy in mining the interrelationships between spectral bands, thus affecting the performance of self-supervised learning tasks.
[0106] Therefore, dimensionality reduction techniques were employed in the design of the spectral feature cube builder to process HSI, effectively reducing the number of spectral bands in the image. This was done when processing the raw HSI data. During the process, a method based on an optimized clustering framework is used to select the top n spectral bands that best represent the data characteristics. s Each spectral channel was used. Then, a sliding window technique was employed to generate a spectrum similar to the original data X. i The corresponding HSI cube The process of this sliding window operation can be referred to in formula (1).
[0107] Finally, the hyperspectral cube S was obtained. i It is used to guide the network to complete the self-supervised learning process, avoids the interference of redundant information on the model's learning effect, and enables the model to better grasp the mapping relationship between normal samples from appearance space and spectral space.
[0108] Part Three: Implicit Self-Supervised Hyperspectral Characterization Anomaly Detection Module
[0109] (1) Implicit self-supervised hyperspectral characterization network
[0110] The core of the proposed method is the Implicit Self-Supervised Hyperspectral Representation Network (ISSHRN). Implicit Neural Representation (INR) is a method that uses neural networks to parameterize signals as continuous functions. It can simulate the continuous nature of natural objects and accurately capture complex details by cleverly adjusting parameters. In the field of HSI processing, because HSI has continuous narrowband image data with high spectral resolution, it can be more accurately assumed that continuous functions can truly reflect the state of the image. Since the exact form of this continuous function cannot be determined, the concept of INR is introduced. This invention introduces INR into the field of hyperspectral anomaly detection for the first time. After inputting training image samples, the goal of ISSHRN is to reconstruct the background spectrum with high accuracy, that is, from the training sample X... i The background spectral information was recovered and the corresponding HSI cube was reconstructed. This process maps samples from color space to spectral space, revealing the deep connection between sample appearance and spectrum, and capturing high-level semantic information of normal targets. This enhances the network's ability to represent normal targets, providing a reconstruction effect that clearly distinguishes between normal and abnormal targets, thereby improving the identification ability and robustness of the anomaly detection model.
[0111] Specifically, the spatial features of the input samples are first extracted using a feature encoder. Then, using deep features M in and the number of target spectral bands n s As input, a spectral image that meets the target band number requirement is reconstructed, as shown in Equation (2).
[0112]
[0113] Despite spatial features M in While containing spatial information, the spectral information crucial for representing continuous-spectrum images was not fully extracted. Therefore, this study further enhances deep features using Spectral Feature Refinement Technique (SFRT) and Spatial-Spectral Synergistic Attention Mechanism (SSSAM). Building upon this, the number of target spectral bands is treated as the coordinates of a continuous implicit function, thus mapping deep features onto spectral intensity, achieving a precise projection from deep features to spectral intensity.
[0114] 1) Spectral Feature Refinement Techniques: Spectral feature refinement techniques aim to delve deeper into the correlation between space and spectrum, utilizing spectral interpolation to characterize the spatial-spectral interactions of features. For deep features... The analysis is conducted along both horizontal and vertical directions to uncover the intrinsic relationship between spatial and spectral dimensions, which are represented as follows: and Next, bilinear interpolation was used to construct the upsampled horizontal spectral features. and vertical spectral characteristics This enables richer feature representations, as shown in formula (3).
[0115]
[0116] Where BI(·) represents bilinear interpolation. The upsampled spectral features are converted into a feature map. By splicing features, a three-dimensional depth feature suitable for implicit neural representation is constructed, as shown in Equation (4).
[0117]
[0118] Among them, M out Conv represents the output characteristics after processing with spectral feature refinement techniques. 3D (·) and [·,·] represent a 3D convolutional layer with the LeakyReLU activation function and a splicing operation, respectively.
[0119] 2) Spatial-Spectral Co-attention Mechanism: The spatial-spectral co-attention mechanism focuses on achieving global attention in both spatial and spectral dimensions. It utilizes this mechanism to explore similarities between channels, thereby enhancing the comprehensiveness of spatial-spectral information. Unlike the standard Transformer approach, this invention uses a three-dimensional tensor as input. To adapt to the processing of the three-dimensional tensor, the three-dimensional features are converted into two-dimensional unit embeddings, which are then mapped through a multilayer perceptron, as shown in equation (5).
[0120] v = MLP(Reshape(M out (5)
[0121] In this context, MLP(·) and Reshape(·) represent the multilayer perceptron and the reshaping operation, respectively. This indicates the reshaped unit embedding.
[0122] In the attention computation process, the unit embeddings are first transformed into key embeddings and memory embeddings to depict the correspondence between deep features. Unlike the self-attention mechanism, this process reveals the interaction between channels by generating attention maps for each channel, thereby enhancing the richness of feature information. Since the unit embeddings compress spatial-spectral information into a one-dimensional form, the attention maps can deeply mine the connections between spectral information and achieve efficient computation of channel similarity, as shown in Equation (6).
[0123]
[0124] in, These represent the query, key, and value matrices required to construct the attention matrix. It refers to the relevant weights. The superscript T denotes batch matrix multiplication and batch transpose operations, respectively. Finally, the unit embedding v of the output is processed using a multilayer perceptron. * The embedding is then reshaped into a three-dimensional feature form, as shown in formula (7).
[0125] M out * =Reshape(MLP(v * (7)
[0126] (2) Background credibility enhancement learning mechanism
[0127] Obtaining clean and high-quality normal samples is extremely challenging when processing hyperspectral anomaly detection datasets. Ensuring that the interference of anomalous targets is eliminated during network training is the core of achieving self-supervised learning and a crucial factor that must be carefully considered during model construction. Based on the aforementioned analysis, a background credibility enhancement learning mechanism is designed to ensure that the network is not affected by anomalous targets during the training phase. This mechanism effectively improves the model's ability to learn features from normal samples, thereby enhancing overall detection performance.
[0128] To fully utilize the low-rank attributes of the normal background to delve deeper into background features, a clustering operation will be performed on the query matrix to obtain n. Cs Each category and its corresponding cluster center vector This effectively extracted the key features of the background.
[0129] After determining the cluster centers, the augmented query matrix and augmented key matrix were constructed using a row concatenation method. Where m = k × k. Then, calculate the attention distribution according to Formula 8:
[0130]
[0131] Among them, AP The value of is between [0,1], and K′T represents the transpose of the augmented bond matrix. This is the scaling factor used to adjust the proportions. Attention Distribution A P pass To reveal the global attention relationships between features, where A P_t =A P [1:m;1:m], and This reflects the attention relationship between features and cluster centers, where A P_c =A P [1:m;m+1:m+n c ].
[0132] To reduce the interference of abnormal targets on the optimization process, a background credibility enhancement learning mechanism was designed. Specifically, a clustering algorithm is used in A... P_c In the matrix, the probability p that each pixel belongs to the background is calculated. i And based on the p of all pixels i The values are selected, and the top γ percentile pixels are used to construct the background weight map M. bw In M bw In this model, pixels with a value of "0" represent the foreground region, while other pixels are considered background. The closer a pixel value is to "1", the higher its probability of being background. Furthermore, by calculating the local average of the attention scores, the local energy of the attention distribution is evaluated, thereby creating a local attention energy map E. i,j For details, please refer to formula (9):
[0133]
[0134] Where κ∈[i±τ,j±τ] represents the spatial local neighborhood at position (i,j), and the neighborhood contains N... a 1 pixel This represents the correlation between the pixel at position (i,j) and its neighboring pixels.
[0135] Based on this, the weight of a pixel belonging to the foreground is calculated. in Background weight map M bw The larger the pixel value, the greater the probability that the pixel belongs to the foreground.
[0136] (3) Training loss and anomaly detection
[0137] The training objective of the network is to achieve high-quality reconstruction of the background spectrum. To this end, the accuracy of background region reconstruction is used as the metric of the loss function during the training phase, and the root mean square error is selected as the indicator to evaluate the quality difference between the reconstructed image and the real image. The specific calculation process is detailed in formula (10).
[0138]
[0139] Among them, M bw The diagram represents the confidence background weight map, and ⊙ represents the dot product of the corresponding spatial locations. represents the normalization factor; S represents the true hyperspectral image cube. Represents the reconstructed hyperspectral image cube; the subscript F indicates the Frobenius norm.
[0140] During the testing phase, the image is divided into non-intersecting spatial cubes, and anomaly detection is performed using self-supervised differential recovery features and the spatial distribution characteristics of anomalous targets. Specifically, if a pixel has both a high confidence foreground weight and a large recovery error, then the probability of that pixel being identified as an anomalous target increases accordingly. Therefore, the anomaly score is calculated as shown in formula (11).
[0141]
[0142] Where ADs represent outlier scores; S i,j This represents the hyperspectral data to be detected; This represents reconstructed hyperspectral data; χ is a weighting coefficient used to adjust the relationship between the two; ω i,j This indicates the weight of a pixel belonging to the foreground region.
[0143] In summary, compared with existing hyperspectral anomaly detection technologies, this invention has stronger detection accuracy and higher robustness. First, it employs a cross-modal implicit mapping method based on self-supervised learning to mine potential hyperspectral representations from visible light data, effectively overcoming the limitations of traditional methods when labels are lacking or hyperspectral acquisition is restricted. Second, the proposed implicit neural representation reconstruction framework can jointly model spatial and spectral features, overcoming the problem of limited spatial receptive field in existing pixel-based or local structure-based modeling methods, and improving the ability to handle complex background interference.
[0144] Furthermore, this invention innovatively introduces a background credibility enhancement mechanism, adaptively adjusting the model's focus on different regions based on clustering characteristics, significantly reducing the false positive rate. This mechanism differs from existing methods that rely on manual prior knowledge or rule setting, possessing better generalization and data-driven characteristics. Therefore, this invention can more robustly identify real anomalies in unsupervised scenarios, and has broader practical application potential.
[0145] Example 2:
[0146] Based on Example 1, but with some differences, the present invention will now be described in conjunction with specific examples and accompanying drawings as follows: the method and system for anomaly detection based on background confidence enhancement implicit self-supervised hyperspectral characterization.
[0147] To comprehensively evaluate the hyperspectral anomaly detection performance of this invention, five representative real hyperspectral datasets were selected for experimental verification: Gulfport, Texas Coast-1, Texas Coast-2, Los Angeles, and San Diego. These datasets cover diverse land cover types and complex background environments, demonstrating both good representativeness and challenge.
[0148] Please see Figure 2 , Figure 2 The figures demonstrate typical detection results of the present invention in different scenarios. In the figures: the first row is the original pseudo-color image; the second row is a manually annotated reference image; and the third row is the detection image obtained by the method of the present invention.
[0149] As can be clearly observed from the figure, the present invention can effectively extract the target area and significantly suppress background interference under various complex background conditions. The detection results are highly consistent with the reference figure, with almost no obvious false alarms or missed detections, demonstrating excellent target salience and robustness.
[0150] Furthermore, the quantitative detection scores on the five datasets were 99.12%, 99.80%, 99.93%, 99.63%, and 99.57%, respectively, with an average score of 99.61%. The overall detection performance was stable and efficient, indicating that the method achieved a good balance between accurate detection and low false alarm rate.
[0151] In summary, combining Figure 2 It is evident that the anomaly detection method proposed in this invention exhibits significant advantages in terms of target enhancement, background suppression, and detection accuracy, demonstrating strong practicality, scalability, and significant progress compared to existing technologies.
[0152] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and its improved concept, should be covered within the scope of protection of the present invention.
Claims
1. A hyperspectral characterization anomaly detection method based on background confidence enhancement implicit self-supervised hyperspectral characterization, characterized in that, Includes the following steps: S1. Design a color space mapping generator to construct color space mapping data for network training; S2. Design a spectral feature cube builder to provide real hyperspectral image data required to guide network training; S3. Design an implicit self-supervised hyperspectral characterization anomaly detection module, specifically including: S3.1 Construct an implicit self-supervised hyperspectral representation network, introducing implicit neural representation into the field of hyperspectral anomaly detection. After inputting training image samples, the background spectrum is reconstructed with high precision by combining a color space mapping generator and a spectral feature cube builder. The background spectral information is recovered from the training samples, and the corresponding real hyperspectral image cube is reconstructed. This realizes the mapping of samples from color space to spectral space, so as to reveal the deep connection between the appearance and spectrum of the samples and capture the high-level semantic information of normal targets. S3.2 Design a background credibility enhancement learning mechanism to obtain pure and high-quality normal samples when processing hyperspectral anomaly detection datasets, ensuring that the network is not affected by abnormal targets during the training phase; S3.3 Design training loss and anomaly detection. During the training phase, the accuracy of background region reconstruction is used as the metric for the loss function, and the root mean square error is used as the indicator to evaluate the quality difference between the reconstructed image and the real image. During the testing phase, the image is divided into non-intersecting spatial cubes, and anomaly detection is performed using self-supervised differential recovery features and the spatial distribution characteristics of abnormal targets.
2. The anomaly detection method based on implicit self-supervised hyperspectral characterization with enhanced background confidence as described in claim 1, characterized in that, S1 specifically includes the following: For the RGB color image corresponding to hyperspectral image Color space mapping data was obtained using a data augmentation method employing a k×k sliding window. Enhance the network's ability to extract spectral and spatial information; among which, hsi m and hsi n The horizontal and vertical pixel counts of the image are respectively represented; k represents the size of the color space mapping data block; i∈1,…,n represents the index number of the color space mapping data, and n is the number of data blocks. The size of the block must satisfy k<hsi m And k < hsi n ; The sample set is expanded using a sliding window method to provide image samples for training an implicit self-supervised hyperspectral representation network; the specific process of the sliding window is shown in equation (1): Among them, w i and h i represents the initial position index of the i-th color space mapping data; p represents the step size of the sliding window.
3. The anomaly detection method based on background confidence enhancement implicit self-supervised hyperspectral characterization as described in claim 2, characterized in that, S2 specifically includes the following: Dimensionality reduction techniques are used to process real hyperspectral images, reducing the number of spectral bands in the images; this is applied to the processing of original real hyperspectral image data. During the process, a method based on an optimized clustering framework is used to select the top n spectral bands that best represent the data characteristics. s One spectral channel; Generating data X using the sliding window technique, compared to the original data X. i Corresponding real hyperspectral image cube The resulting real hyperspectral image cube S i It is used to guide the network to complete the self-supervised learning process, avoid the interference of redundant information on the model's learning effect, and enable the model to master the mapping relationship between the appearance space and spectral space of normal samples.
4. The anomaly detection method based on implicit self-supervised hyperspectral characterization with enhanced background confidence as described in claim 3, characterized in that, S3.1 specifically includes the following: Deep features of the input samples are extracted using a feature encoder. Using the obtained deep features M in and the number of target spectral bands n s As input, a spectral image that meets the target band quantity requirement is reconstructed, and the process is shown in equation (2): in, A cube representing the reconstructed true hyperspectral image; Deep features are enhanced by spectral feature refinement technology and spatial-spectral collaborative attention mechanism. The number of target spectral bands is regarded as the coordinates of a continuous implicit function, and then the deep features are mapped to spectral intensity, so as to achieve accurate projection from deep features to spectral intensity.
5. The anomaly detection method based on background confidence enhancement implicit self-supervised hyperspectral characterization as described in claim 4, characterized in that, The spectral feature refinement technique is used to analyze the correlation between space and spectrum. It uses spectral interpolation to characterize the space-spectral interaction of features, specifically including: For deep features The analysis is conducted along both horizontal and vertical directions to uncover the intrinsic relationship between spatial and spectral dimensions, which are represented as follows: and Using bilinear interpolation, the horizontal spectral features of the upsampled data are constructed respectively. and vertical spectral characteristics The process, represented by rich features, is shown in equation (3): Where BI(·) represents bilinear interpolation; Convert the upsampled spectral features into a feature map. By fusing features through splicing operations, a three-dimensional depth feature suitable for implicit neural representation is constructed, as shown in Equation (4): Among them, M out This represents the output characteristics after processing using spectral feature refinement techniques; Conv 3D (·) and [·,·] represent a 3D convolutional layer with the LeakyReLU activation function and a splicing operation, respectively.
6. The anomaly detection method based on background confidence enhancement implicit self-supervised hyperspectral characterization as described in claim 5, characterized in that, The aforementioned spatial-spectral collaborative attention mechanism focuses on achieving global attention across both spatial and spectral dimensions. It utilizes spatial-spectral attention mechanisms to analyze similarities between various channels, thereby enhancing the comprehensiveness of spatial-spectral information. Specifically, it includes: Using a three-dimensional tensor as input, the three-dimensional features are converted into two-dimensional unit embeddings and mapped through a multilayer perceptron, as shown in equation (5): v=MLP(Reshape(M out )) (5) Where MLP(·) and Reshape(·) represent the multilayer perceptron and reshaping operation, respectively; This indicates the reshaped unit embedding; During the attention computation process, unit embeddings are transformed into key embeddings and memory embeddings to depict the correspondence between deep features. Attention maps for each channel are generated to reveal the interaction between channels, thereby enhancing the richness of feature information. By compressing spatial-spectral information into a one-dimensional form through unit embedding, and using attention maps to deeply mine the connections between spectral information, efficient calculation of channel similarity is achieved. The process is shown in Equation (6): in, These represent the query, key, and value matrices, respectively. For relevant weights; The superscript T and represent batch matrix multiplication and batch transpose operations, respectively; Using a multilayer perceptron to process the output of the unit embedding v * The embedding is reshaped into a three-dimensional feature form, as shown in equation (7): M out * =Reshape(MLP(v * )) (7) Among them, M out * This represents the reshaped 3D features.
7. The anomaly detection method based on implicit self-supervised hyperspectral characterization with enhanced background confidence as described in claim 1, characterized in that, S3.2 specifically includes the following: Clustering is performed on the query matrix to leverage the low-rank properties of the normal background to delve deeper into background features, obtaining n Cs Each category and its corresponding cluster center vector Then, key features of the background are extracted; After determining the cluster centers, the augmented query matrix and augmented key matrix are constructed using the row concatenation method. Where m = k × k; Calculate the attention distribution according to equation (8): Among them, A P The value of is between [0,1]; K′T represents the transpose of the augmented bond matrix; Scaling factor used to adjust the scale; Attention Distribution A P pass Represents the global attention relationship between features, where A P_t =A P [1:m; 1:m]; through This represents the attention relationship between features and cluster centers, where A P_c =A P [1:m;m+1:m+n c ]; Using clustering algorithms in A P_c In the matrix, the probability p that each pixel belongs to the background is calculated. i And based on the p of all pixels i Values, select the top γ percentile pixels to construct the background weight map M. bw ; in M bw In this context, pixels with a value of "0" represent the foreground area, while other pixels are considered the background. The closer a pixel value is to "1", the higher its likelihood of being used as the background. By calculating the local average of attention scores, the local energy of the attention distribution is evaluated, thereby creating a local attention energy map E. i,j The process is shown in equation (9): Where κ∈[i±τ,j±τ] represents the spatial local neighborhood at position (i,j), and the neighborhood contains N... a 1 pixel This represents the correlation between the pixel at position (i,j) and its neighboring pixels in the image; Based on this, the weight ω of a pixel belonging to the foreground region is calculated. i,j The calculation formula is shown in equation (10): in, Represents the background weight map M bw The larger the pixel value, the greater the probability that the pixel belongs to the foreground.
8. The anomaly detection method based on implicit self-supervised hyperspectral characterization with enhanced background confidence as described in claim 7, characterized in that, The calculation process of the training loss described in S3.3 is shown in equation (11): Among them, M bw Represents the confidence background weight map; ⊙ represents the dot product of the corresponding spatial locations; represents the normalization factor; S represents the true hyperspectral image cube. Represents the reconstructed hyperspectral image cube; the subscript F indicates the Frobenius norm.
9. The anomaly detection method based on implicit self-supervised hyperspectral characterization with enhanced background confidence as described in claim 8, characterized in that, The judgment principle for anomaly detection described in S3.3 is as follows: If a pixel has both a high confidence foreground weight and a large recovery error, then the probability of that pixel being judged as an anomalous target also increases; the formula for calculating the anomalous score is shown in equation (12): Where ADs represent outlier scores; S i,j This represents the hyperspectral data to be detected; This represents reconstructed hyperspectral data; χ is a weighting coefficient used to adjust the relationship between the two; ω i,j This indicates the weight of a pixel belonging to the foreground region.
10. The background confidence enhancement implicit self-supervised hyperspectral characterization anomaly detection system based on the method described in any one of claims 1-9, characterized in that, include: The color space mapping generator module constructs color space mapping data for network training. The spectral feature cube builder module provides real hyperspectral image data required to guide network training; Design an implicit self-supervised hyperspectral characterization anomaly detection module, specifically including: Implicit self-supervised hyperspectral representation network unit is used to introduce implicit neural representation into the field of hyperspectral anomaly detection. After inputting training image samples, it combines a color space mapping generator and a spectral feature cube builder to reconstruct the background spectrum with high precision, recover the background spectral information from the training samples, and reconstruct the corresponding real hyperspectral image cube, realizing the mapping of samples from color space to spectral space, so as to reveal the deep connection between the appearance and spectrum of the sample and capture the high-level semantic information of normal targets. The background credibility enhancement learning mechanism unit is used to obtain clean and high-quality normal samples when processing hyperspectral anomaly detection datasets, ensuring that the network is not affected by anomalous targets during the training phase. The training loss and anomaly detection unit are designed to use the accuracy of background region reconstruction as the metric for the loss function during the training phase, and the root mean square error is used as the indicator to evaluate the quality difference between the reconstructed image and the real image. During the testing phase, the image is divided into non-intersecting spatial cubes, and anomaly detection is performed using self-supervised differential recovery features and the spatial distribution characteristics of anomalous targets.
Citation Information
Cited By
Hyperspectral anomaly detection and classification method and system for traditional Chinese medicinal materials and computer equipment
CN122135115A